Browse Source

fix(headless): fail closed on malformed preset records and unroundtrippable args

Address the ds-review-bot v6 review on PR #3849:

- A persisted `agent-preset/selected` record whose `data.agentPreset` is not a
  non-empty string used to read as "no preset", letting the run continue under
  this bundle's composition. It now fails closed.
- A live identity now has to be known to `sessionPersistence` (`stat`), since
  an Agent registered only in memory has no write handle and would flush
  nothing.
- Tool arguments that JSON cannot round-trip (an overflowing number such as
  `1e400`) keep their raw text instead of the `null` JSON.stringify reports.
lsdsjy 2 tuần trước cách đây
mục cha
commit
7a4a55f30b

+ 2 - 2
.agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.md
-2026-09-09-headless-machine-readable-run-surface.md: 9e0e6165b195de8dc96dc61512a56af7164c8bf6
-2026-09-09-headless-machine-readable-run-surface.zh.md: 18cf93a13d7dc03a09c70d14d352f8f5f690e025
+2026-09-09-headless-machine-readable-run-surface.md: 7bcca9226a8620d1eae83ecb269f680e6d1c54d4
+2026-09-09-headless-machine-readable-run-surface.zh.md: ce7047a4113c7ff794a3cbb88976c4760764ca2f

+ 3 - 3
.agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.md

@@ -54,7 +54,7 @@ Projection rules:
 - Text and reasoning are projected only from a committed `assistant/message`, never from live attempt deltas. A retried or discarded attempt appends `assistant/attempt`, which the projection ignores, so the stream never carries content the durable log does not contain ([publish state only at its commit point](../../../../packages/AGENTS.md)).
 - Each committed content block becomes exactly one `text` or `thinking` event in content order; `tool-call` blocks are not projected because the `tool/call` event owns them. `user/message` echoes and internal session events (title, model selection, projection, checkpoint, goal, subagent) are not projected.
 - A `tool/result` is projected only when its `surfaceOp` is `append`. A compaction replacement of an older result is history, and projecting it would emit a call id with no matching `tool_call`.
-- Every projected string and object key is bounded at 8 KiB, an event with a cut value carries `truncated: true`, and one serialized event line is bounded at 32 KiB — an over-long event keeps its scalar fields, drops structured ones, and at the extreme reduces to `type` and `truncated`, while a payload nested 64 levels or deeper is cut at that depth so no legal input can overflow the bounding recursion. This includes the process-level `error` event; a literal `__proto__` argument key is copied as data rather than through the inherited setter, and an empty tool-argument string projects as `{}` to match the executor. The terminal `final` event is deliberately unbounded: it carries the same lossless answer the default mode prints.
+- Every projected string and object key is bounded at 8 KiB, an event with a cut value carries `truncated: true`, and one serialized event line is bounded at 32 KiB — an over-long event keeps its scalar fields, drops structured ones, and at the extreme reduces to `type` and `truncated`, while a payload nested 64 levels or deeper is cut at that depth so no legal input can overflow the bounding recursion. This includes the process-level `error` event; a literal `__proto__` argument key is copied as data rather than through the inherited setter, an empty tool-argument string projects as `{}` to match the executor, and arguments JSON cannot round-trip (an overflowing number such as `1e400`) keep their raw text instead of the `null` serialization would report. The terminal `final` event is deliberately unbounded: it carries the same lossless answer the default mode prints.
 - Text and reasoning arrive when the step commits, not per token; default-mode stderr reasoning remains the only live text channel. A turn that fails in-turn still ends the stream with `final` and no `error` event, so a supervisor classifies that run from the exit code and the `turn_end` reason even when the stream is well formed.
 - `usage` appears on `step_end`, matching the token accounting a provider reports per step.
 - Raw session events stay out of scope. A debug escape hatch can be added later without changing this vocabulary.
@@ -65,7 +65,7 @@ The runtime owns identity. A run without `--session-id` mints `session-<uuid>` a
 
 `--session-id <id>` is adopt-or-create: observe the persisted session, resume it when it exists, create it otherwise. Create-only would fail the second run, because the JSONL store rejects an existing log id ([session persistence](../../implemented/architecture/2026-06-14-session-persistence.md)). The id is opaque, so the runner validates non-emptiness on the trimmed value and passes the caller's exact string through, whitespace included.
 
-Adoption compares the persisted session's recorded cwd with the process cwd, since sessions are organized per project directory ([project session directories](../../implemented/architecture/2026-07-24-project-session-directories.md)). A mismatch exits 1 with a `dsh:` diagnostic instead of silently continuing a conversation rooted elsewhere, and a session that recorded no cwd is rejected for the same reason. A session running under an agent preset is rejected because this bundle composes no preset roster: resuming it here would run it under the headless tools and prompts instead of the composition its log records. The check reads the preset the log currently records — the creation header advanced by any `agent-preset/selected` event — because a blank session may switch preset after creation while the header stays a creation fact. A session linked to a parent or subagent — including a user fork — is rejected. All checks run when a live Agent already holds the requested id, so a live identity cannot bypass them. Two live processes cannot write one id; the store's write lease already rejects the second writer. The runner reads the observation through the composed `sessionQuery` service and fails loudly when `--session-id` is requested without it, or when the requested identity would lack the `sessionPersistence` service that makes it durable.
+Adoption compares the persisted session's recorded cwd with the process cwd, since sessions are organized per project directory ([project session directories](../../implemented/architecture/2026-07-24-project-session-directories.md)). A mismatch exits 1 with a `dsh:` diagnostic instead of silently continuing a conversation rooted elsewhere, and a session that recorded no cwd is rejected for the same reason. A session running under an agent preset is rejected because this bundle composes no preset roster: resuming it here would run it under the headless tools and prompts instead of the composition its log records. The check reads the preset the log currently records — the creation header advanced by any `agent-preset/selected` event — because a blank session may switch preset after creation while the header stays a creation fact, and a malformed selection record fails closed rather than reading as no preset. A session linked to a parent or subagent — including a user fork — is rejected. All checks run when a live Agent already holds the requested id, so a live identity cannot bypass them. Two live processes cannot write one id; the store's write lease already rejects the second writer. The runner reads the observation through the composed `sessionQuery` service and fails loudly when `--session-id` is requested without it, or when the requested identity would lack the `sessionPersistence` service that makes it durable; a live identity must also carry a stored record, because one registered only in memory would flush nothing.
 
 ## Consequences
 
@@ -74,7 +74,7 @@ What landed: `src/startup.ts` parses `--json` and `--session-id <id>`, treats an
 - Default mode is unchanged: a text-only run writes one final assistant line to stdout and nothing to stderr, and exit status still follows the terminal reason.
 - `--json` stdout parses line by line as JSON, starts with `session`, ends with `final`, and contains no plain text. Stderr carries no reasoning in this mode.
 - A step that retries publishes `text` and `thinking` only for the attempt that commits, so a discarded attempt leaves no trace in the stream.
-- Two consecutive runs with the same `--session-id` share history. A run whose cwd differs from the persisted session, that recorded no cwd, that is a subagent or forked session, or that runs under an agent preset, exits 1 with a diagnostic, whether the identity is live or persisted.
+- Two consecutive runs with the same `--session-id` share history. A run whose cwd differs from the persisted session, that recorded no cwd, that is a subagent or forked session, that runs under an agent preset, that carries a malformed preset record, or whose live identity has no stored record, exits 1 with a diagnostic, whether the identity is live or persisted.
 - A piped task with no positional task is honored, a whitespace-only positional is rejected instead of consuming the pipe, and an interactive invocation without a task still fails with the usage error.
 - Unit coverage lands in `packages/bundle/headless/tests/startup.spec.ts`, `tests/headless.spec.ts`, and `tests/json-stream.spec.ts`. The product headless profile expectation test in `apps/cli/tests/profiles/headless/tests/headless.expected.e2e.ts` covers both output modes end to end.
 

+ 3 - 3
.agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.zh.md

@@ -54,7 +54,7 @@ dsh --profile headless [--json] [--session-id <id>] [<task>... | -]
 - 文本与推理只从已提交的 `assistant/message` 投影,绝不来自实时的 attempt 增量。被重试或丢弃的 attempt 会追加 `assistant/attempt`,投影直接忽略,因此事件流永远不会承载持久化日志中不存在的内容([只在提交点发布状态](../../../../packages/AGENTS.md))。
 - 每个已提交的内容块按内容顺序变成恰好一条 `text` 或 `thinking` 事件;`tool-call` 块不投影,因为 `tool/call` 事件已经拥有它。`user/message` 回显和内部会话事件(标题、模型选择、投影、检查点、目标、子 agent)都不投影。
 - `tool/result` 仅在其 `surfaceOp` 为 `append` 时投影。压缩对旧结果的替换属于历史,投影它会产生没有对应 `tool_call` 的 call id。
-- 每个被投影的字符串与对象键都限制在 8 KiB;被截断的事件带 `truncated: true`,单条序列化事件行限制在 32 KiB——超长事件保留标量字段、丢弃结构化字段,极端情况下只剩 `type` 与 `truncated`,嵌套达到 64 层及以上的负载会在该深度被截断,因此任何合法输入都不会让限界递归溢出。进程级 `error` 事件同样受限;字面量 `__proto__` 参数键会作为数据复制,而不经过继承的 setter;空工具参数字符串会投影为 `{}`,与执行器保持一致。终止 `final` 事件刻意不做限长:它承载与默认模式相同的无损答案。
+- 每个被投影的字符串与对象键都限制在 8 KiB;被截断的事件带 `truncated: true`,单条序列化事件行限制在 32 KiB——超长事件保留标量字段、丢弃结构化字段,极端情况下只剩 `type` 与 `truncated`,嵌套达到 64 层及以上的负载会在该深度被截断,因此任何合法输入都不会让限界递归溢出。进程级 `error` 事件同样受限;字面量 `__proto__` 参数键会作为数据复制,而不经过继承的 setter;空工具参数字符串会投影为 `{}`,与执行器保持一致;JSON 无法往返的参数(例如溢出为 `Infinity` 的 `1e400`)保留原始文本,而不是 `null` 序列化结果。终止 `final` 事件刻意不做限长:它承载与默认模式相同的无损答案。
 - 文本与推理在步骤提交时到达,而不是逐 token 到达;默认模式的 stderr 推理仍是唯一的实时文本通道。轮次内失败的运行仍以 `final` 结束且没有 `error` 事件,因此即使事件流格式良好,监督进程也要用退出码与 `turn_end` 原因来分类该次运行。
 - `usage` 出现在 `step_end` 上,对应 provider 每步上报的 token 计量。
 - 原始会话事件不在范围内。调试用的逃生口可以以后再加,不必改动这套词汇表。
@@ -65,7 +65,7 @@ dsh --profile headless [--json] [--session-id <id>] [<task>... | -]
 
 `--session-id <id>` 是采用或创建:先观察持久化会话,存在就 resume,不存在就 create。只创建会让第二次运行失败,因为 JSONL 存储拒绝已存在的日志 id(见 [session persistence](../../implemented/architecture/2026-06-14-session-persistence.zh.md))。标识是不透明的,因此 runner 只在 trim 后的值上校验非空,并把调用方的原始字符串(含空白字符)原样传下去。
 
-采用时会比较持久化会话记录的 cwd 与进程 cwd,因为会话按项目目录组织(见 [project session directories](../../implemented/architecture/2026-07-24-project-session-directories.zh.md))。不一致时以 `dsh:` 诊断退出 1,而不是静默续接一个根目录在别处的会话;未记录 cwd 的会话出于同样理由被拒绝。运行在 agent preset 下的会话被拒绝,因为本 bundle 不组合任何 preset roster:在这里 resume 它,会用 headless 的工具与提示词运行它,而不是它日志当前记录的组合。该检查读取日志当前记录的 preset——创建 header 再叠加任何 `agent-preset/selected` 事件——因为空白会话可能在创建后切换 preset,而 header 始终只是创建事实。带父会话或子 agent 关联的会话——包括用户 fork 出的会话——被拒绝。当某个存活 Agent 已经持有请求的 id 时,上述检查全部执行,因此存活身份无法绕过它们。两个存活进程不能写同一个 id;存储的写租约已经会拒绝第二个写入者。runner 通过已组合的 `sessionQuery` 服务读取观察结果,并在请求 `--session-id` 却没有该服务时显式失败;若所请求的身份缺少让它持久化的 `sessionPersistence` 服务,同样显式失败。
+采用时会比较持久化会话记录的 cwd 与进程 cwd,因为会话按项目目录组织(见 [project session directories](../../implemented/architecture/2026-07-24-project-session-directories.zh.md))。不一致时以 `dsh:` 诊断退出 1,而不是静默续接一个根目录在别处的会话;未记录 cwd 的会话出于同样理由被拒绝。运行在 agent preset 下的会话被拒绝,因为本 bundle 不组合任何 preset roster:在这里 resume 它,会用 headless 的工具与提示词运行它,而不是它日志当前记录的组合。该检查读取日志当前记录的 preset——创建 header 再叠加任何 `agent-preset/selected` 事件——因为空白会话可能在创建后切换 preset,而 header 始终只是创建事实;畸形的选择记录会失败关闭,而不会读成「无 preset」。带父会话或子 agent 关联的会话——包括用户 fork 出的会话——被拒绝。当某个存活 Agent 已经持有请求的 id 时,上述检查全部执行,因此存活身份无法绕过它们。两个存活进程不能写同一个 id;存储的写租约已经会拒绝第二个写入者。runner 通过已组合的 `sessionQuery` 服务读取观察结果,并在请求 `--session-id` 却没有该服务时显式失败;若所请求的身份缺少让它持久化的 `sessionPersistence` 服务,同样显式失败;存活身份还必须已有持久化记录,因为仅注册在内存中的身份不会写入任何内容。
 
 ## 后果
 
@@ -74,7 +74,7 @@ dsh --profile headless [--json] [--session-id <id>] [<task>... | -]
 - 默认模式不变:纯文本运行向 stdout 写一行最终助手消息、stderr 无输出,退出码仍跟随终端原因。
 - `--json` 的 stdout 逐行可解析为 JSON,以 `session` 开头、以 `final` 结尾,不含纯文本。该模式下 stderr 不承载推理。
 - 发生重试的步骤只为最终提交的 attempt 发布 `text` 与 `thinking`,因此被丢弃的 attempt 不会在事件流中留下任何痕迹。
-- 两次连续的相同 `--session-id` 运行共享历史。cwd 不一致、未记录 cwd、属于子 agent 或 fork 会话,或运行在 agent preset 下的运行都以诊断退出 1,无论身份是存活还是持久化的。
+- 两次连续的相同 `--session-id` 运行共享历史。cwd 不一致、未记录 cwd、属于子 agent 或 fork 会话、运行在 agent preset 下、preset 记录畸形,或存活身份没有持久化记录的运行都以诊断退出 1,无论身份是存活还是持久化的。
 - 无位置参数但 stdin 有管道输入时任务被采纳,只有空白的位置参数会被拒绝而不会消费管道,交互式无任务调用仍以用法错误失败。
 - 单元覆盖落在 `packages/bundle/headless/tests/startup.spec.ts`、`tests/headless.spec.ts` 与 `tests/json-stream.spec.ts`。`apps/cli/tests/profiles/headless/tests/headless.expected.e2e.ts` 的产品 headless profile 期望测试端到端覆盖两种输出模式。
 

+ 2 - 2
packages/bundle/headless/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/bundle/headless/README.md
-README.md: 138da22830a7ce93635418912266b816540f9db1
-README.zh.md: 3388ea5e729d4e4aeb1d434427768d05165bb22a
+README.md: c21084314349471258b64584d00266449dc57c88
+README.zh.md: 63129a4803b2dcde24866fb7a112ecce84fc1b62

+ 3 - 3
packages/bundle/headless/README.md

@@ -51,11 +51,11 @@ The generated [configuration catalog](../../../docs/config-catalog.md#deepseek-a
 
 ### Choosing the session identity
 
-Every invocation defaults to a fresh `session-<uuid>` identity. Pass `--session-id <id>` to name it yourself: the runner adopts the persisted Session with that id when one exists, and creates it otherwise. Both paths require the composed `sessionPersistence` service, so a profile that omits it fails loudly instead of returning an id whose history dies with the process. The identity is opaque, so the exact string is used, whitespace included. Adoption is scoped to the current working directory and refuses a Session that is a subagent or forked session, that recorded no working directory, or that runs under an agent preset this profile does not compose — the check reads the preset the Session log currently records, so a Session that switched preset while blank is rejected too. A supervisor therefore cannot silently drive someone else's conversation under a different composition; any mismatch fails before the task runs.
+Every invocation defaults to a fresh `session-<uuid>` identity. Pass `--session-id <id>` to name it yourself: the runner adopts the persisted Session with that id when one exists, and creates it otherwise. Both paths require the composed `sessionPersistence` service, so a profile that omits it fails loudly instead of returning an id whose history dies with the process; a live identity must also carry a stored record, because an Agent registered only in memory would flush nothing. The identity is opaque, so the exact string is used, whitespace included. Adoption is scoped to the current working directory and refuses a Session that is a subagent or forked session, that recorded no working directory, that runs under an agent preset this profile does not compose, or whose preset record is malformed — the check reads the preset the Session log currently records, so a Session that switched preset while blank is rejected too. A supervisor therefore cannot silently drive someone else's conversation under a different composition; any mismatch fails before the task runs.
 
 ### Machine-readable output
 
-`--json` replaces the final-text stdout line with a newline-delimited JSON event stream, while stderr keeps only the `dsh:` diagnostics. The stream opens with `session` (carrying the identity the run used) and closes with `final`, and carries `status`, `text`, `thinking`, `tool_call`, and `tool_result` events in between. `text` and `thinking` are projected from committed assistant messages, so a retried or discarded attempt never reaches the stream; they arrive when the step commits, not per token, and default-mode stderr reasoning remains the only live text channel. The terminal `final` event carries the same lossless answer as the default mode and is not capped; every other string and object key is capped at 8 KiB and flagged with `truncated`, and one event line is capped at 32 KiB — an over-long event keeps its scalar fields, drops structured ones, and at the extreme reduces to `type` and `truncated`, while a payload nested 64 levels or deeper is cut at that depth. An empty tool-argument string projects as `{}`, matching what the executor runs. A process-level failure outside a turn writes an `error` event and ends the stream without `final`, in addition to the `dsh:` stderr line. A turn that fails in-turn still ends with a `final` event (often empty) and no `error` event, so a well-formed stream can still describe a failed run: treat exit code 1 and the `turn_end` reason as the failure signal.
+`--json` replaces the final-text stdout line with a newline-delimited JSON event stream, while stderr keeps only the `dsh:` diagnostics. The stream opens with `session` (carrying the identity the run used) and closes with `final`, and carries `status`, `text`, `thinking`, `tool_call`, and `tool_result` events in between. `text` and `thinking` are projected from committed assistant messages, so a retried or discarded attempt never reaches the stream; they arrive when the step commits, not per token, and default-mode stderr reasoning remains the only live text channel. The terminal `final` event carries the same lossless answer as the default mode and is not capped; every other string and object key is capped at 8 KiB and flagged with `truncated`, and one event line is capped at 32 KiB — an over-long event keeps its scalar fields, drops structured ones, and at the extreme reduces to `type` and `truncated`, while a payload nested 64 levels or deeper is cut at that depth. An empty tool-argument string projects as `{}`, matching what the executor runs, while arguments JSON cannot round-trip — an overflowing number such as `1e400` — keep their raw text instead of the `null` serialization would report. A process-level failure outside a turn writes an `error` event and ends the stream without `final`, in addition to the `dsh:` stderr line. A turn that fails in-turn still ends with a `final` event (often empty) and no `error` event, so a well-formed stream can still describe a failed run: treat exit code 1 and the `turn_end` reason as the failure signal.
 
 ### When to use it
 
@@ -142,7 +142,7 @@ These limits tell you when headless does not fit and what it needs from the `dsh
 - **No pre-token heartbeat** — in default mode stderr stays silent until the provider emits a non-empty reasoning delta; a delayed first token exposes no earlier progress signal.
 - **Reasoning enters stderr logs** — in default mode, redirection and supervisors may retain substantially more and potentially sensitive model output; route stderr to a controlled sink when needed.
 - **Default stdout carries only the final answer** — a run without an assistant message prints an empty stdout line and exits 1; intermediate tool output is not printed unless you opt into `--json`.
-- **Adoption is cwd-, ownership-, and preset-scoped** — `--session-id` refuses a Session recorded in another working directory, one that recorded no working directory, one that is a subagent or forked session, or one that runs under an agent preset this profile does not compose, and requires the composed Session query and persistence services.
+- **Adoption is cwd-, ownership-, and preset-scoped** — `--session-id` refuses a Session recorded in another working directory, one that recorded no working directory, one that is a subagent or forked session, or one that runs under an agent preset this profile does not compose or whose preset record is malformed, and requires the composed Session query and persistence services plus a stored record for a live identity.
 - **The event stream is a projection, not the log** — `--json` caps every string except the terminal `final` at 8 KiB and omits events the projection does not model, so it is not a lossless copy of the Session log.
 
 <a id="dev-note"></a>

+ 3 - 3
packages/bundle/headless/README.zh.md

@@ -51,11 +51,11 @@ agent(智能体)会完成该任务,把提供方的每个非空推理增量
 
 ### 选择 Session 标识
 
-每次调用默认使用全新的 `session-<uuid>` 标识。传入 `--session-id <id>` 可自行命名:该 id 对应的持久化 Session 存在时 runner 会沿用,否则创建;两条路径都要求已组合 `sessionPersistence` 服务,因此缺少该服务的 profile 会显式失败,而不会返回一个历史随进程消失的 id。标识是不透明的,因此会原样使用调用方给出的字符串,包括空白字符。沿用被限定在当前工作目录内,并拒绝子 agent 或 fork 会话、未记录工作目录的会话,以及运行在本 profile 不组合的 agent preset 下的会话——该检查读取 Session 日志当前记录的 preset,因此在空白期切换过 preset 的会话同样会被拒绝。因此监督进程无法在另一套组合下悄悄驱动他人的会话;任一不匹配都会在任务运行前失败。
+每次调用默认使用全新的 `session-<uuid>` 标识。传入 `--session-id <id>` 可自行命名:该 id 对应的持久化 Session 存在时 runner 会沿用,否则创建;两条路径都要求已组合 `sessionPersistence` 服务,因此缺少该服务的 profile 会显式失败,而不会返回一个历史随进程消失的 id;存活身份还必须已有持久化记录,因为仅注册在内存中的 Agent 不会写入任何内容。标识是不透明的,因此会原样使用调用方给出的字符串,包括空白字符。沿用被限定在当前工作目录内,并拒绝子 agent 或 fork 会话、未记录工作目录的会话、运行在本 profile 不组合的 agent preset 下的会话,以及 preset 记录畸形的会话——该检查读取 Session 日志当前记录的 preset,因此在空白期切换过 preset 的会话同样会被拒绝。因此监督进程无法在另一套组合下悄悄驱动他人的会话;任一不匹配都会在任务运行前失败。
 
 ### 机器可读输出
 
-`--json` 用按行 JSON 事件流取代 stdout 的最终文本行,stderr 仅保留 `dsh:` 诊断信息。事件流以 `session`(携带本次运行使用的标识)开头、以 `final` 结尾,其间为 `status`、`text`、`thinking`、`tool_call` 与 `tool_result` 事件。`text` 与 `thinking` 只从已提交的 assistant 消息投影,因此被重试或丢弃的尝试不会进入事件流;它们在步骤提交时到达,而不是逐 token 到达,默认模式的 stderr 推理仍是唯一的实时文本通道。终止 `final` 事件携带与默认模式相同的无损答案,不做限长;其他每个字符串与对象键上限为 8 KiB,超出时标记 `truncated`,单条事件行上限为 32 KiB——超长事件保留标量字段、丢弃结构化字段,极端情况下只剩 `type` 与 `truncated`,嵌套达到 64 层及以上的负载会在该深度被截断。空工具参数字符串会投影为 `{}`,与执行器实际运行的值一致。轮次之外的进程级失败会写出 `error` 事件并在没有 `final` 的情况下结束事件流,同时向 stderr 写入 `dsh:` 行。轮次内失败的运行仍会以 `final` 事件(通常为空)结束且没有 `error` 事件,因此格式良好的事件流也可能描述一次失败的运行:请把退出码 1 与 `turn_end` 原因作为失败信号。
+`--json` 用按行 JSON 事件流取代 stdout 的最终文本行,stderr 仅保留 `dsh:` 诊断信息。事件流以 `session`(携带本次运行使用的标识)开头、以 `final` 结尾,其间为 `status`、`text`、`thinking`、`tool_call` 与 `tool_result` 事件。`text` 与 `thinking` 只从已提交的 assistant 消息投影,因此被重试或丢弃的尝试不会进入事件流;它们在步骤提交时到达,而不是逐 token 到达,默认模式的 stderr 推理仍是唯一的实时文本通道。终止 `final` 事件携带与默认模式相同的无损答案,不做限长;其他每个字符串与对象键上限为 8 KiB,超出时标记 `truncated`,单条事件行上限为 32 KiB——超长事件保留标量字段、丢弃结构化字段,极端情况下只剩 `type` 与 `truncated`,嵌套达到 64 层及以上的负载会在该深度被截断。空工具参数字符串会投影为 `{}`,与执行器实际运行的值一致;而 JSON 无法往返的参数——例如溢出为 `Infinity` 的数字 `1e400`——会保留原始文本,而不是 `null` 序列化结果。轮次之外的进程级失败会写出 `error` 事件并在没有 `final` 的情况下结束事件流,同时向 stderr 写入 `dsh:` 行。轮次内失败的运行仍会以 `final` 事件(通常为空)结束且没有 `error` 事件,因此格式良好的事件流也可能描述一次失败的运行:请把退出码 1 与 `turn_end` 原因作为失败信号。
 
 ### 何时使用
 
@@ -142,7 +142,7 @@ runner 不向请求前缀添加任何内容;它只是把一条用户消息驱
 - **首个 token 前没有心跳**——默认模式下,提供方发出第一个非空推理增量前 stderr 保持静默;延迟首个 token 的提供方不会更早给出进度信号。
 - **推理进入 stderr 日志**——默认模式下,重定向与监督进程可能保留更多且可能敏感的模型输出;需要时应把 stderr 路由到受控位置。
 - **默认 stdout 只承载最终答案**——没有 assistant 消息的运行向 stdout 打印空行并以 1 退出;中间工具输出不会打印,除非显式启用 `--json`。
-- **沿用受 cwd、归属与 preset 限制**——`--session-id` 会拒绝记录在其他工作目录、未记录工作目录、属于子 agent 或 fork 会话,或运行在本 profile 不组合的 agent preset 下的 Session,并要求已组合的 Session 查询与持久化服务。
+- **沿用受 cwd、归属与 preset 限制**——`--session-id` 会拒绝记录在其他工作目录、未记录工作目录、属于子 agent 或 fork 会话,或运行在本 profile 不组合的 agent preset 下的 Session、preset 记录畸形的 Session,并要求已组合的 Session 查询与持久化服务,且存活身份必须已有持久化记录。
 - **事件流是投影而非日志**——`--json` 除终止 `final` 外把每个字符串与对象键限制在 8 KiB,并省略投影未建模的事件,因此它不是 Session 日志的无损副本。
 
 <a id="dev-note"></a>

+ 20 - 5
packages/bundle/headless/src/index.ts

@@ -200,20 +200,27 @@ function* liveEvents(session: Session): Generator<SessionEvent> {
  * the last `agent-preset/selected` event. The header is only a creation fact;
  * the presets plugin reconstructs a session's composition from the projection.
  */
-function currentPreset(header: AdoptableHeader, events: Iterable<SessionEvent>): string | undefined {
+function currentPreset(header: AdoptableHeader, events: Iterable<SessionEvent>, sessionId: SessionId): string | undefined {
   let preset = header.agentPreset
   for (const event of events) {
     // Owned by dsh-agent-presets, which this bundle does not compose, so the
     // event is read structurally rather than through its module augmentation.
-    const candidate = event as unknown as { type: string; data: { agentPreset: string } }
-    if (candidate.type === 'agent-preset/selected') preset = candidate.data.agentPreset
+    const candidate = event as unknown as { type: string; data?: { agentPreset?: unknown } }
+    if (candidate.type !== 'agent-preset/selected') continue
+    const selected = candidate.data?.agentPreset
+    // A corrupt record must not read as "no preset": that would let the run
+    // continue under this bundle's composition instead of the recorded one.
+    if (typeof selected !== 'string' || selected === '') {
+      throw new Error(`session "${sessionId}" records a malformed agent-preset/selected event and cannot be adopted`)
+    }
+    preset = selected
   }
   return preset
 }
 
 /** Reject a Session the one-shot runner must not adopt. */
 function assertAdoptable(header: AdoptableHeader, events: Iterable<SessionEvent>, sessionId: SessionId): void {
-  const preset = currentPreset(header, events)
+  const preset = currentPreset(header, events, sessionId)
   if (preset !== undefined) {
     // This bundle composes no preset roster, so resuming the session here would
     // silently run it under the headless tools and prompts instead of the
@@ -254,13 +261,21 @@ async function resolveAgent(
   // caller a log a later process can continue. Without a durable log the run
   // would succeed, print the id, and still lose the whole history at exit, so
   // a miscomposed profile fails loud before either path.
-  if (ctx.get('sessionPersistence') === undefined) {
+  const persistence = ctx.get('sessionPersistence')
+  if (persistence === undefined) {
     throw new Error('headless --session-id requires the sessionPersistence service; the Session would not survive this process')
   }
   const live = agents.get(sessionId)
   if (live !== undefined) {
     // A live identity skips adoption, not the rules that make adoption safe.
     assertAdoptable(live.session.header, liveEvents(live.session), sessionId)
+    // The service can be mounted while this particular Agent was registered in
+    // memory (a custom factory or direct `agents.register`); the backend then
+    // holds no write handle for it and `session/flush` stores nothing. A stored
+    // record proves the id is actually persistence-backed.
+    if (await persistence.stat(sessionId) === undefined) {
+      throw new Error(`live session "${sessionId}" has no persisted record, so the one-shot runner cannot promise it survives this process`)
+    }
     return live
   }
   const query = ctx.get('sessionQuery')

+ 8 - 1
packages/bundle/headless/src/json-stream.ts

@@ -144,8 +144,15 @@ export function boundJsonLine(
 /** Parse raw tool-call arguments as the executor does: empty input is `{}`, invalid JSON stays text. */
 function parseArguments(raw: string): unknown {
   if (raw === '') return {}
+  const seen = { nonFinite: false }
   try {
-    return JSON.parse(raw) as unknown
+    const parsed = JSON.parse(raw, (_key, value: unknown) => {
+      // `1e400` is valid JSON but parses to Infinity, which JSON.stringify
+      // reports as null; keep the raw text rather than misdescribe the call.
+      if (typeof value === 'number' && !Number.isFinite(value)) seen.nonFinite = true
+      return value
+    }) as unknown
+    return seen.nonFinite ? raw : parsed
   } catch {
     return raw
   }

+ 37 - 1
packages/bundle/headless/tests/headless.spec.ts

@@ -46,6 +46,8 @@ interface BenchOptions {
   observe?: () => Promise<ObservationStub>
   /** Leave the persistence service unmounted to exercise the fail-loud path. */
   omitPersistence?: boolean
+  /** Mount the service but make it report no stored record for any id. */
+  unbackedPersistence?: boolean
   /** Register a live Agent under `sessionId` before the runner starts. */
   prelive?: boolean
   /** Header facts for that pre-registered live Agent. */
@@ -175,7 +177,11 @@ async function bench(script: Script, options: BenchOptions = {}): Promise<{
   if (options.observe !== undefined) {
     ctx.provide('sessionQuery', { observeSession: () => options.observe!() } as never)
   }
-  if (options.omitPersistence !== true) ctx.provide('sessionPersistence', {} as never)
+  if (options.omitPersistence !== true) {
+    ctx.provide('sessionPersistence', {
+      stat: () => Promise.resolve(options.unbackedPersistence === true ? undefined : { header: {} }),
+    } as never)
+  }
   return {
     ctx,
     output: () => ({ out, err, order: [...order] }),
@@ -536,6 +542,21 @@ describe('headless runner', () => {
     await test.ctx.fiber.dispose()
   })
 
+  it('rejects a live Session the persistence service does not know', async () => {
+    const test = await bench({
+      afterPrompt(session, message) { appendTurn(session, 1, message, 'live', true) },
+    }, {
+      sessionId: 'session-exact',
+      prelive: true,
+      unbackedPersistence: true,
+    })
+    const result = await test.run()
+    expect(result.code).toBe(1)
+    expect(result.err).toContain('has no persisted record')
+    expect(result.out).toBe('')
+    await test.ctx.fiber.dispose()
+  })
+
   it('resumes the persisted Session when the query finds it', async () => {
     const test = await bench({
       afterPrompt(session, message) { appendTurn(session, 1, message, 'resumed answer', true) },
@@ -603,6 +624,21 @@ describe('headless runner', () => {
     await test.ctx.fiber.dispose()
   })
 
+  it('rejects a persisted Session whose preset record names no preset', async () => {
+    const test = await bench({ afterPrompt: () => {} }, {
+      sessionId: 'session-exact',
+      observe: () => Promise.resolve({
+        header: { cwd: process.cwd() },
+        events: [{ type: 'agent-preset/selected', data: {} }],
+        [Symbol.dispose]() {},
+      }),
+    })
+    const result = await test.run()
+    expect(result.code).toBe(1)
+    expect(result.err).toContain('malformed agent-preset/selected event')
+    await test.ctx.fiber.dispose()
+  })
+
   it('rejects a persisted Session that recorded no working directory', async () => {
     const test = await bench({ afterPrompt: () => {} }, {
       sessionId: 'session-exact',

+ 9 - 0
packages/bundle/headless/tests/json-stream.spec.ts

@@ -205,6 +205,15 @@ describe('--json projection', () => {
     })
   })
 
+  it('keeps raw arguments that JSON cannot round-trip, such as an overflowing number', () => {
+    const test = harness({}, 's1')
+    test.emitSession({
+      type: 'tool/call',
+      data: { turn: 1, step: 1, callId: 'inf', name: 'bash', arguments: '{"n":1e400}' },
+    } as unknown as SessionEvent)
+    expect(test.parsed()[1]).toEqual({ type: 'tool_call', callId: 'inf', tool: 'bash', input: '{"n":1e400}' })
+  })
+
   it('keeps a literal __proto__ key and bounds over-long object keys', () => {
     const proto = harness({ maxStringBytes: 32 }, 's1')
     proto.emitSession({