Переглянути джерело

fix(headless): close the v11 review gaps in the json run surface

Address the ds-review-bot v11 review on PR #3849:

- Accumulate every attempt's usage into `step_end.usage`, including a retried
  attempt whose only sample sits in its discarded `assistant/attempt` stream.
- Re-check the adopted Session after the idle wait, so a preset selected in
  that window is still rejected.
- Reject a whitespace-only `sessionId` supplied through a Cordis overlay, which
  bypasses the CLI trim check.
- Narrow the documented error contract: a profile whose own plugins fail to
  load exits before the runner mounts and keeps only stderr diagnostics.
lsdsjy 2 тижнів тому
батько
коміт
980b4b77af

+ 2 - 2
.agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.md
-2026-09-09-headless-machine-readable-run-surface.md: 4afb0e68b0eba335b480f8d954a1afe559741d1e
-2026-09-09-headless-machine-readable-run-surface.zh.md: e55476a733fbda6904b203571b87c48cc3cd2f26
+2026-09-09-headless-machine-readable-run-surface.md: 212fd956b833cedf90804340cff22b85ce9f38aa
+2026-09-09-headless-machine-readable-run-surface.zh.md: 32986ab8f950eb14c687aece7902d0e11b31e572

+ 3 - 3
.agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.md

@@ -46,7 +46,7 @@ Task resolution order: joined positionals, then `-`, then piped stdin. A whitesp
 | `thinking` | `text` | one committed reasoning block |
 | `tool_call` | `callId`, `tool`, `input` | once per call |
 | `tool_result` | `callId`, `status`, `result` | once per appended result |
-| `error` | `message` | process-level failure outside a turn |
+| `error` | `message` | a failure the runner raises outside a turn |
 | `final` | `text` | last line |
 
 Projection rules:
@@ -56,7 +56,7 @@ Projection rules:
 - A `tool/result` is projected only when its `surfaceOp` is `append`. A compaction replacement of an older result is history, and projecting it would emit a call id with no matching `tool_call`.
 - Every projected string and object key is bounded at 8 KiB, an event with a cut value carries `truncated: true`, and one serialized event line, newline included, is bounded at 32 KiB — an over-long event keeps its scalar fields, drops structured ones, and at the extreme reduces to `type` and `truncated`, while a payload nested 64 levels or deeper is cut at that depth so no legal input can overflow the bounding recursion. This includes the process-level `error` event; a literal `__proto__` argument key is copied as data rather than through the inherited setter, an empty tool-argument string projects as `{}` to match the executor, and arguments that JSON cannot round-trip (an overflowing number such as `1e400`) keep their raw text rather than the `null` that `JSON.stringify` would report. The terminal `final` event is deliberately unbounded: it carries the same lossless answer the default mode prints.
 - Text and reasoning arrive when the step commits, not per token; default-mode stderr reasoning remains the only live text channel. A turn that fails in-turn still ends the stream with `final` and no `error` event, so a supervisor classifies that run from the exit code and the `turn_end` reason even when the stream is well formed.
-- `usage` appears on `step_end`, matching the token accounting a provider reports per step.
+- `usage` appears on `step_end`, summing every attempt in the step — including a retried attempt whose only usage sample sits in its discarded `assistant/attempt` stream — so it matches the token accounting the provider billed for that step.
 - Raw session events stay out of scope. A debug escape hatch can be added later without changing this vocabulary.
 
 ### Session identity
@@ -65,7 +65,7 @@ The runtime owns identity. A run without `--session-id` mints `session-<uuid>` a
 
 `--session-id <id>` is adopt-or-create: observe the persisted session, resume it when it exists, create it otherwise. Create-only would fail the second run, because the JSONL store rejects an existing log id ([session persistence](../../implemented/architecture/2026-06-14-session-persistence.md)). The id is opaque, so the runner validates non-emptiness on the trimmed value and passes the caller's exact string through, whitespace included.
 
-Adoption compares the persisted session's recorded cwd with the process cwd, since sessions are organized per project directory ([project session directories](../../implemented/architecture/2026-07-24-project-session-directories.md)). A mismatch exits 1 with a `dsh:` diagnostic instead of silently continuing a conversation rooted elsewhere, and a session that recorded no cwd is rejected for the same reason. A session running under an agent preset is rejected because this bundle composes no preset roster: resuming it here would run it under the headless tools and prompts instead of the composition its log records. The check reads the preset the log currently records — the creation header advanced by any `agent-preset/selected` event — because a blank session may switch preset after creation while the header stays a creation fact, and a malformed selection record fails closed rather than reading as no preset. A session linked to a parent or subagent — including a user fork — is rejected. All checks run when a live Agent already holds the requested id, so a live identity cannot bypass them. Two live processes cannot write one id; the store's write lease already rejects the second writer. The runner reads the observation through the composed `sessionQuery` service and fails loudly when `--session-id` is requested without it — including when a live Agent already holds the id, because a later process has to find it — or when the requested identity would lack the `sessionPersistence` service that makes it durable; a live identity must also carry a stored record, because one registered only in memory would flush nothing.
+Adoption compares the persisted session's recorded cwd with the process cwd, since sessions are organized per project directory ([project session directories](../../implemented/architecture/2026-07-24-project-session-directories.md)). A mismatch exits 1 with a `dsh:` diagnostic instead of silently continuing a conversation rooted elsewhere, and a session that recorded no cwd is rejected for the same reason. A session running under an agent preset is rejected because this bundle composes no preset roster: resuming it here would run it under the headless tools and prompts instead of the composition its log records. The check reads the preset the log currently records — the creation header advanced by any `agent-preset/selected` event — because a blank session may switch preset after creation while the header stays a creation fact, and a malformed selection record fails closed rather than reading as no preset. A session linked to a parent or subagent — including a user fork — is rejected. All checks run when a live Agent already holds the requested id, and again after the runner's idle wait, so a live identity cannot bypass them and a preset selected in that window is still rejected. A whitespace-only `sessionId` is rejected on both the CLI and the direct-config path. Two live processes cannot write one id; the store's write lease already rejects the second writer. The runner reads the observation through the composed `sessionQuery` service and fails loudly when `--session-id` is requested without it — including when a live Agent already holds the id, because a later process has to find it — or when the requested identity would lack the `sessionPersistence` service that makes it durable; a live identity must also carry a stored record, because one registered only in memory would flush nothing.
 
 ## Consequences
 

+ 3 - 3
.agents/notes/implemented/feature/2026-09-09-headless-machine-readable-run-surface.zh.md

@@ -46,7 +46,7 @@ dsh --profile headless [--json] [--session-id <id>] [<task>... | -]
 | `thinking` | `text` | 一个已提交的推理块 |
 | `tool_call` | `callId`、`tool`、`input` | 每次调用一条 |
 | `tool_result` | `callId`、`status`、`result` | 每次追加的结果一条 |
-| `error` | `message` | 轮次之外的进程级失败 |
+| `error` | `message` | runner 在轮次之外抛出的失败 |
 | `final` | `text` | 最后一行 |
 
 投影规则:
@@ -56,7 +56,7 @@ dsh --profile headless [--json] [--session-id <id>] [<task>... | -]
 - `tool/result` 仅在其 `surfaceOp` 为 `append` 时投影。压缩对旧结果的替换属于历史,投影它会产生没有对应 `tool_call` 的 call id。
 - 每个被投影的字符串与对象键都限制在 8 KiB;被截断的事件带 `truncated: true`,单条序列化事件行(含换行)限制在 32 KiB——超长事件保留标量字段、丢弃结构化字段,极端情况下只剩 `type` 与 `truncated`,嵌套达到 64 层及以上的负载会在该深度被截断,因此任何合法输入都不会让限界递归溢出。进程级 `error` 事件同样受限;字面量 `__proto__` 参数键会作为数据复制,而不经过继承的 setter;空工具参数字符串会投影为 `{}`,与执行器保持一致;JSON 无法往返的参数(例如溢出为 `Infinity` 的 `1e400`)保留原始文本,而不是 `JSON.stringify` 会报告的 `null`。终止 `final` 事件刻意不做限长:它承载与默认模式相同的无损答案。
 - 文本与推理在步骤提交时到达,而不是逐 token 到达;默认模式的 stderr 推理仍是唯一的实时文本通道。轮次内失败的运行仍以 `final` 结束且没有 `error` 事件,因此即使事件流格式良好,监督进程也要用退出码与 `turn_end` 原因来分类该次运行。
-- `usage` 出现在 `step_end` 上,对应 provider 每步上报的 token 计量。
+- `usage` 出现在 `step_end` 上,累计该步的每一次 attempt——包括仅在被丢弃的 `assistant/attempt` 流中留下用量样本的重试——与 provider 为该步计费的 token 计量一致。
 - 原始会话事件不在范围内。调试用的逃生口可以以后再加,不必改动这套词汇表。
 
 ### 会话身份
@@ -65,7 +65,7 @@ dsh --profile headless [--json] [--session-id <id>] [<task>... | -]
 
 `--session-id <id>` 是采用或创建:先观察持久化会话,存在就 resume,不存在就 create。只创建会让第二次运行失败,因为 JSONL 存储拒绝已存在的日志 id(见 [session persistence](../../implemented/architecture/2026-06-14-session-persistence.zh.md))。标识是不透明的,因此 runner 只在 trim 后的值上校验非空,并把调用方的原始字符串(含空白字符)原样传下去。
 
-采用时会比较持久化会话记录的 cwd 与进程 cwd,因为会话按项目目录组织(见 [project session directories](../../implemented/architecture/2026-07-24-project-session-directories.zh.md))。不一致时以 `dsh:` 诊断退出 1,而不是静默续接一个根目录在别处的会话;未记录 cwd 的会话出于同样理由被拒绝。运行在 agent preset 下的会话被拒绝,因为本 bundle 不组合任何 preset roster:在这里 resume 它,会用 headless 的工具与提示词运行它,而不是它日志当前记录的组合。该检查读取日志当前记录的 preset——创建 header 再叠加任何 `agent-preset/selected` 事件——因为空白会话可能在创建后切换 preset,而 header 始终只是创建事实;畸形的选择记录会失败关闭,而不会读成「无 preset」。带父会话或子 agent 关联的会话——包括用户 fork 出的会话——被拒绝。当某个存活 Agent 已经持有请求的 id 时,上述检查全部执行,因此存活身份无法绕过它们。两个存活进程不能写同一个 id;存储的写租约已经会拒绝第二个写入者。runner 通过已组合的 `sessionQuery` 服务读取观察结果,并在请求 `--session-id` 却没有该服务时显式失败——即便本进程已有存活 Agent 持有该 id,因为后续进程仍需找到它;若所请求的身份缺少让它持久化的 `sessionPersistence` 服务,同样显式失败;存活身份还必须已有持久化记录,因为仅注册在内存中的身份不会写入任何内容。
+采用时会比较持久化会话记录的 cwd 与进程 cwd,因为会话按项目目录组织(见 [project session directories](../../implemented/architecture/2026-07-24-project-session-directories.zh.md))。不一致时以 `dsh:` 诊断退出 1,而不是静默续接一个根目录在别处的会话;未记录 cwd 的会话出于同样理由被拒绝。运行在 agent preset 下的会话被拒绝,因为本 bundle 不组合任何 preset roster:在这里 resume 它,会用 headless 的工具与提示词运行它,而不是它日志当前记录的组合。该检查读取日志当前记录的 preset——创建 header 再叠加任何 `agent-preset/selected` 事件——因为空白会话可能在创建后切换 preset,而 header 始终只是创建事实;畸形的选择记录会失败关闭,而不会读成「无 preset」。带父会话或子 agent 关联的会话——包括用户 fork 出的会话——被拒绝。当某个存活 Agent 已经持有请求的 id 时,上述检查全部执行,并在 runner 等待 idle 后再次执行,因此存活身份无法绕过它们,在该窗口内选中的 preset 也仍会被拒绝。纯空白的 `sessionId` 在 CLI 与直接配置两条路径上都会被拒绝。两个存活进程不能写同一个 id;存储的写租约已经会拒绝第二个写入者。runner 通过已组合的 `sessionQuery` 服务读取观察结果,并在请求 `--session-id` 却没有该服务时显式失败——即便本进程已有存活 Agent 持有该 id,因为后续进程仍需找到它;若所请求的身份缺少让它持久化的 `sessionPersistence` 服务,同样显式失败;存活身份还必须已有持久化记录,因为仅注册在内存中的身份不会写入任何内容。
 
 ## 后果
 

+ 2 - 2
packages/bundle/headless/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write packages/bundle/headless/README.md
-README.md: 2eabc77bca147a2e754284cbecb938a1430ab546
-README.zh.md: 999c2327245ab75ca8557718338e4b8ca52d1990
+README.md: 8bdad30de4de3e1e2c4921c2f879454c3440822f
+README.zh.md: 1675c026f4b82752d2e7a5ae1a51b320b8eb0f27

+ 1 - 1
packages/bundle/headless/README.md

@@ -55,7 +55,7 @@ Every invocation defaults to a fresh `session-<uuid>` identity. Pass `--session-
 
 ### Machine-readable output
 
-`--json` replaces the final-text stdout line with a newline-delimited JSON event stream, while stderr keeps only the `dsh:` diagnostics. The stream opens with `session` (carrying the identity the run used) and closes with `final`, and carries `status`, `text`, `thinking`, `tool_call`, and `tool_result` events in between. `text` and `thinking` are projected from committed assistant messages, so a retried or discarded attempt never reaches the stream; they arrive when the step commits, not per token, and default-mode stderr reasoning remains the only live text channel. The terminal `final` event carries the same lossless answer as the default mode and is not capped; every other string and object key is capped at 8 KiB and flagged with `truncated`, and one event line, its newline included, is capped at 32 KiB — an over-long event keeps its scalar fields, drops structured ones, and at the extreme reduces to `type` and `truncated`, while a payload nested 64 levels or deeper is cut at that depth. An empty tool-argument string projects as `{}`, matching what the executor runs, while arguments that JSON cannot round-trip — an overflowing number such as `1e400` — keep their raw text rather than the `null` that `JSON.stringify` would report. A process-level failure outside a turn writes an `error` event and ends the stream without `final`, in addition to the `dsh:` stderr line. A turn that fails in-turn still ends with a `final` event (often empty) and no `error` event, so a well-formed stream can still describe a failed run: treat exit code 1 and the `turn_end` reason as the failure signal.
+`--json` replaces the final-text stdout line with a newline-delimited JSON event stream, while stderr keeps only the `dsh:` diagnostics. The stream opens with `session` (carrying the identity the run used) and closes with `final`, and carries `status`, `text`, `thinking`, `tool_call`, and `tool_result` events in between. `text` and `thinking` are projected from committed assistant messages, so a retried or discarded attempt never reaches the stream; they arrive when the step commits, not per token, and default-mode stderr reasoning remains the only live text channel. The terminal `final` event carries the same lossless answer as the default mode and is not capped; every other string and object key is capped at 8 KiB and flagged with `truncated`, and one event line, its newline included, is capped at 32 KiB — an over-long event keeps its scalar fields, drops structured ones, and at the extreme reduces to `type` and `truncated`, while a payload nested 64 levels or deeper is cut at that depth. An empty tool-argument string projects as `{}`, matching what the executor runs, while arguments that JSON cannot round-trip — an overflowing number such as `1e400` — keep their raw text rather than the `null` that `JSON.stringify` would report. A failure the runner raises outside a turn writes an `error` event and ends the stream without `final`, in addition to the `dsh:` stderr line; a profile whose own plugins fail to load exits before the runner mounts, so that case keeps only the loader's stderr diagnostics. A turn that fails in-turn still ends with a `final` event (often empty) and no `error` event, so a well-formed stream can still describe a failed run: treat exit code 1 and the `turn_end` reason as the failure signal.
 
 ### When to use it
 

+ 1 - 1
packages/bundle/headless/README.zh.md

@@ -55,7 +55,7 @@ agent(智能体)会完成该任务,把提供方的每个非空推理增量
 
 ### 机器可读输出
 
-`--json` 用按行 JSON 事件流取代 stdout 的最终文本行,stderr 仅保留 `dsh:` 诊断信息。事件流以 `session`(携带本次运行使用的标识)开头、以 `final` 结尾,其间为 `status`、`text`、`thinking`、`tool_call` 与 `tool_result` 事件。`text` 与 `thinking` 只从已提交的 assistant 消息投影,因此被重试或丢弃的尝试不会进入事件流;它们在步骤提交时到达,而不是逐 token 到达,默认模式的 stderr 推理仍是唯一的实时文本通道。终止 `final` 事件携带与默认模式相同的无损答案,不做限长;其他每个字符串与对象键上限为 8 KiB,超出时标记 `truncated`,单条事件行(含换行)上限为 32 KiB——超长事件保留标量字段、丢弃结构化字段,极端情况下只剩 `type` 与 `truncated`,嵌套达到 64 层及以上的负载会在该深度被截断。空工具参数字符串会投影为 `{}`,与执行器实际运行的值一致;而 JSON 无法往返的参数——例如溢出为 `Infinity` 的数字 `1e400`——会保留原始文本,而不是 `JSON.stringify` 会报告的 `null`。轮次之外的进程级失败会写出 `error` 事件并在没有 `final` 的情况下结束事件流,同时向 stderr 写入 `dsh:` 行。轮次内失败的运行仍会以 `final` 事件(通常为空)结束且没有 `error` 事件,因此格式良好的事件流也可能描述一次失败的运行:请把退出码 1 与 `turn_end` 原因作为失败信号。
+`--json` 用按行 JSON 事件流取代 stdout 的最终文本行,stderr 仅保留 `dsh:` 诊断信息。事件流以 `session`(携带本次运行使用的标识)开头、以 `final` 结尾,其间为 `status`、`text`、`thinking`、`tool_call` 与 `tool_result` 事件。`text` 与 `thinking` 只从已提交的 assistant 消息投影,因此被重试或丢弃的尝试不会进入事件流;它们在步骤提交时到达,而不是逐 token 到达,默认模式的 stderr 推理仍是唯一的实时文本通道。终止 `final` 事件携带与默认模式相同的无损答案,不做限长;其他每个字符串与对象键上限为 8 KiB,超出时标记 `truncated`,单条事件行(含换行)上限为 32 KiB——超长事件保留标量字段、丢弃结构化字段,极端情况下只剩 `type` 与 `truncated`,嵌套达到 64 层及以上的负载会在该深度被截断。空工具参数字符串会投影为 `{}`,与执行器实际运行的值一致;而 JSON 无法往返的参数——例如溢出为 `Infinity` 的数字 `1e400`——会保留原始文本,而不是 `JSON.stringify` 会报告的 `null`。runner 在轮次之外抛出的失败会写出 `error` 事件并在没有 `final` 的情况下结束事件流,同时向 stderr 写入 `dsh:` 行;若 profile 自身的插件加载失败,进程会在 runner 挂载前退出,该情形只保留 loader 的 stderr 诊断。轮次内失败的运行仍会以 `final` 事件(通常为空)结束且没有 `error` 事件,因此格式良好的事件流也可能描述一次失败的运行:请把退出码 1 与 `turn_end` 原因作为失败信号。
 
 ### 何时使用
 

+ 11 - 0
packages/bundle/headless/src/index.ts

@@ -329,6 +329,12 @@ async function run(ctx: Context, config: Config, io: HeadlessIo): Promise<void>
   // Early process shutdown can dispose the tree while settlement is pending.
   if (agents === undefined || defaultModel === undefined || sessions === undefined) return
 
+  // A Cordis overlay sets the row directly and bypasses the CLI trim check, so
+  // the same public setting must fail here rather than become a blank identity.
+  if (config.sessionId !== undefined && config.sessionId.trim() === '') {
+    throw new Error('headless-runner: sessionId must not be blank')
+  }
+
   const task = config.task === undefined || config.task === '-'
     ? await internals.readStdin()
     : config.task
@@ -356,6 +362,11 @@ async function run(ctx: Context, config: Config, io: HeadlessIo): Promise<void>
     })).agent
     : await resolveAgent(ctx, agents, sessionId, agentOptions, setup)
   await agent.whenIdle()
+  if (config.sessionId !== undefined) {
+    // A live Session can select an agent preset while the runner awaits idle;
+    // re-read its log so the rejection cannot be outrun by that timing.
+    assertAdoptable(agent.session.header, liveEvents(agent.session), sessionId)
+  }
   const firstSeq = agent.session.seq
   const projection = config.json === true ? projectJsonRun(ctx, agent, io.stdout) : undefined
   const stopReasoning = projection === undefined ? streamReasoning(ctx, agent, io.stderr) : undefined

+ 44 - 2
packages/bundle/headless/src/json-stream.ts

@@ -9,6 +9,7 @@
 
 import type { Context } from '@deepseek-ai/cordis'
 import type { Agent } from '@deepseek-ai/dsh-agent'
+import { lastAssistantStreamChunk } from '@deepseek-ai/dsh-llm/assistant-stream'
 import type { SessionEvent } from '@deepseek-ai/dsh-session'
 
 /** Default per-string and per-key cap applied to every bounded projected payload. */
@@ -176,6 +177,42 @@ function resultText(blocks: readonly { type: string; text?: string }[]): string
     .join('')
 }
 
+/** Provider token accounting for one step, as carried by `assistant/message`. */
+type StepUsage = NonNullable<SessionEvent<'assistant/message'>['data']['usage']>
+
+/** The usage a provider reported in one attempt's stream, if any. */
+function streamUsage(stream: SessionEvent<'assistant/message'>['data']['stream']): StepUsage | undefined {
+  return lastAssistantStreamChunk(stream, 'usage')?.usage
+}
+
+/**
+ * Accumulate one step's usage across its attempts. A retried attempt keeps its
+ * usage only in its `assistant/attempt` stream, so the committed message alone
+ * would under-report billed tokens; an optional bucket is summed only when
+ * every contribution reports it.
+ * @param total - usage accumulated so far.
+ * @param next - usage reported by the next attempt.
+ * @returns the accumulated usage, or `undefined` when neither side reports any.
+ */
+function addUsage(total: StepUsage | undefined, next: StepUsage | undefined): StepUsage | undefined {
+  if (next === undefined) return total
+  if (total === undefined) return next
+  const sum = (a: number | undefined, b: number | undefined): number | undefined =>
+    a === undefined || b === undefined ? undefined : a + b
+  const totalTokens = sum(total.totalTokens, next.totalTokens)
+  const cacheReadTokens = sum(total.cacheReadTokens, next.cacheReadTokens)
+  const cacheWriteTokens = sum(total.cacheWriteTokens, next.cacheWriteTokens)
+  const reasoningTokens = sum(total.reasoningTokens, next.reasoningTokens)
+  return {
+    inputTokens: total.inputTokens + next.inputTokens,
+    outputTokens: total.outputTokens + next.outputTokens,
+    ...totalTokens === undefined ? {} : { totalTokens },
+    ...cacheReadTokens === undefined ? {} : { cacheReadTokens },
+    ...cacheWriteTokens === undefined ? {} : { cacheWriteTokens },
+    ...reasoningTokens === undefined ? {} : { reasoningTokens },
+  }
+}
+
 /**
  * Project one Agent's run as newline-delimited JSON on `sink`.
  *
@@ -198,7 +235,7 @@ export function projectJsonRun(
 ): JsonProjection {
   const maxStringBytes = options.maxStringBytes ?? MAX_STRING_BYTES
   let disposed = false
-  let stepUsage: SessionEvent<'assistant/message'>['data']['usage']
+  let stepUsage: StepUsage | undefined
 
   const write = (event: Record<string, unknown>): void => {
     sink.write(`${boundJsonLine(event, maxStringBytes)}\n`)
@@ -213,8 +250,13 @@ export function projectJsonRun(
       case 'step/start':
         write({ type: 'status', phase: 'step_start', turn: event.data.turn, step: event.data.step })
         return
+      case 'assistant/attempt':
+        // A retried attempt has no committed message; its billed tokens live
+        // only in the stream, so accumulate them into the step's total.
+        stepUsage = addUsage(stepUsage, streamUsage(event.data.stream))
+        return
       case 'assistant/message':
-        stepUsage = event.data.usage
+        stepUsage = addUsage(stepUsage, event.data.usage ?? streamUsage(event.data.stream))
         for (const block of event.data.message.content) {
           if (block.type === 'reasoning') write({ type: 'thinking', text: block.text })
           else if (block.type === 'text') write({ type: 'text', text: block.text })

+ 33 - 3
packages/bundle/headless/tests/headless.spec.ts

@@ -54,6 +54,8 @@ interface BenchOptions {
   prelive?: boolean
   /** Header facts for that pre-registered live Agent. */
   preliveMeta?: { cwd?: string; origin?: 'subagent'; agentPreset?: string }
+  /** Run when the live path reads persistence, e.g. to mutate the live log. */
+  onStat?: (agent: Agent) => void
 }
 
 const frameStates = new WeakMap<Agent, { attemptId: ReturnType<typeof LlmAttemptId>; revision: number; index: number }>()
@@ -180,9 +182,13 @@ async function bench(script: Script, options: BenchOptions = {}): Promise<{
     const observe = options.observe ?? (() => Promise.reject(new SessionQueryError('missing', 'SESSION_QUERY_SESSION_NOT_FOUND')))
     ctx.provide('sessionQuery', { observeSession: () => observe() } as never)
   }
+  let preliveAgent: Agent | undefined
   if (options.omitPersistence !== true) {
     ctx.provide('sessionPersistence', {
-      stat: () => Promise.resolve(options.unbackedPersistence === true ? undefined : { header: {} }),
+      stat: () => {
+        if (options.onStat !== undefined && preliveAgent !== undefined) options.onStat(preliveAgent)
+        return Promise.resolve(options.unbackedPersistence === true ? undefined : { header: {} })
+      },
     } as never)
   }
   return {
@@ -197,10 +203,10 @@ async function bench(script: Script, options: BenchOptions = {}): Promise<{
         ctx.provide('appExit', (code: number) => { order.push('exit'); resolve(code) })
       })
       if (options.prelive === true || options.preliveMeta !== undefined) {
-        await ctx.agents.create({
+        preliveAgent = (await ctx.agents.create({
           sessionId: brandString<SessionId>(options.sessionId ?? 'session-exact'),
           meta: { cwd: process.cwd(), ...options.preliveMeta },
-        })
+        })).agent
       }
       apply(ctx, {
         ...options.useStdin === true ? {} : { task: options.task ?? 'do the thing' },
@@ -695,6 +701,15 @@ describe('headless runner', () => {
     await test.ctx.fiber.dispose()
   })
 
+  it('rejects a whitespace-only session identity from configuration', async () => {
+    const test = await bench({ afterPrompt: () => {} }, { sessionId: '   ' })
+    const result = await test.run()
+    expect(result.code).toBe(1)
+    expect(result.err).toContain('sessionId must not be blank')
+    expect(result.out).toBe('')
+    await test.ctx.fiber.dispose()
+  })
+
   it('reuses a live Agent already registered under the requested identity', async () => {
     const test = await bench({
       afterPrompt(session, message) { appendTurn(session, 1, message, 'live answer', true) },
@@ -759,6 +774,21 @@ describe('headless runner', () => {
     await test.ctx.fiber.dispose()
   })
 
+  it('rejects a preset selected while the runner awaits idle', async () => {
+    const test = await bench({
+      afterPrompt(session, message) { appendTurn(session, 1, message, 'live', true) },
+    }, {
+      sessionId: 'session-exact',
+      prelive: true,
+      onStat: (agent) => { selectPreset(agent.session, 'minimal') },
+    })
+    const result = await test.run()
+    expect(result.code).toBe(1)
+    expect(result.err).toContain('runs under agent preset "minimal"')
+    expect(result.out).toBe('')
+    await test.ctx.fiber.dispose()
+  })
+
   it('rejects a preset appended after the observation snapshot was taken', async () => {
     const test = await bench({ afterPrompt: () => {} }, {
       sessionId: 'session-exact',

+ 92 - 0
packages/bundle/headless/tests/json-stream.spec.ts

@@ -34,6 +34,19 @@ function assistantMessage(content: unknown[], usage?: unknown): SessionEvent {
   } as unknown as SessionEvent
 }
 
+/** One discarded attempt whose stream reports the given usage sample. */
+function attemptWithUsage(usage: unknown): SessionEvent {
+  return {
+    type: 'assistant/attempt',
+    data: { turn: 1, step: 1, stream: [{ type: 'chunk', chunk: { type: 'usage', usage } }] },
+  } as unknown as SessionEvent
+}
+
+/** One step boundary event closing the accumulated usage window. */
+function stepEnd(): SessionEvent {
+  return { type: 'step/end', data: { turn: 1, step: 1 } } as unknown as SessionEvent
+}
+
 /** One tool result event with the given surface placement. */
 function toolResult(
   callId: string,
@@ -300,6 +313,85 @@ describe('--json projection', () => {
       .toMatchObject({ type: 'tool_call', truncated: true })
   })
 
+  it('sums retried attempt usage into the step total and drops an unshared bucket', () => {
+    const test = harness()
+    test.emitSession(attemptWithUsage({ inputTokens: 10, outputTokens: 2, totalTokens: 12, cacheReadTokens: 4 }))
+    test.emitSession(assistantMessage(
+      [{ type: 'text', text: 'ok' }],
+      { inputTokens: 3, outputTokens: 1, totalTokens: 4 },
+    ))
+    test.emitSession(stepEnd())
+    expect(test.parsed().at(-1)).toEqual({
+      type: 'status', phase: 'step_end', turn: 1, step: 1,
+      usage: { inputTokens: 13, outputTokens: 3, totalTokens: 16 },
+    })
+  })
+
+  it('sums every optional bucket both attempts report', () => {
+    const test = harness()
+    const sample = {
+      inputTokens: 2, outputTokens: 1, totalTokens: 3,
+      cacheReadTokens: 1, cacheWriteTokens: 1, reasoningTokens: 1,
+    }
+    test.emitSession(attemptWithUsage(sample))
+    test.emitSession(attemptWithUsage(sample))
+    test.emitSession(stepEnd())
+    expect(test.parsed().at(-1)).toMatchObject({
+      usage: {
+        inputTokens: 4, outputTokens: 2, totalTokens: 6,
+        cacheReadTokens: 2, cacheWriteTokens: 2, reasoningTokens: 2,
+      },
+    })
+  })
+
+  it('drops a bucket only the later attempt reports', () => {
+    const test = harness()
+    test.emitSession(attemptWithUsage({ inputTokens: 1, outputTokens: 1 }))
+    test.emitSession(attemptWithUsage({ inputTokens: 1, outputTokens: 1, cacheWriteTokens: 2 }))
+    test.emitSession(stepEnd())
+    const usage = (test.parsed().at(-1) as { usage: Record<string, unknown> }).usage
+    expect(usage).toEqual({ inputTokens: 2, outputTokens: 2 })
+    expect(usage).not.toHaveProperty('cacheWriteTokens')
+  })
+
+  it('reads a message usage sample from its stream when the field is absent', () => {
+    const test = harness()
+    test.emitSession({
+      type: 'assistant/message',
+      data: {
+        turn: 1,
+        step: 1,
+        stream: [{ type: 'chunk', chunk: { type: 'usage', usage: { inputTokens: 5, outputTokens: 1 } } }],
+        message: {
+          role: 'assistant',
+          content: [{ type: 'text', text: 'hi' }],
+          source: { kind: 'model', provider: 'p', model: 'm' },
+        },
+      },
+    } as unknown as SessionEvent)
+    test.emitSession(stepEnd())
+    expect(test.parsed().at(-1)).toMatchObject({ usage: { inputTokens: 5, outputTokens: 1 } })
+  })
+
+  it('keeps an earlier attempt sample when the committed message reports none', () => {
+    const test = harness()
+    test.emitSession(attemptWithUsage({ inputTokens: 7, outputTokens: 3 }))
+    test.emitSession(assistantMessage([{ type: 'text', text: 'ok' }]))
+    test.emitSession(stepEnd())
+    expect(test.parsed().at(-1)).toMatchObject({ usage: { inputTokens: 7, outputTokens: 3 } })
+  })
+
+  it('ignores an attempt that reports no usage sample', () => {
+    const test = harness()
+    test.emitSession({
+      type: 'assistant/attempt',
+      data: { turn: 1, step: 1, stream: [] },
+    } as unknown as SessionEvent)
+    test.emitSession(assistantMessage([{ type: 'text', text: 'ok' }], { inputTokens: 1, outputTokens: 1 }))
+    test.emitSession(stepEnd())
+    expect(test.parsed().at(-1)).toMatchObject({ usage: { inputTokens: 1, outputTokens: 1 } })
+  })
+
   it('writes the terminal final event without bounding its answer', () => {
     const test = harness({ maxStringBytes: 4 })
     test.projection.finish('abcdefgh')