Ver Fonte

Merge branch 'feat/plugin-mgmt-1-boot' into feat/plugin-mgmt-2-manager

Yichen Jiang há 2 semanas atrás
pai
commit
c22d66ffd8
100 ficheiros alterados com 2611 adições e 193 exclusões
  1. 2 2
      .agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.i18n.yaml
  2. 7 2
      .agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.md
  3. 7 2
      .agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.zh.md
  4. 2 2
      .agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.i18n.yaml
  5. 1 1
      .agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.md
  6. 1 1
      .agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.zh.md
  7. 2 2
      .agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.i18n.yaml
  8. 1 1
      .agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.md
  9. 1 1
      .agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.zh.md
  10. 2 2
      .agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.i18n.yaml
  11. 2 2
      .agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.md
  12. 2 2
      .agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.zh.md
  13. 2 2
      .agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.i18n.yaml
  14. 1 1
      .agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.md
  15. 1 1
      .agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.zh.md
  16. 2 2
      .agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.i18n.yaml
  17. 18 14
      .agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.md
  18. 18 14
      .agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.zh.md
  19. 2 2
      .agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.i18n.yaml
  20. 149 25
      .agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.md
  21. 149 25
      .agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.zh.md
  22. 2 2
      .agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.i18n.yaml
  23. 9 5
      .agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.md
  24. 9 5
      .agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.zh.md
  25. 6 0
      .agents/notes/implemented/architecture/2026-09-05-canonical-feedback-log.i18n.yaml
  26. 33 0
      .agents/notes/implemented/architecture/2026-09-05-canonical-feedback-log.md
  27. 33 0
      .agents/notes/implemented/architecture/2026-09-05-canonical-feedback-log.zh.md
  28. 6 0
      .agents/notes/implemented/architecture/2026-09-05-nonofficial-feedback-otel.i18n.yaml
  29. 33 0
      .agents/notes/implemented/architecture/2026-09-05-nonofficial-feedback-otel.md
  30. 33 0
      .agents/notes/implemented/architecture/2026-09-05-nonofficial-feedback-otel.zh.md
  31. 6 0
      .agents/notes/implemented/architecture/2026-09-05-read-only-session-migration-preparation.i18n.yaml
  32. 170 0
      .agents/notes/implemented/architecture/2026-09-05-read-only-session-migration-preparation.md
  33. 170 0
      .agents/notes/implemented/architecture/2026-09-05-read-only-session-migration-preparation.zh.md
  34. 6 0
      .agents/notes/implemented/architecture/2026-09-06-embedded-stream-record-readers.i18n.yaml
  35. 50 0
      .agents/notes/implemented/architecture/2026-09-06-embedded-stream-record-readers.md
  36. 50 0
      .agents/notes/implemented/architecture/2026-09-06-embedded-stream-record-readers.zh.md
  37. 6 0
      .agents/notes/implemented/bug-fix/2026-08-31-win32-picker-path-string-read.i18n.yaml
  38. 25 0
      .agents/notes/implemented/bug-fix/2026-08-31-win32-picker-path-string-read.md
  39. 25 0
      .agents/notes/implemented/bug-fix/2026-08-31-win32-picker-path-string-read.zh.md
  40. 6 0
      .agents/notes/implemented/bug-fix/2026-09-05-nested-terminal-cards.i18n.yaml
  41. 35 0
      .agents/notes/implemented/bug-fix/2026-09-05-nested-terminal-cards.md
  42. 35 0
      .agents/notes/implemented/bug-fix/2026-09-05-nested-terminal-cards.zh.md
  43. 6 0
      .agents/notes/implemented/bug-fix/2026-09-05-pi-ai-upgrade-compatibility.i18n.yaml
  44. 25 0
      .agents/notes/implemented/bug-fix/2026-09-05-pi-ai-upgrade-compatibility.md
  45. 25 0
      .agents/notes/implemented/bug-fix/2026-09-05-pi-ai-upgrade-compatibility.zh.md
  46. 6 0
      .agents/notes/implemented/bug-fix/2026-09-05-session-reference-model-budget.i18n.yaml
  47. 25 0
      .agents/notes/implemented/bug-fix/2026-09-05-session-reference-model-budget.md
  48. 25 0
      .agents/notes/implemented/bug-fix/2026-09-05-session-reference-model-budget.zh.md
  49. 6 0
      .agents/notes/implemented/bug-fix/2026-09-05-session-reference-spill-reuse.i18n.yaml
  50. 41 0
      .agents/notes/implemented/bug-fix/2026-09-05-session-reference-spill-reuse.md
  51. 41 0
      .agents/notes/implemented/bug-fix/2026-09-05-session-reference-spill-reuse.zh.md
  52. 6 0
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.i18n.yaml
  53. 27 0
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md
  54. 27 0
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.zh.md
  55. 2 2
      .agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.i18n.yaml
  56. 2 2
      .agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.md
  57. 2 2
      .agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.zh.md
  58. 2 2
      .agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.i18n.yaml
  59. 3 3
      .agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.md
  60. 3 4
      .agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.zh.md
  61. 2 2
      .agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.i18n.yaml
  62. 10 8
      .agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.md
  63. 10 8
      .agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.zh.md
  64. 2 2
      .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.i18n.yaml
  65. 12 8
      .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.md
  66. 11 7
      .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.zh.md
  67. 6 0
      .agents/notes/implemented/process/2026-09-03-semantic-issue-templates-and-policy.i18n.yaml
  68. 39 0
      .agents/notes/implemented/process/2026-09-03-semantic-issue-templates-and-policy.md
  69. 39 0
      .agents/notes/implemented/process/2026-09-03-semantic-issue-templates-and-policy.zh.md
  70. 6 0
      .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.i18n.yaml
  71. 48 0
      .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.md
  72. 48 0
      .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.zh.md
  73. 6 0
      .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.i18n.yaml
  74. 46 0
      .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.md
  75. 46 0
      .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.zh.md
  76. 6 0
      .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.i18n.yaml
  77. 25 0
      .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.md
  78. 25 0
      .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.zh.md
  79. 6 0
      .agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.i18n.yaml
  80. 27 0
      .agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.md
  81. 27 0
      .agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.zh.md
  82. 6 0
      .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.i18n.yaml
  83. 31 0
      .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.md
  84. 31 0
      .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.zh.md
  85. 6 0
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml
  86. 105 0
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
  87. 105 0
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md
  88. 6 0
      .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.i18n.yaml
  89. 26 0
      .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.md
  90. 26 0
      .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.zh.md
  91. 6 0
      .agents/notes/proposed/architecture/2026-09-02-system-prompt-as-surface-node.i18n.yaml
  92. 78 0
      .agents/notes/proposed/architecture/2026-09-02-system-prompt-as-surface-node.md
  93. 78 0
      .agents/notes/proposed/architecture/2026-09-02-system-prompt-as-surface-node.zh.md
  94. 6 0
      .agents/notes/proposed/feature/2026-09-02-in-history-system-prompt-replacement.i18n.yaml
  95. 74 0
      .agents/notes/proposed/feature/2026-09-02-in-history-system-prompt-replacement.md
  96. 74 0
      .agents/notes/proposed/feature/2026-09-02-in-history-system-prompt-replacement.zh.md
  97. 92 0
      .agents/skills/dsh-speed-up-perf/SKILL.md
  98. 16 13
      .github/ISSUE_TEMPLATE/bug.md
  99. 0 1
      .github/ISSUE_TEMPLATE/config.yml
  100. 4 11
      .github/ISSUE_TEMPLATE/feature.md

+ 2 - 2
.agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.md
-2026-07-08-tool-output-spill-files.md: 4e18e9887ab636d31dfe40b1cd38e078e4b30c23
-2026-07-08-tool-output-spill-files.zh.md: b900c168d81253189c7229cc77d6400c614aa89c
+2026-07-08-tool-output-spill-files.md: a80b6cdcb0e731a687d511d38ea288a68a298173
+2026-07-08-tool-output-spill-files.zh.md: a995eafc7878b342a2164c5116da1f43ada791a5

+ 7 - 2
.agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.md

@@ -22,7 +22,7 @@ A thin spill storage seam plus a default spill policy plugin, in a new `packages
 | `@deepseek-ai/dsh-spill-local` | Local backend: private, session-scoped file storage on the host filesystem. |
 | `@deepseek-ai/dsh-spill-local` | Local backend: private, session-scoped file storage on the host filesystem. |
 | `@deepseek-ai/dsh-spill-policy` | Tool-result policy plugin: wraps final text results after dispatch and replaces oversized results with a retained preview plus a spill locator. |
 | `@deepseek-ai/dsh-spill-policy` | Tool-result policy plugin: wraps final text results after dispatch and replaces oversized results with a retained preview plus a spill locator. |
 
 
-There is no dedicated model-facing Consumer package. The Consumer is the existing `ctx.tools` execution pipeline: `dsh-spill-policy` consumes final tool results through the `tools/post-execute` waterfall, and the model follows the backend-supplied retrieval hint for the returned locator.
+The tool-result Consumer is `dsh-spill-policy`, which consumes final tool results through the `tools/post-execute` waterfall. The model follows the backend-supplied retrieval hint for the returned locator. [Session-reference spill reuse](../bug-fix/2026-09-05-session-reference-spill-reuse.md) adds a direct storage consumer with separate preview, provenance, and failure semantics; it does not change the tool-result policy.
 
 
 ### Spill seam
 ### Spill seam
 
 
@@ -33,10 +33,15 @@ interface SpillStore {
   saveText(input: SaveTextSpill): Promise<SpillRef>
   saveText(input: SaveTextSpill): Promise<SpillRef>
 }
 }
 
 
-interface SpillSource {
+type SpillSource = {
+  kind: 'tool'
   toolName: string
   toolName: string
   callId: ToolCallId
   callId: ToolCallId
   label: string
   label: string
+} | {
+  kind: 'session-reference'
+  sessionId: SessionId
+  label: string
 }
 }
 
 
 interface SaveTextSpill {
 interface SaveTextSpill {

+ 7 - 2
.agents/notes/implemented/architecture/2026-07-08-tool-output-spill-files.zh.md

@@ -22,7 +22,7 @@ Status: implemented
 | `@deepseek-ai/dsh-spill-local` | 本地后端:在宿主文件系统中提供私有、会话作用域的文件存储。 |
 | `@deepseek-ai/dsh-spill-local` | 本地后端:在宿主文件系统中提供私有、会话作用域的文件存储。 |
 | `@deepseek-ai/dsh-spill-policy` | 工具结果策略插件:包装分发后的最终文本结果,并以保留预览和 spill 定位符替换超大结果。 |
 | `@deepseek-ai/dsh-spill-policy` | 工具结果策略插件:包装分发后的最终文本结果,并以保留预览和 spill 定位符替换超大结果。 |
 
 
-系统不增加专用的面向模型消费方包。消费方是现有 `ctx.tools` 执行流水线:`dsh-spill-policy` 通过 `tools/post-execute` waterfall(瀑布式事件)使用最终工具结果,模型则按照后端随定位符返回的检索提示读取内容。
+工具结果消费方是 `dsh-spill-policy`,它通过 `tools/post-execute` waterfall(瀑布式事件)使用最终工具结果。模型按照后端随定位符返回的检索提示读取内容。[会话引用 spill 复用](../bug-fix/2026-09-05-session-reference-spill-reuse.zh.md)增加一个直接存储消费方,采用独立的预览、来源信息与失败语义;它不改变工具结果策略。
 
 
 ### spill seam
 ### spill seam
 
 
@@ -33,10 +33,15 @@ interface SpillStore {
   saveText(input: SaveTextSpill): Promise<SpillRef>
   saveText(input: SaveTextSpill): Promise<SpillRef>
 }
 }
 
 
-interface SpillSource {
+type SpillSource = {
+  kind: 'tool'
   toolName: string
   toolName: string
   callId: ToolCallId
   callId: ToolCallId
   label: string
   label: string
+} | {
+  kind: 'session-reference'
+  sessionId: SessionId
+  label: string
 }
 }
 
 
 interface SaveTextSpill {
 interface SaveTextSpill {

+ 2 - 2
.agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.md
-2026-07-14-provider-routed-llm-adapters.md: 11f99a51dc4f4dc3670d29d691f366567d75f836
-2026-07-14-provider-routed-llm-adapters.zh.md: e2e5526c4bb7a29480dde37af1c4e493ebe0e111
+2026-07-14-provider-routed-llm-adapters.md: 4cc3cc3cdceeefdfea864ccf3465e528e32159c7
+2026-07-14-provider-routed-llm-adapters.zh.md: 098a6af7c951c6dee8a1b4f948bc69725e8d189c

+ 1 - 1
.agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.md

@@ -42,7 +42,7 @@ Assistant messages carry the request's `provider` and `model`, plus an optional
 
 
 A terminal successful `finish` chunk may carry replay state as a `ReplayEnvelope`: opaque response-level metadata plus optional per-block entries aligned with the emitted block sequence. `BlockAssembler` makes one keep/drop decision for content and metadata — when max-token assembly drops a tool call, the envelope loses the entry at the same position — so the state the loop attaches to the assembled assistant message's model source always describes the stored blocks, per the [max-token replay-state alignment decision](../../archived/bug-fix/2026-08-15-max-token-replay-state-alignment.md). The loop exposes no response-rewrite hook. Error and aborted responses do not produce a normal assistant message and therefore do not enter future model history.
 A terminal successful `finish` chunk may carry replay state as a `ReplayEnvelope`: opaque response-level metadata plus optional per-block entries aligned with the emitted block sequence. `BlockAssembler` makes one keep/drop decision for content and metadata — when max-token assembly drops a tool call, the envelope loses the entry at the same position — so the state the loop attaches to the assembled assistant message's model source always describes the stored blocks, per the [max-token replay-state alignment decision](../../archived/bug-fix/2026-08-15-max-token-replay-state-alignment.md). The loop exposes no response-rewrite hook. Error and aborted responses do not produce a normal assistant message and therefore do not enter future model history.
 
 
-The pi-ai replay state fills that envelope with a versioned, minimal projection of its successful `AssistantMessage`: a response half (source API/provider/model, response id/model, stop reason) and per-block text, thinking, and tool-call signatures. It does not duplicate text or tool arguments already carried by Harness content blocks, and it omits diagnostics, timestamps, usage, and errors. On a later request, `LlmRuntime` gives replay state to the target adapter only when the historical provider and target provider are currently owned by the same adapter instance. That adapter combines the logged Harness content with replay state when it can restore the historical response, and owns any required cross-model or cross-provider conversion. Durable content stays authoritative: an adapter receiving replay state it cannot use — an unknown kind or version, malformed metadata, or a block shape that no longer matches the content — degrades that message to provider-neutral conversion with a diagnostic; a different adapter receives only provider-neutral content plus provider/model fields.
+The pi-ai replay state fills that envelope with a versioned, minimal projection of its successful `AssistantMessage`: a response half (source API/provider/model, response id/model, optional provider-native effort, stop reason) and per-block text, thinking, and tool-call signatures. The [pi-ai upgrade compatibility decision](../bug-fix/2026-09-05-pi-ai-upgrade-compatibility.md#decision) defines requested-versus-native Anthropic model identity and effort preservation. It does not duplicate text or tool arguments already carried by Harness content blocks, and it omits diagnostics, timestamps, usage, and errors. On a later request, `LlmRuntime` gives replay state to the target adapter only when the historical provider and target provider are currently owned by the same adapter instance. That adapter combines the logged Harness content with replay state when it can restore the historical response, and owns any required cross-model or cross-provider conversion. Durable content stays authoritative: an adapter receiving replay state it cannot use — an unknown kind or version, malformed metadata, or a block shape that no longer matches the content — degrades that message to provider-neutral conversion with a diagnostic; a different adapter receives only provider-neutral content plus provider/model fields.
 
 
 This state is model-visible replay input and therefore follows the existing [reconstructable-request rule](2026-07-05-reconstructable-requests.md): it is present in both the terminal `finish` chunk and the assembled `assistant/message` model source that drives derivation. Resume and fork preserve it verbatim. Compaction that shadows the assistant message also removes its replay state from the active surface; the summary is ordinary provider-neutral content.
 This state is model-visible replay input and therefore follows the existing [reconstructable-request rule](2026-07-05-reconstructable-requests.md): it is present in both the terminal `finish` chunk and the assembled `assistant/message` model source that drives derivation. Resume and fork preserve it verbatim. Compaction that shadows the assistant message also removes its replay state from the active surface; the summary is ordinary provider-neutral content.
 
 

+ 1 - 1
.agents/notes/implemented/architecture/2026-07-14-provider-routed-llm-adapters.zh.md

@@ -42,7 +42,7 @@ pi-ai 的通用流选项不支持停止序列。若 Harness `stop` 选项已定
 
 
 成功的终止 `finish` 分片可以以 `ReplayEnvelope` 形式携带回放状态:不透明的响应级元数据,加上与发射块序列对齐的可选逐块条目。`BlockAssembler` 对内容与元数据只做一次保留/丢弃决定——max-token 组装丢弃工具调用时,数据同一位置的条目一并丢弃——因此 agent loop 附加到已组装助手消息模型来源中的状态始终描述存储的块,见 [max-token 回放状态对齐决定](../../archived/bug-fix/2026-08-15-max-token-replay-state-alignment.md)。agent loop 不公开响应改写钩子。错误或中止响应不会生成正常助手消息,因此不会进入后续模型历史。
 成功的终止 `finish` 分片可以以 `ReplayEnvelope` 形式携带回放状态:不透明的响应级元数据,加上与发射块序列对齐的可选逐块条目。`BlockAssembler` 对内容与元数据只做一次保留/丢弃决定——max-token 组装丢弃工具调用时,数据同一位置的条目一并丢弃——因此 agent loop 附加到已组装助手消息模型来源中的状态始终描述存储的块,见 [max-token 回放状态对齐决定](../../archived/bug-fix/2026-08-15-max-token-replay-state-alignment.md)。agent loop 不公开响应改写钩子。错误或中止响应不会生成正常助手消息,因此不会进入后续模型历史。
 
 
-pi-ai 回放状态用其成功 `AssistantMessage` 的带版本最小投影填充该结构:一个响应半区(源 API/提供方/模型、响应 ID/模型、停止原因),以及逐块的文本签名、thinking 签名和工具调用签名。它不会重复 Harness 内容块中已有的文本或工具参数,也不包含诊断信息、时间戳、用量或错误。后续请求中,只有历史提供方和目标提供方当前归同一个适配器实例所有时,`LlmRuntime` 才会把回放状态交给目标适配器。适配器在能够恢复历史响应时,将 Harness 记录的内容与回放状态组合,并负责所需的跨模型或跨提供方转换。持久化内容保持权威:适配器收到无法使用的回放状态——未知 kind 或版本、格式错误的元数据、或与内容不再匹配的块结构——会把该消息降级为提供方无关转换并带出诊断;其他适配器只能收到提供方无关的内容以及提供方/模型字段。
+pi-ai 回放状态用其成功 `AssistantMessage` 的带版本最小投影填充该结构:一个响应半区(源 API/提供方/模型、响应 ID/模型、可选的提供方原生 effort、停止原因),以及逐块的文本签名、thinking 签名和工具调用签名。[pi-ai 升级兼容性决定](../bug-fix/2026-09-05-pi-ai-upgrade-compatibility.zh.md#decision) 定义请求模型与 Anthropic 原生模型身份的区别,以及 effort 的保留规则。它不会重复 Harness 内容块中已有的文本或工具参数,也不包含诊断信息、时间戳、用量或错误。后续请求中,只有历史提供方和目标提供方当前归同一个适配器实例所有时,`LlmRuntime` 才会把回放状态交给目标适配器。适配器在能够恢复历史响应时,将 Harness 记录的内容与回放状态组合,并负责所需的跨模型或跨提供方转换。持久化内容保持权威:适配器收到无法使用的回放状态——未知 kind 或版本、格式错误的元数据、或与内容不再匹配的块结构——会把该消息降级为提供方无关转换并带出诊断;其他适配器只能收到提供方无关的内容以及提供方/模型字段。
 
 
 该状态属于模型可见的回放输入,因此遵循现有的[请求可重建规则](2026-07-05-reconstructable-requests.zh.md):它同时存在于终止 `finish` 分片和驱动派生的已组装 `assistant/message` 模型来源中。恢复和 fork 会原样保留该状态。压缩(compaction)遮蔽助手消息时,也会从活动 surface 中移除其回放状态;摘要属于普通的提供方无关内容。
 该状态属于模型可见的回放输入,因此遵循现有的[请求可重建规则](2026-07-05-reconstructable-requests.zh.md):它同时存在于终止 `finish` 分片和驱动派生的已组装 `assistant/message` 模型来源中。恢复和 fork 会原样保留该状态。压缩(compaction)遮蔽助手消息时,也会从活动 surface 中移除其回放状态;摘要属于普通的提供方无关内容。
 
 

+ 2 - 2
.agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.md
-2026-07-20-canonical-tool-output-contract.md: 52b56a44a747a07e13e05511d6d40a85fcd9696d
-2026-07-20-canonical-tool-output-contract.zh.md: a78defb2b45143d46fed2bedf760df12744621a5
+2026-07-20-canonical-tool-output-contract.md: f2c17325f77b93675086c39dd5a7693854b8521a
+2026-07-20-canonical-tool-output-contract.zh.md: d81c1730aab66df2dcf4eea4515f260df917ffb4

+ 1 - 1
.agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.md

@@ -34,7 +34,7 @@ type ToolExecutionResult =
 
 
 `tools/post-execute` has two mutually exclusive successful projections. Replacing `content` changes only Native/model presentation and preserves the canonical value and metadata. Replacing `value` revalidates the replacement and recomputes both presentation projections. A block removes the value and becomes a failure. Content replacement is therefore not a confidentiality mechanism: policy that must prevent programmatic access blocks the call or replaces the value.
 `tools/post-execute` has two mutually exclusive successful projections. Replacing `content` changes only Native/model presentation and preserves the canonical value and metadata. Replacing `value` revalidates the replacement and recomputes both presentation projections. A block removes the value and becomes a failure. Content replacement is therefore not a confidentiality mechanism: policy that must prevent programmatic access blocks the call or replaces the value.
 
 
-Canonical values are execution-local. The agent loop persists `tool/result` with only `content`, `error`, and optional `meta`; PTC mode's `tool/code-dispatch` persists the sub-call's rendered `content` and `isError`. Neither event stores the canonical intermediate value, so replay reproduces presentation but cannot reconstruct the programmatic result. When a tool declares `presentationMeta`, it is computed only for a direct surface call; a nested Code dispatch gets no metadata or result card. The outer `run_code` card instead reads final post-policy content and declares no presentation metadata. Generic and tool-owned spill projections similarly skip nested dispatches, whose canonical value never enters model context.
+Canonical values are execution-local. The agent loop persists `tool/result` with only `content`, `error`, and optional `meta`; PTC mode's `tool/code-dispatch` persists the sub-call's rendered `content` and `isError`. Neither event stores the canonical intermediate value, so replay reproduces presentation but cannot reconstruct the programmatic result. When a tool declares `presentationMeta`, it is computed only for a direct surface call; a nested Code dispatch gets no metadata. The Client can derive [nested terminal cards](../bug-fix/2026-09-05-nested-terminal-cards.md) from raw arguments and rendered content without that metadata. The outer `run_code` card instead reads final post-policy content and declares no presentation metadata. Generic and tool-owned spill projections similarly skip nested dispatches, whose canonical value never enters model context.
 
 
 The first-party tools preserve their existing Native text while returning domain DTOs:
 The first-party tools preserve their existing Native text while returning domain DTOs:
 
 

+ 1 - 1
.agents/notes/implemented/architecture/2026-07-20-canonical-tool-output-contract.zh.md

@@ -34,7 +34,7 @@ type ToolExecutionResult =
 
 
 `tools/post-execute` 为成功结果提供两种互斥的投影方式。替换 `content` 只改变 Native/模型展示,并保留规范值和元数据。替换 `value` 会重新校验替代值,并重新计算两份展示投影。阻止操作会移除值并转为失败。因此,替换内容并不是保密机制:必须阻止程序化访问的策略,应当阻止调用或替换值。
 `tools/post-execute` 为成功结果提供两种互斥的投影方式。替换 `content` 只改变 Native/模型展示,并保留规范值和元数据。替换 `value` 会重新校验替代值,并重新计算两份展示投影。阻止操作会移除值并转为失败。因此,替换内容并不是保密机制:必须阻止程序化访问的策略,应当阻止调用或替换值。
 
 
-规范值仅存在于执行期间。agent loop(智能体循环)持久化的 `tool/result` 只包含 `content`、`error` 和可选的 `meta`;PTC mode 的 `tool/code-dispatch` 持久化子调用渲染后的 `content` 与 `isError`。两个事件都不存储规范中间值,因此回放可以重现展示,却无法重建程序化结果。当工具声明 `presentationMeta` 时,系统只会为直接的外层调用计算它;嵌套 Code 分发没有元数据或结果卡片。外层 `run_code` 卡片则读取最终的 post-policy 内容,并且不声明展示元数据。通用以及工具自有的 spill 投影同样跳过嵌套分发,因为它们的规范值永远不会进入模型上下文。
+规范值仅存在于执行期间。agent loop(智能体循环)持久化的 `tool/result` 只包含 `content`、`error` 和可选的 `meta`;PTC mode 的 `tool/code-dispatch` 持久化子调用渲染后的 `content` 与 `isError`。两个事件都不存储规范中间值,因此回放可以重现展示,却无法重建程序化结果。当工具声明 `presentationMeta` 时,系统只会为直接的外层调用计算它;嵌套 Code 分发没有元数据。Client 可以从原始参数与渲染后的内容派生[嵌套 terminal 卡片](../bug-fix/2026-09-05-nested-terminal-cards.zh.md),无需这些元数据。外层 `run_code` 卡片则读取最终的 post-policy 内容,并且不声明展示元数据。通用以及工具自有的 spill 投影同样跳过嵌套分发,因为它们的规范值永远不会进入模型上下文。
 
 
 第一方工具在保持现有 Native 文本不变的同时返回领域 DTO:
 第一方工具在保持现有 Native 文本不变的同时返回领域 DTO:
 
 

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.md
-2026-08-10-session-log-version-mechanism.md: 98eb220e49457d3a2783edefce13d053c9b94025
-2026-08-10-session-log-version-mechanism.zh.md: 8853e6ed9a1a1ed43193cfe0949fffdbb9b4a26f
+2026-08-10-session-log-version-mechanism.md: 0f7f70b5ad6ecb2445729b1aa61fb3b295fc4ddb
+2026-08-10-session-log-version-mechanism.zh.md: a6d58505ad9fdb6068a1afe48f7250f020d20a9e

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.md

@@ -14,13 +14,13 @@ Session logs must be upgradable after release, and the runtime that ships first
 
 
 **The writer decides bumps, not the reader.** A bump is required exactly when an old runtime could no longer handle a new log with full semantic correctness. "Parses without error" is not the bar: silently skipping content that shapes reconstruction is a wrong read. Only structural changes qualify — header shape, event envelope, core event semantics, the surface mechanism (`SurfaceEventType` set, `SurfaceOp` variants). When unsure, bump: a near-identity upgrader is almost free, a missed bump silently corrupts old readers.
 **The writer decides bumps, not the reader.** A bump is required exactly when an old runtime could no longer handle a new log with full semantic correctness. "Parses without error" is not the bar: silently skipping content that shapes reconstruction is a wrong read. Only structural changes qualify — header shape, event envelope, core event semantics, the surface mechanism (`SurfaceEventType` set, `SurfaceOp` variants). When unsure, bump: a near-identity upgrader is almost free, a missed bump silently corrupts old readers.
 
 
-**Read rules by direction.** Equal version: read normally. Newer than the reader: refuse, name the direction ("written by a newer harness — upgrade"), and point at the raw log artifact so the user can still see the text (`SessionFormatUnsupportedError`, distinct from `SessionPersistenceCorruptionError` because nothing is damaged). Older than the reader: every event-body operation first runs the complete adjacent chain in memory, leaves the source path, bytes, and inode unchanged, exclusively publishes only the final current generation under its canonical versioned filename, and reopens it before current restoration. Header-only listing remains non-mutating and reports the numerically highest canonical generation. Catalog generation and module initialization reject a missing adjacent step, so a published first-party build never exposes a partial historical chain. Retained lower generations are not automatic fallback or a downgrade compatibility promise.
+**Read rules by direction.** Equal version: read normally. Newer than the reader: refuse, name the direction ("written by a newer harness — upgrade"), and point at the raw log artifact so the user can still see the text (`SessionFormatUnsupportedError`, distinct from `SessionPersistenceCorruptionError` because nothing is damaged). Older than the reader: every event-body operation first runs the complete adjacent chain in memory and leaves the source path, bytes, and inode unchanged. Read handles may consume that current logical result directly; a write open exclusively publishes the final current generation under its canonical versioned filename before append. Header-only listing remains non-mutating and reports the numerically highest canonical generation. Catalog generation and module initialization reject a missing adjacent step, so a published first-party build never exposes a partial historical chain. Retained lower generations are not automatic fallback or a downgrade compatibility promise.
 
 
 **A per-event `ignorable` marker covers vocabulary growth, so ordinary event additions never bump the version.** The event vocabulary is decided by which plugins are mounted, which a single version integer cannot describe. A reader meeting an unrecognized event type refuses to interpret the log unless the event carries `ignorable: true` in its envelope. The default is *required*: forgetting the marker over-refuses a resumable session (an inconvenience), while a default of ignorable would make the same mistake silently resume a gutted one (a safety failure). The architecture makes this sound: model-visible content flows only through the three `surfaceOp`-marked surface event types plus the `request/header`/`request/context` folds, so the dangerous unknowns are exactly the non-surface events that change how the rest of the log is read (`session/end-seed` is the existing example).
 **A per-event `ignorable` marker covers vocabulary growth, so ordinary event additions never bump the version.** The event vocabulary is decided by which plugins are mounted, which a single version integer cannot describe. A reader meeting an unrecognized event type refuses to interpret the log unless the event carries `ignorable: true` in its envelope. The default is *required*: forgetting the marker over-refuses a resumable session (an inconvenience), while a default of ignorable would make the same mistake silently resume a gutted one (a safety failure). The architecture makes this sound: model-visible content flows only through the three `surfaceOp`-marked surface event types plus the `request/header`/`request/context` folds, so the dangerous unknowns are exactly the non-surface events that change how the rest of the log is read (`session/end-seed` is the existing example).
 
 
 ## Consequences
 ## Consequences
 
 
-What shipped in v0 (release 0812): direction-aware refusal with the raw-log path; the unknown-event guard against a generated known-vocabulary list (`KNOWN_SESSION_EVENT_TYPES`, emitted by `gen-persistence-catalog` from every `SessionEventMap` merge and kept fresh by `verify-persistence-catalog`); the `ignorable` envelope field accepted by seed validation, JSONL, and the BFF wire schema. V1 adds the static adjacent catalog, the identity v0-to-v1 edge, header-only descriptors, exact-generation JSONL publication, and current-only restoration described in [Released Session formats](2026-08-31-released-session-format-migrations.md). V2 keeps the physical codec neutral to ordinary event vocabulary and payload additions: the adjacent edge freezes its released source and target inventories, while equal-version restoration applies the installed known-event set and current payload semantics. First-party writers do not set `ignorable` through `Session.append`, while a repository-external plugin is a current consumer; equal-version retention lives in the [external-plugin retention decision](2026-08-30-retain-ignorable-external-session-events.md), and the stricter historical rule lives in the [alpha migration refusal decision](2026-08-31-alpha-historical-unknown-event-refusal.md). The unknown-type guard remains read-side because append-time vocabulary refusal would stall a live session's durability. JSONL classifies foreign versions from the minimal raw header before current-header or event parsing, so a structurally different future format reports the upgrade direction instead of "corrupt".
+What shipped in v0 (release 0812): direction-aware refusal with the raw-log path; the unknown-event guard against a generated known-vocabulary list (`KNOWN_SESSION_EVENT_TYPES`, emitted by `gen-persistence-catalog` from every `SessionEventMap` merge and kept fresh by `verify-persistence-catalog`); the `ignorable` envelope field accepted by seed validation, JSONL, and the BFF wire schema. V1 adds the static adjacent catalog, the identity v0-to-v1 edge, header-only descriptors, exact-generation JSONL publication, and current-only restoration described in [Released Session formats](2026-08-31-released-session-format-migrations.md). [Historical Session read preparation](2026-09-05-read-only-session-migration-preparation.md) owns the JSONL timing between in-memory restoration and write publication. V2 keeps the physical codec neutral to ordinary event vocabulary and payload additions: the adjacent edge freezes its released source and target inventories, while equal-version restoration applies the installed known-event set and current payload semantics. First-party writers do not set `ignorable` through `Session.append`, while a repository-external plugin is a current consumer; equal-version retention lives in the [external-plugin retention decision](2026-08-30-retain-ignorable-external-session-events.md), and the stricter historical rule lives in the [alpha migration refusal decision](2026-08-31-alpha-historical-unknown-event-refusal.md). The unknown-type guard remains read-side because append-time vocabulary refusal would stall a live session's durability. JSONL classifies foreign versions from the minimal raw header before current-header or event parsing, so a structurally different future format reports the upgrade direction instead of "corrupt".
 
 
 ## Alternatives considered
 ## Alternatives considered
 
 

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-10-session-log-version-mechanism.zh.md

@@ -14,13 +14,13 @@ Session log 在发布后必须能升级格式,而最先发布的运行时决
 
 
 **升不升版本由写入方决定,与读取方能力无关。**当且仅当老运行时无法在语义上完全正确地处理新日志时才必须升版本。"解析不报错"不是标准:静默跳过影响重建的内容就是读错。只有结构性变更够得上这条线:header 形状、事件信封、核心事件语义、surface 机制(`SurfaceEventType` 集合、`SurfaceOp` 变体)。拿不准就升:近似恒等的升级器几乎没有成本,漏升一次会让老读取器静默读坏。
 **升不升版本由写入方决定,与读取方能力无关。**当且仅当老运行时无法在语义上完全正确地处理新日志时才必须升版本。"解析不报错"不是标准:静默跳过影响重建的内容就是读错。只有结构性变更够得上这条线:header 形状、事件信封、核心事件语义、surface 机制(`SurfaceEventType` 集合、`SurfaceOp` 变体)。拿不准就升:近似恒等的升级器几乎没有成本,漏升一次会让老读取器静默读坏。
 
 
-**读取规则按方向区分。**版本相等:正常读。比读取器新:拒绝,说明方向("由更新的 harness 写入,请升级"),并给出原始日志文件的路径,用户仍能看到文本(`SessionFormatUnsupportedError`,与 `SessionPersistenceCorruptionError` 区分,因为数据没有损坏)。比读取器旧:每个事件正文操作先在内存中运行完整相邻链,保持源路径、字节与 inode 不变,只在规范具名版本文件下排他发布最终当前 generation,再在当前恢复前重新打开。仅 header 的列表保持不变更,并报告数值最高的规范 generation。catalog 生成与模块初始化会拒绝缺失的相邻步骤,因此已发布第一方 build 绝不会暴露不完整历史链。保留的低 generation 不是自动 fallback,也不构成 downgrade compatibility 承诺。
+**读取规则按方向区分。**版本相等:正常读。比读取器新:拒绝,说明方向("由更新的 harness 写入,请升级"),并给出原始日志文件的路径,用户仍能看到文本(`SessionFormatUnsupportedError`,与 `SessionPersistenceCorruptionError` 区分,因为数据没有损坏)。比读取器旧:每个事件正文操作先在内存中运行完整相邻链,并保持源路径、字节与 inode 不变。读句柄可以直接使用该 current 逻辑结果;写 open 则在 append 前把最终 current generation 排他发布到其规范版本文件名。仅 header 的列表保持不变更,并报告数值最高的规范 generation。catalog 生成与模块初始化会拒绝缺失的相邻步骤,因此已发布第一方 build 绝不会暴露不完整历史链。保留的低 generation 不是自动 fallback,也不构成 downgrade compatibility 承诺。
 
 
 **逐事件的 `ignorable` 标记吸收词汇表增长,普通的新增事件永远不用升版本。**事件词汇表由挂载了哪些插件决定,单个版本整数描述不了它。读取器遇到不认识的事件类型时拒绝解读日志,除非该事件的信封带 `ignorable: true`。默认为必需:忘写标记的后果是把一个本可恢复的会话拒绝过头(体验问题),而默认可忽略会让同样的疏忽静默恢复出残缺会话(安全事故)。架构保证了这条规则成立:模型可见内容只经三种带 `surfaceOp` 标记的 surface 事件加 `request/header`、`request/context` 折叠进入重建,危险的未知事件恰好是那些不进 surface 但改变日志其余部分解读方式的事件(`session/end-seed` 是现存例子)。
 **逐事件的 `ignorable` 标记吸收词汇表增长,普通的新增事件永远不用升版本。**事件词汇表由挂载了哪些插件决定,单个版本整数描述不了它。读取器遇到不认识的事件类型时拒绝解读日志,除非该事件的信封带 `ignorable: true`。默认为必需:忘写标记的后果是把一个本可恢复的会话拒绝过头(体验问题),而默认可忽略会让同样的疏忽静默恢复出残缺会话(安全事故)。架构保证了这条规则成立:模型可见内容只经三种带 `surfaceOp` 标记的 surface 事件加 `request/header`、`request/context` 折叠进入重建,危险的未知事件恰好是那些不进 surface 但改变日志其余部分解读方式的事件(`session/end-seed` 是现存例子)。
 
 
 ## 影响
 ## 影响
 
 
-v0(0812 发布)交付的内容:分方向的拒绝并带原始日志路径;基于生成的已知词汇清单(`KNOWN_SESSION_EVENT_TYPES`,由 `gen-persistence-catalog` 从所有 `SessionEventMap` 声明合并生成,`verify-persistence-catalog` 保证新鲜)的未知事件守卫;`ignorable` 信封字段被种子校验、JSONL 和 BFF 线上 schema 接受。V1 添加静态相邻 catalog、恒等 v0-to-v1 迁移边、仅 header descriptor、精确代际 JSONL 发布与[已发布 Session 格式](2026-08-31-released-session-format-migrations.zh.md)定义的当前专用恢复。V2 让物理 codec 对普通事件词汇与 payload 新增项保持中立:相邻迁移边冻结 released source 与 target 清单,同版本恢复则应用已安装的 known-event set 与当前 payload 语义。第一方 writer 不通过 `Session.append` 设置 `ignorable`,而一个仓库外插件仍依赖该字段;同版本保留由[外部插件保留决策](2026-08-30-retain-ignorable-external-session-events.zh.md)定义,更严格的历史规则由 [alpha 迁移拒绝决策](2026-08-31-alpha-historical-unknown-event-refusal.zh.md)定义。未知类型守卫仍只在读取侧生效,因为 append 时的词汇拒绝会中断活跃 Session 的持久化。JSONL 会在当前 header 或事件解析前从最小原始 header 分类外来版本,因此结构完全不同的未来格式会报告升级方向而不是"损坏"。
+v0(0812 发布)交付的内容:分方向的拒绝并带原始日志路径;基于生成的已知词汇清单(`KNOWN_SESSION_EVENT_TYPES`,由 `gen-persistence-catalog` 从所有 `SessionEventMap` 声明合并生成,`verify-persistence-catalog` 保证新鲜)的未知事件守卫;`ignorable` 信封字段被种子校验、JSONL 和 BFF 线上 schema 接受。V1 添加静态相邻 catalog、恒等 v0-to-v1 迁移边、仅 header descriptor、精确代际 JSONL 发布与[已发布 Session 格式](2026-08-31-released-session-format-migrations.zh.md)定义的当前专用恢复。[历史 Session 只读迁移准备](2026-09-05-read-only-session-migration-preparation.zh.md)负责内存恢复与写入发布之间的 JSONL 时序。V2 让物理 codec 对普通事件词汇与 payload 新增项保持中立:相邻迁移边冻结 released source 与 target 清单,同版本恢复则应用已安装的 known-event set 与当前 payload 语义。第一方 writer 不通过 `Session.append` 设置 `ignorable`,而一个仓库外插件仍依赖该字段;同版本保留由[外部插件保留决策](2026-08-30-retain-ignorable-external-session-events.zh.md)定义,更严格的历史规则由 [alpha 迁移拒绝决策](2026-08-31-alpha-historical-unknown-event-refusal.zh.md)定义。未知类型守卫仍只在读取侧生效,因为 append 时的词汇拒绝会中断活跃 Session 的持久化。JSONL 会在当前 header 或事件解析前从最小原始 header 分类外来版本,因此结构完全不同的未来格式会报告升级方向而不是"损坏"。
 
 
 ## 曾考虑的替代方案
 ## 曾考虑的替代方案
 
 

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.md
-2026-08-22-single-dsh-application-launcher.md: 5729ed770b1d26969079e6341d2d2da4b17d0b90
-2026-08-22-single-dsh-application-launcher.zh.md: 45d5b9758fd1f01057de0f34e580249b63f9877e
+2026-08-22-single-dsh-application-launcher.md: 21fc99790c506f4e3f6e8b3ab572121e1cf1a882
+2026-08-22-single-dsh-application-launcher.zh.md: 4175bd675889ec01db46ebca157aff35c9ef9cab

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.md

@@ -34,7 +34,7 @@ Profile manifests own patch reload:
 
 
 Custom profiles default to `live`. A startup profile still applies its bundle, profile, home-level, and invocation `--patch` layers, but it does not watch them after boot. `dsh-base` inserts the module-HMR row disabled; a profile with a tested source-module reload lifecycle must enable it explicitly. None of the shipped profiles enable server module HMR: `patchReload: live` uses the launcher's config-only watcher while the startup profiles install no watcher. SDK and ACP cannot safely replace their server, agents, persistence, or tool registry inside one owned stdio connection.
 Custom profiles default to `live`. A startup profile still applies its bundle, profile, home-level, and invocation `--patch` layers, but it does not watch them after boot. `dsh-base` inserts the module-HMR row disabled; a profile with a tested source-module reload lifecycle must enable it explicitly. None of the shipped profiles enable server module HMR: `patchReload: live` uses the launcher's config-only watcher while the startup profiles install no watcher. SDK and ACP cannot safely replace their server, agents, persistence, or tool registry inside one owned stdio connection.
 
 
-The shipped protocol profiles reserve stdout for protocol frames, expose help without starting transport, and route stdin EOF and signals through bounded root disposal. ACP remains automation-only. The SDK JSON-RPC methods, notification fields, and `initialize.serverInfo.name` remain stable. Full-profile model-visible tool and persistence defaults come from `dsh-base`; `sdk-minimal` owns its explicit defaults. Runnable snapshots own the assembled application outputs.
+The shipped protocol profiles reserve stdout for protocol frames, expose help without starting transport, and route stdin EOF and signals through bounded root disposal. ACP remains automation-only. The SDK JSON-RPC methods, notification fields, and `initialize.serverInfo.name` remain stable. Full-profile model-visible tool and persistence defaults come from `dsh-base`, including its [default editor selection](../simplification/2026-09-05-base-default-file-editor.md); `sdk-minimal` owns its explicit defaults. Runnable snapshots own the assembled application outputs.
 
 
 ### TypeScript SDK customization
 ### TypeScript SDK customization
 
 

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-22-single-dsh-application-launcher.zh.md

@@ -34,7 +34,7 @@ Profile manifest 负责 patch 重载:
 
 
 自定义 profile 默认为 `live`。`startup` profile 仍会应用组合包、profile、home 级与调用时 `--patch` 各层,但启动后不会监视这些文件。`dsh-base` 插入的模块 HMR(热模块替换)配置项默认禁用;具有经过验证的源码模块重载生命周期的 profile 必须显式启用它。随附 profile 均不启用服务器模块 HMR:`patchReload: live` 使用启动器的仅配置 watcher,`startup` profile 则不安装 watcher。SDK 与 ACP 无法在一个自有 stdio 连接内安全替换其服务器、agent、持久化或工具注册表。
 自定义 profile 默认为 `live`。`startup` profile 仍会应用组合包、profile、home 级与调用时 `--patch` 各层,但启动后不会监视这些文件。`dsh-base` 插入的模块 HMR(热模块替换)配置项默认禁用;具有经过验证的源码模块重载生命周期的 profile 必须显式启用它。随附 profile 均不启用服务器模块 HMR:`patchReload: live` 使用启动器的仅配置 watcher,`startup` profile 则不安装 watcher。SDK 与 ACP 无法在一个自有 stdio 连接内安全替换其服务器、agent、持久化或工具注册表。
 
 
-随附协议 profile 将 stdout 保留给协议帧,显示帮助时不启动 transport,并通过有界根节点 dispose(资源释放)处理 stdin EOF 与信号。ACP 继续仅用于自动化。SDK JSON-RPC 方法、通知字段与 `initialize.serverInfo.name` 保持稳定。完整 profile 的模型可见工具与持久化默认值来自 `dsh-base`;`sdk-minimal` 拥有自己的显式默认值。可运行快照负责固定已组装的应用输出。
+随附协议 profile 将 stdout 保留给协议帧,显示帮助时不启动 transport,并通过有界根节点 dispose(资源释放)处理 stdin EOF 与信号。ACP 继续仅用于自动化。SDK JSON-RPC 方法、通知字段与 `initialize.serverInfo.name` 保持稳定。完整 profile 的模型可见工具与持久化默认值来自 `dsh-base`,包括其[默认编辑器选择](../simplification/2026-09-05-base-default-file-editor.zh.md);`sdk-minimal` 拥有自己的显式默认值。可运行快照负责固定已组装的应用输出。
 
 
 ### TypeScript SDK 自定义
 ### TypeScript SDK 自定义
 
 

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.md
-2026-08-23-client-derived-tool-presentation.md: ada444ab06bebd3e50018cea4a8dcde6b13956c4
-2026-08-23-client-derived-tool-presentation.zh.md: 8c1d60d5ed11e79b4428c7245f0db9c6b91c8b2c
+2026-08-23-client-derived-tool-presentation.md: 58a8f23d717580355b852703f448c723d9c3a7ea
+2026-08-23-client-derived-tool-presentation.zh.md: 8e1a02eff57cce7967985a22c5bdcec7c18e6059

+ 18 - 14
.agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.md

@@ -22,6 +22,8 @@ The required result is one raw Session journal and one Client presentation owner
 
 
 ## Decision
 ## Decision
 
 
+The visual-equivalence requirements below exclude the separately approved [nested terminal-card fix](../bug-fix/2026-09-05-nested-terminal-cards.md); all other presentation and ownership constraints remain.
+
 The Session Remote journal sends only raw, validated, persistable Session events. `session.page` and `session.follow` do not parse tool arguments, query the Tools registry, restore a presenter scope, execute `presentCall` or `presentResult`, or construct or clone any tool view.
 The Session Remote journal sends only raw, validated, persistable Session events. `session.page` and `session.follow` do not parse tool arguments, query the Tools registry, restore a presenter scope, execute `presentCall` or `presentResult`, or construct or clone any tool view.
 
 
 The Client Conversation layer continues to own tool call/result identity, pairing, lifecycle, Code Dispatch topology, and stable Chat Nodes. It does not interpret individual tool names or produce terminal, diff, read, search, or web component props.
 The Client Conversation layer continues to own tool call/result identity, pairing, lifecycle, Code Dispatch topology, and stable Chat Nodes. It does not interpret individual tool names or produce terminal, diff, read, search, or web component props.
@@ -49,7 +51,7 @@ The Host `ToolDefinition.presentCall`, `ToolDefinition.presentResult`, `ToolCall
 | Retained | the Session log format, Remote journal lifecycle, and Conversation identity/topology |
 | Retained | the Session log format, Remote journal lifecycle, and Conversation identity/topology |
 | Retained | the existing keyed slot, Generic fallback, and Chat, Details, and Trajectory structure |
 | Retained | the existing keyed slot, Generic fallback, and Chat, Details, and Trajectory structure |
 | Forbidden | a new Client presenter service, parallel registry, or wire renderer id |
 | Forbidden | a new Client presenter service, parallel registry, or wire renderer id |
-| Forbidden | new cards, visual redesign, interaction redesign, or Code Dispatch rich-card enhancements |
+| Forbidden | new cards, visual redesign, interaction redesign, or Code Dispatch rich-card enhancements except the [nested terminal-card exception](../bug-fix/2026-09-05-nested-terminal-cards.md) |
 | Forbidden | compatibility dual-writing, version negotiation, or retention of the old `view` field |
 | Forbidden | compatibility dual-writing, version negotiation, or retention of the old `view` field |
 
 
 ## Terminology
 ## Terminology
@@ -242,13 +244,13 @@ The Chat and Trajectory Tool Definitions read no views. They derive the followin
 
 
 ### Root and Code Dispatch subcalls
 ### Root and Code Dispatch subcalls
 
 
-Host presenter APIs describe top-level calls and results. Code Dispatch subcalls use the Generic, flattened Client presentation; recognizing a subcall name does not grant it a structured card.
+Host presenter APIs describe top-level calls and results. Code Dispatch subcalls retain Generic, flattened presentation for the diff, read, search, and web models covered here; supported terminal calls use the same eligibility rules as roots.
 
 
-Code Dispatch start and result events already carry `parentCallId`. Conversation preserves that existing fact on each child `ToolCallBlock`; root Session calls omit it. The five structured card models accept only blocks without `parentCallId`, while existing renderers that intentionally support nested calls continue receiving the same child block.
+Code Dispatch start and result events already carry `parentCallId`. Conversation preserves that existing fact on each child `ToolCallBlock`; root Session calls omit it. The diff, read, search, and web models accept only blocks without `parentCallId`; the terminal model and existing renderers that intentionally support nested calls accept child blocks.
 
 
-The Details panel delegates the selected block unchanged. The same card models observe `parentCallId` and keep a selected Code Dispatch child on the existing raw fallback, so the Details slot needs no placement field.
+The Details panel delegates the selected block unchanged. Shared card models apply the same terminal eligibility and nonterminal child restrictions in rows and Details, so the Details slot needs no placement field.
 
 
-The keyed slot continues dispatching every subcall by its real tool name. `parentCallId` controls only the terminal, diff, read, search, and web structured models covered by this decision. Existing specialized renderers such as Skill and Cordis, which already read raw blocks, remain unchanged.
+The keyed slot continues dispatching every subcall by its real tool name. `parentCallId` restricts only the diff, read, search, and web structured models covered by this decision. Existing specialized renderers such as Skill and Cordis, which already read raw blocks, remain unchanged.
 
 
 ### Missing call head
 ### Missing call head
 
 
@@ -294,7 +296,7 @@ The title, kind, rawInput, content, and locations from Generic Host `presentCall
 
 
 ### Terminal card
 ### Terminal card
 
 
-The Client terminal model derives existing `TerminalBlock` props from the tool name, call arguments, result content, error, existing `parentCallId`, and Session cwd.
+The Client terminal model derives existing `TerminalBlock` props from the tool name, call arguments, result content, error, and Session cwd, independently of `parentCallId`.
 
 
 | Input | Preserved result |
 | Input | Preserved result |
 |---|---|
 |---|---|
@@ -306,9 +308,9 @@ The Client terminal model derives existing `TerminalBlock` props from the tool n
 | settled persistent `bash`/`pwsh` | Generic flattened result, with no new exit card |
 | settled persistent `bash`/`pwsh` | Generic flattened result, with no new exit card |
 | foreground `terminal_send` | terminal prompt and output |
 | foreground `terminal_send` | terminal prompt and output |
 | background/error `terminal_send` | Generic result |
 | background/error `terminal_send` | Generic result |
-| Code Dispatch child | current flattened Generic form |
+| Code Dispatch child | same terminal eligibility and fallback rules as a root call |
 
 
-Standard shell results continue parsing trailing `[exit code: N]` and `[killed by signal: X]` markers. A parsed marker is removed from the body; timeout, sandbox denial, and markers without a pill remain in the body.
+Standard shell results parse trailing `[exit code: N]` and `[killed by signal: X]` markers. A final recognized spill-policy notice selects Generic output instead: expandable in shell rows and raw in Details, because the exit marker may be displaced or omitted. A parsed marker is removed from the terminal body; timeout, sandbox denial, and markers without a pill remain in the body.
 
 
 Call `description` remains above the card and overrides the collapsed summary. Workdir continues handling absolute, relative, and missing values. Relative paths resolve against the Session cwd while preserving normalization for `.`, `..`, drive letters, and UNC roots.
 Call `description` remains above the card and overrides the collapsed summary. Workdir continues handling absolute, relative, and missing values. Relative paths resolve against the Session cwd while preserving normalization for `.`, `..`, drive letters, and UNC roots.
 
 
@@ -418,7 +420,7 @@ The fixture does not import Host tool packages to compute page presentation and
 | grep/glob | current grouped/path card, truncation, and recovery |
 | grep/glob | current grouped/path card, truncation, and recovery |
 | web_search/web_fetch | current source/summary card and raw body |
 | web_search/web_fetch | current source/summary card and raw body |
 | Todo/Question/Skill/Cordis | current specialized rows |
 | Todo/Question/Skill/Cordis | current specialized rows |
-| Code Dispatch subcall | current Generic/flattened form |
+| Code Dispatch subcall | terminal cards when eligible; diff/read/search/web remain Generic/flattened |
 | Chat and Details | identical card fields for the same call |
 | Chat and Details | identical card fields for the same call |
 | Trajectory | current identity, tree, selection, and details |
 | Trajectory | current identity, tree, selection, and details |
 | Deliverables | current successful-mutation chips and links |
 | Deliverables | current successful-mutation chips and links |
@@ -526,7 +528,7 @@ This change does not promise to preserve differences expressed only through a Ho
 - search produces the pinned grouped/path card and recovery from metadata/content.
 - search produces the pinned grouped/path card and recovery from metadata/content.
 - web produces the pinned sources/fetch summary from metadata/content.
 - web produces the pinned sources/fetch summary from metadata/content.
 - unknown, malformed, error, missing-call, and missing-metadata cases remain Generic.
 - unknown, malformed, error, missing-call, and missing-metadata cases remain Generic.
-- absent and present `parentCallId` cases prove that structured presentation does not reach Code Dispatch descendants.
+- absent and present `parentCallId` cases prove equal terminal eligibility and preserve Generic fallback for diff, read, search, and web descendants.
 - Chat and Details produce identical card fields for the same block.
 - Chat and Details produce identical card fields for the same block.
 
 
 ### Deliverables
 ### Deliverables
@@ -582,12 +584,12 @@ Changes to this decision use `dsh-pre-push-checks` to select commands for the fi
 - Result metadata passes byte-for-byte through the log and Remote to the Client.
 - Result metadata passes byte-for-byte through the log and Remote to the Client.
 - Conversation assembles `ToolCallBlock` only from raw events.
 - Conversation assembles `ToolCallBlock` only from raw events.
 - `ToolCallBlock` contains no Host render-intent fields.
 - `ToolCallBlock` contains no Host render-intent fields.
-- The five structured card models read only raw blocks, their existing `parentCallId`, and Session path facts.
+- The five structured card models read only raw blocks and Session path facts; only diff, read, search, and web use `parentCallId` to reject children.
 - Generic, Todo, Question, Skill, and Cordis rows remain unchanged.
 - Generic, Todo, Question, Skill, and Cordis rows remain unchanged.
 - Deliverables does not depend on render intent and preserves current paths.
 - Deliverables does not depend on render intent and preserves current paths.
 - Text, components, expanded content, states, links, and ordering for all first-party top-level tools remain unchanged.
 - Text, components, expanded content, states, links, and ordering for all first-party top-level tools remain unchanged.
 - Malformed, missing-metadata, error, orphan, and unknown-tool cases continue to fall back safely.
 - Malformed, missing-metadata, error, orphan, and unknown-tool cases continue to fall back safely.
-- Code Dispatch subcalls remain Generic and flattened.
+- Code Dispatch diff, read, search, and web subcalls remain Generic and flattened; terminal subcalls follow root eligibility.
 - Chat, Details, and Trajectory behavior remains unchanged.
 - Chat, Details, and Trajectory behavior remains unchanged.
 - Existing Web browser expected outputs pass without refresh.
 - Existing Web browser expected outputs pass without refresh.
 - Host presenter APIs, implementations, and direct tests remain unchanged.
 - Host presenter APIs, implementations, and direct tests remain unchanged.
@@ -632,7 +634,7 @@ An on-demand RPC would turn one page read into N network calls and would still r
 
 
 ### Allow presentation enhancements
 ### Allow presentation enhancements
 
 
-The Client could produce more rich cards for Code Dispatch subcalls, missing call heads, or history whose Host presenter was unavailable. That would mix an ownership change with product behavior and prevent snapshots from proving equivalence, so this alternative is rejected.
+Bundling richer Code Dispatch cards, missing-call-head inference, or other historical presentation enhancements with the ownership change would prevent snapshots from proving equivalence. This decision rejects that coupling; the [nested terminal-card exception](../bug-fix/2026-09-05-nested-terminal-cards.md) does not relax nonterminal child restrictions.
 
 
 ### Accept temporary Generic degradation
 ### Accept temporary Generic degradation
 
 
@@ -684,6 +686,8 @@ The absence of optional `view` is a prerelease wire-type decision shared by all
 
 
 ## Relationship to Existing Decisions
 ## Relationship to Existing Decisions
 
 
+[Nested terminal cards](../bug-fix/2026-09-05-nested-terminal-cards.md) partially supersedes only the terminal child-card prohibition and its visual-equivalence requirement. This note remains active for raw-journal ownership, Client derivation, and the diff/read/search/web child restrictions.
+
 This note partially supersedes the implementation fact in [Client tool presentation ownership](../../archived/architecture/2026-08-08-client-tool-presentation-ownership.md) that “card models receive Host views.” Its core decisions remain: `ui-tool` owns presentation, business plugins use keyed slots, and Conversation owns only lifecycle and topology.
 This note partially supersedes the implementation fact in [Client tool presentation ownership](../../archived/architecture/2026-08-08-client-tool-presentation-ownership.md) that “card models receive Host views.” Its core decisions remain: `ui-tool` owns presentation, business plugins use keyed slots, and Conversation owns only lifecycle and topology.
 
 
 This note preserves [toolview dissolution](../../archived/architecture/2026-07-23-toolview-dissolution.md): the Client still has one slot registration model and does not restore `ToolViewRegistry`.
 This note preserves [toolview dissolution](../../archived/architecture/2026-07-23-toolview-dissolution.md): the Client still has one slot registration model and does not restore `ToolViewRegistry`.
@@ -699,7 +703,7 @@ This note preserves result metadata from the [canonical tool output contract](20
 ## Deferred
 ## Deferred
 
 
 - A separate explicit decision may evaluate deleting Host presenters if they remain without production consumers; this decision does not prejudge it.
 - A separate explicit decision may evaluate deleting Host presenters if they remain without production consumers; this decision does not prejudge it.
-- Specialized cards for Code Dispatch subcalls require a separate design and visible-snapshot updates; this decision preserves current behavior.
+- Specialized diff, read, search, and web cards for Code Dispatch subcalls require a separate design and visible-snapshot updates; terminal calls are covered by the linked partial supersession.
 - A third-party mutation tool that joins Deliverables requires a new Client-owned contribution; this decision does not create a registry for an absent consumer.
 - A third-party mutation tool that joins Deliverables requires a new Client-owned contribution; this decision does not create a registry for an absent consumer.
 - Distinct Client presentation for same-named providers first requires a stable, non-presentational identity; it must not restore per-page Host views.
 - Distinct Client presentation for same-named providers first requires a stable, non-presentational identity; it must not restore per-page Host views.
 - If Client card-model performance needs measurement, an immutable-block microbenchmark can be added; the shipped architecture already prohibits scanning the Session window.
 - If Client card-model performance needs measurement, an immutable-block microbenchmark can be added; the shipped architecture already prohibits scanning the Session window.

+ 18 - 14
.agents/notes/implemented/architecture/2026-08-23-client-derived-tool-presentation.zh.md

@@ -22,6 +22,8 @@ Host presenter 与 Client keyed renderer 分担展示会形成对同一事件的
 
 
 ## Decision
 ## Decision
 
 
+下述展示对等要求不包含已独立批准的[嵌套 terminal 卡片修复](../bug-fix/2026-09-05-nested-terminal-cards.zh.md);其他展示与所有权约束全部保留。
+
 Session Remote journal 只下发原始、已验证、可持久化的 Session event。`session.page` 和 `session.follow` 不解析工具参数,不查询 Tools registry,不恢复 presenter scope,不执行 `presentCall`/`presentResult`,也不构造或克隆任何 tool view。
 Session Remote journal 只下发原始、已验证、可持久化的 Session event。`session.page` 和 `session.follow` 不解析工具参数,不查询 Tools registry,不恢复 presenter scope,不执行 `presentCall`/`presentResult`,也不构造或克隆任何 tool view。
 
 
 Client Conversation 层继续负责工具调用与结果的 identity、配对、生命周期、Code Dispatch 拓扑和稳定 Chat Node。它不解释具体工具名称,也不生成 terminal、diff、read、search 或 web 组件 props。
 Client Conversation 层继续负责工具调用与结果的 identity、配对、生命周期、Code Dispatch 拓扑和稳定 Chat Node。它不解释具体工具名称,也不生成 terminal、diff、read、search 或 web 组件 props。
@@ -49,7 +51,7 @@ Host 的 `ToolDefinition.presentCall`、`ToolDefinition.presentResult`、`ToolCa
 | 保留 | Session 日志格式、Remote journal 生命周期与 Conversation identity/topology |
 | 保留 | Session 日志格式、Remote journal 生命周期与 Conversation identity/topology |
 | 保留 | 现有 keyed slot、Generic fallback、Chat、Details 与 Trajectory 结构 |
 | 保留 | 现有 keyed slot、Generic fallback、Chat、Details 与 Trajectory 结构 |
 | 禁止 | 新 Client presenter service、平行 registry 或 wire renderer id |
 | 禁止 | 新 Client presenter service、平行 registry 或 wire renderer id |
-| 禁止 | 新卡片、视觉改版、交互改版或 Code Dispatch rich-card 增强 |
+| 禁止 | 新卡片、视觉改版、交互改版或 Code Dispatch rich-card 增强,[嵌套 terminal 卡片例外](../bug-fix/2026-09-05-nested-terminal-cards.zh.md)除外 |
 | 禁止 | 为兼容保留双写、版本协商或旧 `view` 字段 |
 | 禁止 | 为兼容保留双写、版本协商或旧 `view` 字段 |
 
 
 ## 术语
 ## 术语
@@ -242,13 +244,13 @@ Chat 和 Trajectory 的 Tool Definition 都不读取 view,而从事件生成
 
 
 ### Root 与 Code Dispatch 子调用
 ### Root 与 Code Dispatch 子调用
 
 
-Host presenter API 描述顶层 call/result。Code Dispatch 子调用使用 Generic/flattened Client 展示;Client 能识别子调用名称并不赋予它结构化卡片。
+Host presenter API 描述顶层 call/result。本决定覆盖的 diff、read、search 和 web model 对 Code Dispatch 子调用保留 Generic/flattened 展示;受支持的 terminal 调用使用与根调用相同的适用规则。
 
 
-Code Dispatch start 与 result event 已经携带 `parentCallId`。Conversation 在每个 child `ToolCallBlock` 上保留这项现有事实,root Session call 则不携带它。五类结构化 card model 只接受没有 `parentCallId` 的 block,原本有意支持嵌套调用的 renderer 则继续收到同一个 child block。
+Code Dispatch start 与 result event 已经携带 `parentCallId`。Conversation 在每个 child `ToolCallBlock` 上保留这项现有事实,root Session call 则不携带它。diff、read、search 和 web model 只接受没有 `parentCallId` 的 block;terminal model 与原本有意支持嵌套调用的 renderer 接受 child block。
 
 
-Details panel 原样委托选中的 block。同一组 card model 读取 `parentCallId`,让选中的 Code Dispatch child 保持现有 raw fallback,因此 Details slot 不需要 placement 字段。
+Details panel 原样委托选中的 block。共享 card model 在行与 Details 中应用相同的 terminal 适用规则和非 terminal 子调用限制,因此 Details slot 不需要 placement 字段。
 
 
-keyed slot 仍按每个子调用的真实 tool name 分发;`parentCallId` 只控制本决定覆盖的 terminal/diff/read/search/web 结构化模型。Skill、Cordis 等已经直接读取 raw block 的专用 renderer 保持现状。
+keyed slot 仍按每个子调用的真实 tool name 分发;`parentCallId` 只限制本决定覆盖的 diff/read/search/web 结构化模型。Skill、Cordis 等已经直接读取 raw block 的专用 renderer 保持现状。
 
 
 ### 缺失调用头
 ### 缺失调用头
 
 
@@ -294,7 +296,7 @@ Generic Host `presentCall` 的 title、kind、rawInput、content 与 locations 
 
 
 ### Terminal 卡片
 ### Terminal 卡片
 
 
-Client terminal model 从工具名称、调用参数、结果 content、error、现有 `parentCallId` 与 Session cwd 派生现有 `TerminalBlock` props。
+Client terminal model 从工具名称、调用参数、结果 content、error 与 Session cwd 派生现有 `TerminalBlock` props,不依赖 `parentCallId`。
 
 
 | 输入 | 保持的结果 |
 | 输入 | 保持的结果 |
 |---|---|
 |---|---|
@@ -306,9 +308,9 @@ Client terminal model 从工具名称、调用参数、结果 content、error、
 | persistent `bash`/`pwsh` settled | Generic flattened result,不新增 exit card |
 | persistent `bash`/`pwsh` settled | Generic flattened result,不新增 exit card |
 | `terminal_send` 前台 | terminal prompt 与 output |
 | `terminal_send` 前台 | terminal prompt 与 output |
 | `terminal_send` background/error | Generic 结果 |
 | `terminal_send` background/error | Generic 结果 |
-| Code Dispatch child | 当前 flattened Generic 形态 |
+| Code Dispatch child | 与根调用相同的 terminal 适用规则与 fallback 规则 |
 
 
-标准 shell 结果继续解析末尾 `[exit code: N]` 与 `[killed by signal: X]`。已解析的 marker 从正文移除;timeout、sandbox denial 与没有 pill 的 marker 留在正文。
+标准 shell 结果解析末尾 `[exit code: N]` 与 `[killed by signal: X]`。末尾已识别的 spill 策略提示会改用 Generic 输出:在 shell 行中可展开,在 Details 中显示原文,因为退出标记可能被移位或省略。已解析的 marker 从 terminal 正文移除;timeout、sandbox denial 与没有 pill 的 marker 留在正文。
 
 
 调用 `description` 继续显示在 card 上方并覆盖折叠摘要。workdir 继续按绝对、相对和缺失三种情况处理;相对路径基于 Session cwd,且保留 `.`、`..`、盘符与 UNC root 的归一化。
 调用 `description` 继续显示在 card 上方并覆盖折叠摘要。workdir 继续按绝对、相对和缺失三种情况处理;相对路径基于 Session cwd,且保留 `.`、`..`、盘符与 UNC root 的归一化。
 
 
@@ -418,7 +420,7 @@ fixture 不导入 Host 工具包来计算页面展示,也不保留 presenter 
 | grep/glob | 当前 grouped/path card、截断与 recovery |
 | grep/glob | 当前 grouped/path card、截断与 recovery |
 | web_search/web_fetch | 当前来源/摘要 card 与原始正文 |
 | web_search/web_fetch | 当前来源/摘要 card 与原始正文 |
 | Todo/Question/Skill/Cordis | 当前专用行 |
 | Todo/Question/Skill/Cordis | 当前专用行 |
-| Code Dispatch subcall | 当前 Generic/flattened 形态 |
+| Code Dispatch subcall | 满足条件时显示 terminal 卡片;diff/read/search/web 保持 Generic/flattened |
 | Chat 与 Details | 同一调用使用相同 card fields |
 | Chat 与 Details | 同一调用使用相同 card fields |
 | Trajectory | 当前 identity、树、选择和 details |
 | Trajectory | 当前 identity、树、选择和 details |
 | Deliverables | 当前成功 mutation chips 与链接 |
 | Deliverables | 当前成功 mutation chips 与链接 |
@@ -526,7 +528,7 @@ Host registry 允许不同 scope 为同一 tool name 提供不同定义;Sessio
 - search 用 meta/content 得到已固定的 grouped/path card 与 recovery。
 - search 用 meta/content 得到已固定的 grouped/path card 与 recovery。
 - web 用 meta/content 得到已固定的 sources/fetch summary。
 - web 用 meta/content 得到已固定的 sources/fetch summary。
 - unknown、malformed、error、missing-call 与 missing-meta 继续 Generic。
 - unknown、malformed、error、missing-call 与 missing-meta 继续 Generic。
-- `parentCallId` 缺失与存在的用例证明结构化展示不会到达 Code Dispatch descendant。
+- `parentCallId` 缺失与存在的用例证明 terminal 适用规则一致,并保留 diff、read、search 和 web descendant 的 Generic fallback。
 - Chat 与 Details 对同一 block 得到相同 card fields。
 - Chat 与 Details 对同一 block 得到相同 card fields。
 
 
 ### Deliverables
 ### Deliverables
@@ -582,12 +584,12 @@ Host registry 允许不同 scope 为同一 tool name 提供不同定义;Sessio
 - result meta 逐字节通过日志与 Remote 到达 Client。
 - result meta 逐字节通过日志与 Remote 到达 Client。
 - Conversation 只从 raw event 组装 ToolCallBlock。
 - Conversation 只从 raw event 组装 ToolCallBlock。
 - ToolCallBlock 不含 Host render-intent 字段。
 - ToolCallBlock 不含 Host render-intent 字段。
-- 五类结构化 card model 只读 raw block、其现有 `parentCallId` 与 Session path facts。
+- 五类结构化 card model 只读 raw block 与 Session path facts;只有 diff、read、search 和 web 使用 `parentCallId` 拒绝子调用。
 - Generic、Todo、Question、Skill 与 Cordis 行行为不变。
 - Generic、Todo、Question、Skill 与 Cordis 行行为不变。
 - Deliverables 不依赖 render intent 且保持当前 paths。
 - Deliverables 不依赖 render intent 且保持当前 paths。
 - 所有第一方顶层工具的文本、组件、展开内容、状态、链接与排序不变。
 - 所有第一方顶层工具的文本、组件、展开内容、状态、链接与排序不变。
 - malformed、missing-meta、error、orphan 与 unknown-tool 继续安全 fallback。
 - malformed、missing-meta、error、orphan 与 unknown-tool 继续安全 fallback。
-- Code Dispatch 子调用保持 Generic/flattened。
+- Code Dispatch 的 diff、read、search 和 web 子调用保持 Generic/flattened;terminal 子调用遵循根调用适用规则。
 - Chat、Details 与 Trajectory 行为不变。
 - Chat、Details 与 Trajectory 行为不变。
 - 现有 Web browser expected 无需刷新即可通过。
 - 现有 Web browser expected 无需刷新即可通过。
 - Host presenter API、实现与直接测试不变。
 - Host presenter API、实现与直接测试不变。
@@ -632,7 +634,7 @@ read 行结构、applied diff、search 分组、web sources 和有效 truncation
 
 
 ### 允许展示增强
 ### 允许展示增强
 
 
-Client 可以为 Code Dispatch 子调用、缺失 call head 或 Host presenter 不可用的历史生成更多 rich card,但这会混淆 ownership 变化与产品行为,并使快照无法证明对等,因此拒绝。
+将更丰富的 Code Dispatch 卡片、缺失 call head 的推断或其他历史展示增强与所有权变更捆绑,会使快照无法证明对等。本决定拒绝这种捆绑;[嵌套 terminal 卡片例外](../bug-fix/2026-09-05-nested-terminal-cards.zh.md)不放宽非 terminal 子调用限制。
 
 
 ### 接受临时 Generic 退化
 ### 接受临时 Generic 退化
 
 
@@ -684,6 +686,8 @@ optional `view` 的缺失是所有 consumer 共同遵守的预发布 wire 类型
 
 
 ## 与现有决策的关系
 ## 与现有决策的关系
 
 
+[嵌套 terminal 卡片](../bug-fix/2026-09-05-nested-terminal-cards.zh.md)仅部分取代 terminal 子调用卡片禁令及其展示对等要求。本文继续负责原始 journal 所有权、Client 派生以及 diff/read/search/web 子调用限制。
+
 本文部分取代 [Client 工具展示所有权](../../archived/architecture/2026-08-08-client-tool-presentation-ownership.md) 中“card model 接收 Host view”的实现事实;`ui-tool` 拥有展示、业务插件使用 keyed slot、Conversation 只拥有生命周期与拓扑的核心决定保持不变。
 本文部分取代 [Client 工具展示所有权](../../archived/architecture/2026-08-08-client-tool-presentation-ownership.md) 中“card model 接收 Host view”的实现事实;`ui-tool` 拥有展示、业务插件使用 keyed slot、Conversation 只拥有生命周期与拓扑的核心决定保持不变。
 
 
 本文保留 [toolview 溶解](../../archived/architecture/2026-07-23-toolview-dissolution.md) 的决定:Client 仍只有 slot 注册模型,不恢复 `ToolViewRegistry`。
 本文保留 [toolview 溶解](../../archived/architecture/2026-07-23-toolview-dissolution.md) 的决定:Client 仍只有 slot 注册模型,不恢复 `ToolViewRegistry`。
@@ -699,7 +703,7 @@ optional `view` 的缺失是所有 consumer 共同遵守的预发布 wire 类型
 ## Deferred
 ## Deferred
 
 
 - Host presenter 若长期没有生产消费者,可由另一项明确决策评估删除;本决定不预判。
 - Host presenter 若长期没有生产消费者,可由另一项明确决策评估删除;本决定不预判。
-- Code Dispatch 子调用若要专用卡片,需单独设计并更新可见快照;本决定保持现状。
+- Code Dispatch 子调用的 diff、read、search 和 web 专用卡片仍需独立设计并更新可见快照;terminal 调用由链接的部分取代决策负责。
 - 第三方 mutation tool 若要加入 Deliverables,需新增 Client-owned 贡献;本决定不为尚无消费者的扩展性建 registry。
 - 第三方 mutation tool 若要加入 Deliverables,需新增 Client-owned 贡献;本决定不为尚无消费者的扩展性建 registry。
 - 同名 provider 若要不同 Client 展示,需先定义稳定、非展示性的 identity;不得恢复按页 Host view。
 - 同名 provider 若要不同 Client 展示,需先定义稳定、非展示性的 identity;不得恢复按页 Host view。
 - Client card model 若需量化性能,可以增加 immutable-block 微基准;已交付架构禁止扫描 Session window。
 - Client card model 若需量化性能,可以增加 immutable-block 微基准;已交付架构禁止扫描 Session window。

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.md
-2026-08-31-released-session-format-migrations.md: 2c75d0b57a0b513c218b6a67b8c2b31c7cae4d0f
-2026-08-31-released-session-format-migrations.zh.md: d88c643cbfaf7f3d4f52ca6e5fa244a26917eb99
+2026-08-31-released-session-format-migrations.md: 592322c0e4c1b2fa52dcf71652f3878f43a8c8ca
+2026-08-31-released-session-format-migrations.zh.md: ba2317903845739cda8da1c01c2f959c7a2ccd50

+ 149 - 25
.agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.md

@@ -1,4 +1,4 @@
-# Agent Note: Released Session formats migrate on body read through adjacent pure edges
+# Agent Note: Released Session formats migrate through stateful streaming stages
 
 
 Status: implemented
 Status: implemented
 
 
@@ -6,52 +6,176 @@ English | [中文](2026-08-31-released-session-format-migrations.zh.md)
 
 
 ## Problem
 ## Problem
 
 
-Session format v0 shipped in an alpha release, so a structural writer change can no longer treat existing JSONL as disposable pre-release state. Stored event bodies reach consumers through read or write `SessionHandle` instances used by resume, query, export, fork, and continuation paths. Migrating only one consumer would let callers observe different logical generations or fail only when a later writer reaches the old file.
+Session format v0 shipped in an alpha release, so a structural writer change can no longer treat existing JSONL as disposable pre-release state. The first whole-artifact migration implementation made those logs convertible, but its data model turned a 116 MB real Session into an operation that exhausted a 16 GB Node process before returning a handle.
 
 
-Migration must retain the exact source path, bytes, and inode, including a torn physical tail, while giving every published format one unambiguous canonical filename. Plain JSONL and Zstandard are encoding choices for the same logical format and must not create parallel migration implementations.
+### Whole-artifact performance failure
+
+- Zstandard input was split into 317,540 frames and each frame used a separate asynchronous decompression call. The implementation retained every plaintext frame and then concatenated them before JSON parsing, creating the same number of Promise, thread-pool, and native decode transitions.
+- Physical Decode materialized a complete plaintext Buffer, one complete string, every JSONL row, expanded source events, migrated target events, encoded target rows, a joined target string, and target physical Buffers at overlapping points in the same request.
+- Every codec and migration edge called `snapshotSessionFormatJson()` or `snapshotSessionFormatArtifact()`. These operations detached, recursively copied, and deeply froze whole headers, rows, payloads, and event arrays before and after adjacent migrations.
+- Released packed Assistant chunks expanded into about 9.14 million logical v0/v1 events before v1-to-v2 folded them into 72,784 current events. The whole-artifact API required both representations and the old-to-new sequence map to coexist.
+- Encoding built the complete JSONL and compressed output in memory. The successful path then decoded the staged target, decoded the committed target, and decoded it again in persistence to construct the business object; it also reread the source for a full fingerprint comparison.
+- Per-frame `await` calls did not provide useful bounded scheduling. The pre-migration reader instead reused one synchronous decoder and yielded from the outer loop about every 500 ms, avoiding hundreds of thousands of asynchronous transitions.
+
+### The interfaces prevented local fixes from composing
+
+`SessionFormatCodec` decoded and encoded complete arrays, each adjacent `SessionFormatMigration` accepted and returned a complete `SessionFormatArtifact`, and the compiled chain could only hand one materialized artifact to the next edge. A faster physical decoder therefore still encountered source-row arrays, expanded-event arrays, per-edge snapshots, and target-row arrays downstream.
+
+The migrations are stateful even though the API presented them as one-shot functions. v0-to-v1 tracks message and retry identity. v1-to-v2 buffers one unsettled Assistant attempt, tracks events blocked behind it, and maintains old-to-new sequence references. Wrapping that state in closures or push/finish helper objects made the runtime structure different from the static declarations and made production, Worker verification, fixtures, and replay use different entry paths.
 
 
 ## Decision
 ## Decision
 
 
-`SESSION_FORMAT_VERSION` is a monotonic current-writer integer. One profile-independent pure package owns each adjacent `vN -> vN+1` conversion. `@deepseek-ai/dsh-session-format` supplies only lossless snapshots, unique gap-free planning, header-only conversion, and whole-artifact composition; `@deepseek-ai/dsh-session-format-catalog` statically imports the complete chain independently of mounted Cordis plugins. Historical codecs and normalizers live in the named edge package, while current Session and persistence code accept only the latest logical types.
+The Session format packages use a stateful synchronous Stage API. Static migration declarations describe one adjacent version edge and create a new stage for each restored artifact. A stage owns that artifact's mutable state; no stage instance is shared across Sessions.
 
 
-Each edge freezes strict source and target semantics, while its target physical codec remains vocabulary-neutral so ordinary event growth can stay within one format version. The catalog restores the final generation through the installed peer `@deepseek-ai/dsh-session` and its current `KNOWN_SESSION_EVENT_TYPES`, preventing a frozen historical edge from becoming the current vocabulary owner.
+### Stage and Context protocol
 
 
-The JSONL provider completes ensure-current work before `open` returns a handle for a stored Session. It selects the highest canonical generation, migrates a supported historical body, and decodes the current result from one physical snapshot; the public `SessionPersistence` and `SessionHandle` interfaces contain no migration operations. Header-only `stat` and `list` rescan Session directories, translate supported historical headers in memory, and never publish a successor. `create` checks canonical filenames independently of header readability, so every existing generation reserves its Session id.
+```text
+interface SessionFormatMigrationContext {
+  emitEvent(event: SessionFormatEvent): void
+  emitRun(run: SessionFormatEventRun): void
+}
 
 
-Cancellation belongs to the `open`, `stat`, or `list` call that supplied it. Discovery, stable reads, decoding, and pre-publication checks observe that signal; once an immutable successor is published and its directory entry is synced, later cancellation does not delete the committed generation.
+interface SessionFormatMigrationStage {
+  readonly headerInheritedEventCount?: number
+  transformEvent(
+    event: SessionFormatEvent,
+    context: SessionFormatMigrationContext,
+  ): void
+  transformRun(
+    run: SessionFormatEventRun,
+    context: SessionFormatMigrationContext,
+  ): void
+  finish(context: SessionFormatMigrationContext): number
+}
+```
 
 
-The configured JSONL encoding owns one full suffix, `.jsonl` or `.jsonl.zstd`. Migration reads a stable exact source, decodes the recoverable logical prefix, composes every required edge in memory, validates and syncs a same-directory temporary stage for only the final target, rechecks the source fingerprint, publishes that previously absent target without overwrite, syncs the namespace, and reopens it through current validation before returning a handle. The source never moves or changes; only disposable temporary stages may be moved, linked, or removed. Migration does not synthesize interrupted-turn events: agent-loop appends those repairs through the write handle, while read-only query paths balance them in memory.
+`SessionFormatMigrationContext.emitEvent()` and `emitRun()` are synchronous. The producer declares whether it emits a scalar event or a compact run, so the hot path never infers the category from properties on a parsed file object. The caller owns scheduling and supplies the context to each operation instead of injecting a callback into the stage constructor. One input may emit zero, one, or many outputs without allocating a temporary return array or retaining an internal output queue.
 
 
-Canonical filenames encode the physical format generation: v0 is `session.jsonl` or `session.jsonl.zstd`; every positive generation is lowercase `session.vN.jsonl` or `session.vN.jsonl.zstd`. `dsh-session-format` owns the raw basename rule (`sessionFormatLogFilename`, `parseSessionFormatLogFilename`); the JSONL provider, the session-log export archive, and recorded-session fixtures append only the compression suffix. Publication never renames, replaces, or deletes a committed generation path. If the target already exists, it is accepted only as a regular current-format file with exactly the expected bytes; any other target refuses. Lower generations remain for operator inspection or explicit copying, but normal runtime operations select the numerically highest canonical name and never use retained predecessors as automatic fallback, restore, or downgrade support.
+`SessionFormatMigration` remains an immutable declaration: version numbers, header migration, target-header validation, and `createStage()`. `CompiledSessionFormatChain` validates a unique gap-free edge sequence once, creates per-artifact stages in source-to-target order, and connects them with context objects in reverse order. `finish()` settles stages in source-to-target order so each stage can emit its tail before the downstream stage closes.
 
 
-The current-format fast path classifies the header from one stable source snapshot, invokes no historical converter or generation write, and passes that snapshot to current decoding without another file read. The decoded log enters the existing bounded revision-keyed memo for an immediate observe-to-resume handoff, while `stat` and `list` deliberately rescan. Multiple edges leave the original generation unchanged and publish only the final target; intermediate versions exist only in memory. A source fingerprint recheck restarts migration when content changes, and exclusive target publication accepts a racing winner only when its bytes match exactly. Cross-process append fencing remains outside this guarantee.
+```text
+JSONL record
+  → released physical row decoder
+  → v0-to-v1 stage
+  → v1-to-v2 stage
+  → current event collector
+```
 
 
-The first edge, `@deepseek-ai/dsh-session-format-v0-to-v1`, is intentionally identity-shaped: aside from the version and bounded historical normalizations already accepted by v0, it preserves logical headers, events, sequence numbers, references, timestamps, payloads, and the configured compression choice. The exact `session.jsonl[.zstd]` source remains byte- and inode-identical, while the current writer encodes the new `session.v1.jsonl[.zstd]` successor. This exercises the complete publication lifecycle before a cardinality-changing format needs it.
+The chain contains no `flatMap`, spread expansion, intermediate event array, or scheduler. The final event collector expands a compact run only after every migration stage has had the opportunity to consume it directly.
 
 
-Projection-cache records bind their fold to the Session header's `formatVersion`. The `session_projcache` v7 reader may load predecessor domain records structurally, but a record without the format generation cannot seed a current Session; the authoritative log refolds it and the next checkpoint writes the complete current identity. This prevents a cache row produced before a bounded normalizer or cardinality-changing edge from bypassing that migration.
+### Physical codecs and packed runs
 
 
-## Consequences
+Each released codec creates a row decoder with explicit `strict` or `recoverable` recovery. The decoder validates and emits one event or one codec-owned `SessionFormatEventRun` at a time through separate context methods. v0-to-v1 and v1-to-v2 implement both `transformEvent()` and `transformRun()`, so packed Assistant chunks can reach the folding edge without first becoming millions of ordinary events.
+
+The v0-to-v1 edge preserves logical headers, sequence numbers, references, timestamps, and payloads except for bounded released-v0 normalizations. It translates the retired `steering/message` and `compact/*` event names, accepts a released `llm/retry` after its matching `step/end`, deterministically supplies a missing `llm/retry.retryId` per turn/step/provider/policy chain, and supplies one deterministic `compactionId` across a legacy compaction group that omitted it. The v1-to-v2 edge owns attempt folding and reference remapping, and emits only settled current events. It splits a legacy goal-sourced user message into `goal/change` plus the original model-visible message. It also inserts an interrupted `turn/end` for the bounded released restart in which an open turn with no open step is followed by a non-empty `next-turn` inbox splice and the next numbered `turn/start`.
+
+The catalog exposes one `createRestore()` operation for production, Worker, fixture, and replay callers. Recovery policy and final validation policy are chosen once at restore creation. Historical production uses recoverable source parsing with transformed-current validation; this validates the released current result after migration, while input that is already current receives only codec validation. Worker and fixture verification use strict parsing with full installed current restoration. A migration-stage or transformed-current validation refusal remains `SessionFormatUnsupportedMigrationError`; physical decoding failures remain corruption. Test support keeps only fixture-specific token and envelope materialization.
+
+### JSONL integration
+
+The JSONL provider scans frame boundaries once, reuses one Zstandard decoder, parses complete JSONL records incrementally, and feeds rows directly into the catalog restore. The outer loop yields at a bounded cadence; there is no per-frame `await` and no complete plaintext or source-row array.
 
 
-Reading event bodies with a newer build may durably add a higher generation. The exact old generation remains available, but the runtime thereafter selects the highest canonical filename; retention does not promise that an older build can safely downgrade or that the newer build will fall back when the successor is corrupt. A read-only filesystem reports an actionable migration failure instead of returning an in-memory current view that differs from disk.
+Current encoding is record based. The provider serializes about 1 MiB of plaintext per main-thread slice, streams it through one Zstandard context with source-error propagation, writes compressed output in 4 MiB batches to an exclusively created same-directory temporary file, and syncs it before publication. A process-wide scheduler admits at most two full verification Workers and hands a released permit directly to the oldest waiter.
 
 
-JSONL publication uses POSIX hard-link creation plus directory sync, and Windows uses no-overwrite `MoveFileExW` with write-through. A competing writer that wins target creation is accepted only when the committed bytes exactly match. One process-local writer per Session is the supported concurrency model. A future per-Session cross-process lock can close the remaining source-check-to-publication race without changing the format edge interface.
+Preparation forwards cancellation through source reads and observes it at the existing approximately 500 ms Decode yield boundary. Once `publish()` starts, encode, Worker verification, and publication do not receive caller cancellation and run to settlement; write open checks its caller signal again afterward. A published generation is never rolled back.
 
 
-Retained generations are not a live-stream write-ahead log. A future optional WAL sidecar may preserve unfinished assistant streams across a hard crash. Explicit generation inspection or copying, retention tooling, compression conversion, and streamed whole-artifact transformation are separate features; automatic fallback and downgrade compatibility are not implied future work.
+The Stage pipeline ends at one prepared current artifact. [Historical Session read preparation](2026-09-05-read-only-session-migration-preparation.md) defines how read open consumes that artifact immediately while write open performs encode, verification, and publication before returning append access.
 
 
-This note supersedes the continue-only persistence rule and the deferred-chain status in [Session log versioning](2026-08-10-session-log-version-mechanism.md). That note remains the authority for when to bump the version and for ordinary equal-version `ignorable` event behavior.
+### Durable format and publication rules
+
+Canonical filenames encode physical format generation: v0 is `session.jsonl[.zstd]` and positive generations use `session.vN.jsonl[.zstd]`. Migration never moves, replaces, or deletes a committed generation and writes only the final current target; intermediate versions exist only as stage state.
+
+POSIX publication uses hard-link creation plus directory sync. Windows uses no-overwrite, write-through `MoveFileExW`. An existing target is accepted only when its verified migration prefix equals the staged bytes; any append tail belongs to current-generation reading rather than migration winner verification.
+
+Existing write handles retain the process-local claim and kernel-backed cross-process `SessionWriteLease`. Header-only `stat` and `list` translate supported historical headers without opening the body or publishing a generation. Projection-cache records bind their fold to the Session header's format version so a cache row cannot bypass a cardinality-changing migration.
+
+## Problem-to-solution mapping
+
+| Whole-artifact problem | Implemented mechanism | Result |
+|---|---|---|
+| One asynchronous decode call per Zstandard frame | One reusable decoder; outer 500 ms scheduling cadence | Removes 317,540 async transitions |
+| Complete plaintext, string, and row arrays | Incremental JSONL parser and row decoder | Retains only one cross-chunk record fragment |
+| Complete event array between every edge | Context-connected stateful stages | No intermediate version event arrays |
+| Packed chunks expand before folding | `SessionFormatEventRun` plus `transformRun()` | 9.14 million source events need not materialize |
+| Whole-artifact snapshot and deep freeze at every edge | Stage-owned exclusive values and final validation | Removes repeated recursive copy/freeze |
+| One-shot migration functions hide state | Per-artifact stage classes from immutable declarations | State ownership and concurrency are explicit |
+| Bulk current encode builds whole strings and Buffers | Record encoder, 1 MiB input slices, 4 MiB write batches | Bounds allocation and main-thread slices |
+| Verification repeats on the main thread | At most two complete-generation Workers | Keeps verification CPU off the main thread |
+| Production and fixture migration use different APIs | Catalog `createRestore()` with explicit policies | One decoder/chain implementation |
 
 
 ## Verification
 ## Verification
 
 
-Release verification runs the committed Session-format corpus gate over every versioned persisted-or-projected `session*.jsonl` fixture under `snapshots/`, `packages/`, and `scripts/snapshots/python-sdk-single-exe/`. Fixture-only omitted envelopes and request-header tokens are materialized before the real static catalog; every fixture reaches the current v1 view through current restoration or historical migration. Released-v0 replay inputs remain suffixless, while fresh v1 writer outputs use `session.v1.jsonl` for a parent and `session.<ordinal>.v1.jsonl` for children. Record and refresh preserve every completed generation, including generations of a child role absent from a later run. Malformed historical fixtures are repaired at their source rather than admitted through path-dependent replay policy. The continuing gate discovers the corpus dynamically and fails every restoration refusal; separate assembled JSONL tests own exact physical-byte migration.
+### Benchmark input and meanings
+
+The benchmark uses Node v24.18.0 and one 116,228,655-byte v0 Zstandard log containing 317,540 frames and 454,151 physical rows. The old reader restores 9,143,111 expanded v0 events. Migration produces 72,784 current v2 events with artifact SHA-256 `fa16ff9472ca350595a3112c20a3db79655bc2673973469987ecaf2a57ebd17c`.
+
+Runs use built artifacts under plain Node, one process per sample, and a 16 GB V8 heap limit. “Retained heap” is measured after forced GC while the restored Session remains live. Values below are three-run medians except the whole-artifact failure, which consistently cannot reach a handle.
+
+### Physical Decode
+
+| Data path | Decode time | Peak RSS | Scheduling |
+|---|---:|---:|---|
+| Pre-migration optimized reader | 1.553s | 916MB | One decoder; 2–3 outer yields |
+| Whole-artifact migration | 7.527s | 7,219MB | 317,540 async decoder calls |
+| Streaming Stage path | 1.467s | 908MB | One decoder; 2 outer yields |
+
+### Historical-file cold open
+
+| Version | Time to restored Session | CPU time | Peak RSS | Retained heap | Restored events | Outcome |
+|---|---:|---:|---:|---:|---:|---|
+| Pre-migration high-performance v0 reader | 4.594s | 6.048s | 2.720GB | 2.016GB | 9,143,111 | Reads v0; does not migrate |
+| Whole-artifact migration | >72.8s | — | ≥7.219GB during Decode | — | — | OOM before a handle |
+| Streaming Stage migration with serial publication | 6.241s | 8.493s | 2.107GB | 477MB | 72,784 | Publishes and opens v2 |
+
+The old reader has lower one-time wall time because it performs no format conversion or durable publication. It also keeps the 9.14-million-event representation live. The Stage path pays encode and verification once, then retains the folded v2 state.
+
+### Current-format cold open
+
+| Version reading its current format | Time to restored Session | Peak RSS | Retained heap |
+|---|---:|---:|---:|
+| Old reader on v0 | 4.594s | 2.720GB | 2.016GB |
+| Whole-artifact-era reader on v2 | 1.273s | 1.107GB | 476MB |
+| Streaming Stage reader on v2 | 1.284s | 1.109GB | 476MB |
+
+The current-v2 fast path remains performance-equivalent. The architectural change does not route current data through historical stages.
+
+### Streaming serial migration breakdown
+
+This table records the serial open flow measured for this Stage decision. The current preparation-first scheduling and its measurements are owned by [Historical Session read preparation](2026-09-05-read-only-session-migration-preparation.md).
+
+| Phase | Median |
+|---|---:|
+| Source Decode and migration | 2.784s |
+| Encode, write, and sync | 0.956s |
+| Full staged-file Worker verification | 1.415s |
+| Source recheck and no-overwrite publication | 0.106s |
+| Committed-prefix verification and header reopen | 0.046s |
+| Generation ensure-current total | 5.318s |
+| Final current decode observed by persistence | 0.620s |
+| Session restoration | 0.594s |
+| End-to-end restored Session | 6.241s |
+
+The generation breakdown and end-to-end table come from separate instrumented runs, so rounded rows are not expected to sum exactly.
+
+Format, catalog, edge, JSONL, fixture, replay, and built-Worker tests cover both encodings, packed runs, header-only classification, torn tails, migration refusal, deterministic legacy normalization, source changes, target collisions, write leases, and Worker failure.
+
+## Consequences
+
+At least one final current-event array remains necessary because Session restoration and Agent execution retain complete history. The Stage architecture removes full source and intermediate target arrays; it does not promise memory proportional to a page window.
+
+Decoded scalar `assistant/chunk` rows receive envelope validation and final target validation, but their complete frozen-v1 source payload-member validation is deferred because that per-event check materially affects Decode and migration time on released logs. Packed Assistant runs remain strictly decoded. The scalar check must be restored only with performance evidence that preserves this migration path's measured behavior.
 
 
-Handle-integration verification runs the pure format, catalog, persistence-seam, and JSONL provider suites together: 420 tests cover both encodings, immutable publication races, header-only observation, read and write handles, migration refusal, append after migration, cancellation, and crash-tail behavior with per-file 100% statement, branch, function, and line coverage. Repository typecheck and lint, 113 keyless recorded-session replays with two declared skips, and 28 owner-local expected-output cases also pass on the merged master checkpoint.
+Read-only access consumes the Stage result before durable publication, while write open reuses the same result and waits for publication before append. The persistence scheduling remains independent from the format pipeline.
 
 
-The assembled headless profile test stages `session.jsonl`, resumes it through the shipped composition, observes v1 before Session construction, verifies that the exact v0 bytes and inode remain while `session.v1.jsonl` appears, and proves the next append targets v1. JSONL contract tests exercise raw and Zstandard exclusive publication, torn-tail preservation, source changes, target collisions, future-highest refusal, revision-keyed parsed-log reuse, listing rescans, temporary cleanup, committed reopen, and current-format bypass.
+Lower generations remain for operator inspection. Retention does not promise downgrade compatibility, automatic fallback, or that an older runtime can safely interpret a newer generation.
 
 
 ## Alternatives considered
 ## Alternatives considered
 
 
-- **Migrate only on continuation** — leaves query, export, fork, and suffix consumers on old generations and duplicates restoration policy.
-- **Return a migrated in-memory view without persisting** — lets one process observe state that does not match the highest committed generation and postpones failure until a later writer.
-- **Persist every intermediate version** — consumes space and creates recovery states with no runtime consumer; only the source and final generation are durable.
-- **Let mounted event-owner plugins register migrations** — makes historical readability deployment-dependent; the static catalog must work before feature plugins mount.
-- **Reuse one filename for every current format and relocate its predecessor** — rejected because migration would move or overwrite committed evidence, require collision and retention rules, and make the filename disagree with the stored format. Canonical immutable generation names let discovery select the highest version directly.
+- **Optimize only Zstandard Decode** — restores physical Decode speed but leaves source rows, expanded events, snapshots, intermediate artifacts, and bulk encode in memory.
+- **Synchronous Generator stages** — retain execution frames and batches at each yield. Real-log measurements increased migration time and migrate-complete RSS from about 1.0 GB to about 1.2 GB.
+- **Return arrays from each stage** — preserves the old allocation, traversal, and flattening costs under a new name.
+- **Give each stage an internal output queue** — adds drain, EOF, and error ownership while still retaining intermediate values.
+- **Inject an emit callback through constructors** — forces reverse construction or a partially connected lifecycle. Passing a context to operations keeps stage construction independent of downstream wiring.
+- **Share stateful codec instances globally** — would mix pending attempts, mappings, and counters across concurrent Session restores.
+- **Persist every intermediate format version** — creates durable states with no runtime consumer; only the exact source and final current generation are needed.
+- **Let mounted plugins register migrations** — makes historical readability deployment dependent. The static catalog must restore released formats before feature plugins mount.

+ 149 - 25
.agents/notes/implemented/architecture/2026-08-31-released-session-format-migrations.zh.md

@@ -1,4 +1,4 @@
-# Agent Note: 已发布 Session 格式在读取正文时通过相邻纯迁移边升级
+# Agent Note: 已发布 Session 格式通过有状态流式 Stage 迁移
 
 
 Status: implemented
 Status: implemented
 
 
@@ -6,52 +6,176 @@ Status: implemented
 
 
 ## 问题
 ## 问题
 
 
-Session 格式 v0 已随 alpha 版本发布,因此结构化 writer 变更不能再把已有 JSONL 当作可丢弃的预发布状态。已存储事件正文通过读或写 `SessionHandle` 到达恢复、查询、导出、分叉与继续路径。只迁移一个消费方会让调用方看到不同的逻辑 generation,或只在后续 writer 到达旧文件时失败。
+Session 格式 v0 已随 alpha 版本发布,因此结构化 writer 变更不能再把已有 JSONL 当作可丢弃的预发布状态。第一版 whole-artifact migration 让这些日志在语义上可迁移,但它的数据模型会让一份 116 MB 真实 Session 在返回 handle 前耗尽 16 GB Node 进程。
 
 
-迁移必须保留精确源路径、字节与 inode,包括撕裂的物理尾部,同时为每个已发布格式提供一个无歧义的规范文件名。普通 JSONL 与 Zstandard 是同一逻辑格式的编码选择,不能产生两套并行迁移实现。
+### Whole-artifact 性能问题
+
+- Zstandard 输入包含 317,540 个 frame,每个 frame 都单独执行一次异步解压。实现先保留全部 plaintext frame,再在 JSON 解析前统一拼接,因此创建了同等数量的 Promise、线程池与 native Decode 调度。
+- Physical Decode 会在同一请求的重叠阶段物化完整 plaintext Buffer、完整字符串、全部 JSONL row、展开后的 source events、迁移后的 target events、编码后的 target rows、拼接后的目标字符串与目标 physical Buffer。
+- 每个 codec 与 migration edge 都会调用 `snapshotSessionFormatJson()` 或 `snapshotSessionFormatArtifact()`,在相邻迁移前后递归复制并 deep freeze 完整 header、row、payload 与 event array。
+- 已发布的 packed Assistant chunk 会先展开成约 914 万个 v0/v1 逻辑事件,再由 v1-to-v2 折叠成 72,784 个 current events。Whole-artifact API 要求两种表示和 old-to-new seq map 同时存活。
+- Encode 会在内存中构造完整 JSONL 与压缩输出。成功路径随后 Decode staged target、Decode committed target,并由 persistence 再 Decode 一次以创建业务对象;它还会完整重读 source 以比较 fingerprint。
+- 逐 frame `await` 没有形成有意义的有界调度。迁移前的高性能 reader 会复用一个同步 decoder,只由外层循环约每 500 ms yield 一次,从而避免数十万次异步切换。
+
+### 既有接口使单点优化无法组合
+
+`SessionFormatCodec` 以完整数组 Decode 与 Encode;每条相邻 `SessionFormatMigration` 接收并返回完整 `SessionFormatArtifact`;compiled chain 只能把已经物化的 artifact 交给下一条 edge。因此即使 physical decoder 单点变快,下游仍会重新创建 source-row array、expanded-event array、逐 edge snapshot 与 target-row array。
+
+Migration 实际有状态,但 API 把它们表现为一次性函数。v0-to-v1 需要跟踪 message 与 retry identity;v1-to-v2 需要暂存一个尚未结算的 Assistant attempt、被它阻塞的后续事件,并维护 old-to-new seq 引用。把这些状态隐藏在 closure 或 push/finish helper object 中,会让运行结构与静态声明分离,也让 production、Worker verify、fixture 与 replay 使用不同入口。
 
 
 ## 决策
 ## 决策
 
 
-`SESSION_FORMAT_VERSION` 是单调递增的当前 writer 整数。每个相邻 `vN -> vN+1` 转换由一个与 profile 无关的纯包负责。`@deepseek-ai/dsh-session-format` 只提供无损快照、唯一且无缺口的规划、仅 header 转换与整产物组合;`@deepseek-ai/dsh-session-format-catalog` 静态导入完整链,不依赖已挂载的 Cordis 插件。历史 codec 和归一化器位于具名迁移边包中,而当前 Session 与持久化代码只接纳最新逻辑类型。
+Session format 包采用有状态同步 Stage API。静态 migration declaration 描述一条相邻版本边,并为每次 artifact restore 创建新的 stage。Stage 拥有该 artifact 的可变状态;不同 Session 之间绝不共享 stage instance。
 
 
-每条迁移边都会冻结严格的源与目标语义,其目标物理 codec 则保持词汇中立,使普通事件增长可以留在同一格式版本内。目录通过已安装的 peer `@deepseek-ai/dsh-session` 及其当前 `KNOWN_SESSION_EVENT_TYPES` 还原最终代,避免冻结的历史迁移边反过来成为当前词汇 owner。
+### Stage 与 Context 协议
 
 
-JSONL provider 在 `open` 为已存储 Session 返回句柄前完成 ensure-current 工作。它选择最高规范 generation、迁移受支持的历史正文,并从同一物理快照解码当前结果;公开 `SessionPersistence` 与 `SessionHandle` 接口不包含迁移操作。仅 header 的 `stat` 与 `list` 会重新扫描 Session 目录,在内存中转换受支持的历史 header,且绝不发布后继。`create` 独立于 header 可读性检查规范文件名,因此每个现有 generation 都会占用其 Session id。
+```text
+interface SessionFormatMigrationContext {
+  emitEvent(event: SessionFormatEvent): void
+  emitRun(run: SessionFormatEventRun): void
+}
 
 
-取消属于提供信号的 `open`、`stat` 或 `list` 调用。发现、稳定读取、解码与发布前检查都会观察该信号;不可变后继一旦发布且其目录项已经同步,后续取消不会删除已提交 generation。
+interface SessionFormatMigrationStage {
+  readonly headerInheritedEventCount?: number
+  transformEvent(
+    event: SessionFormatEvent,
+    context: SessionFormatMigrationContext,
+  ): void
+  transformRun(
+    run: SessionFormatEventRun,
+    context: SessionFormatMigrationContext,
+  ): void
+  finish(context: SessionFormatMigrationContext): number
+}
+```
 
 
-配置的 JSONL 编码拥有一个完整后缀:`.jsonl` 或 `.jsonl.zstd`。迁移读取稳定的精确源,解码可恢复逻辑前缀,在内存中组合全部必需迁移边,只为最终目标校验并同步同目录临时 stage,重新检查源 fingerprint,以不覆盖方式发布此前不存在的目标,同步 namespace,并在返回句柄前通过当前格式校验重新打开。源永不移动或改变;只有可丢弃临时 stage 可以被移动、链接或移除。迁移不会合成中断轮次事件:agent-loop 通过写句柄追加这些修复,而只读查询路径在内存中补齐它们。
+`SessionFormatMigrationContext.emitEvent()` 与 `emitRun()` 都是同步操作。Producer 会声明其发出单个事件还是紧凑 run,因此热路径不会根据已解析文件对象的属性推断类别。调度归 caller 所有,context 在每次调用时传入,而不是把 callback 注入 stage constructor。一个输入可以输出零个、一个或多个值,不需要分配临时返回数组,也不需要 stage 内部保留输出队列。
 
 
-规范文件名编码物理格式 generation:v0 是 `session.jsonl` 或 `session.jsonl.zstd`;每个正 generation 都是小写 `session.vN.jsonl` 或 `session.vN.jsonl.zstd`。`dsh-session-format` 拥有原始 basename 规则(`sessionFormatLogFilename`、`parseSessionFormatLogFilename`);JSONL provider、session-log 导出归档与 recorded-session fixture 只追加压缩后缀。发布绝不重命名、替换或删除已提交 generation 路径。目标已经存在时,只有它是普通当前格式文件且字节与预期完全相同时才接受;其他目标都会拒绝。低 generation 为 operator 检查或显式复制而保留,但普通 runtime 操作选择数值最高的规范名称,绝不把保留的前任当作自动 fallback、restore 或 downgrade 支持。
+`SessionFormatMigration` 继续作为 immutable declaration,声明版本号、header migration、target-header validation 与 `createStage()`。`CompiledSessionFormatChain` 只校验一次唯一、无缺口的 edge 序列,按 source-to-target 顺序创建每次 artifact 独占的 stage,再按反方向用 context 连接它们。`finish()` 按 source-to-target 顺序关闭 stage,使每一级都能在下游关闭前发出尾部数据。
 
 
-当前格式快速路径从一个稳定源快照分类 header,不调用历史 converter,不写 generation,并把该快照交给当前格式解码,而不再次读取文件。解码日志进入现有按 revision 为键的有界 memo,供紧接的观察到恢复交接复用,而 `stat` 与 `list` 会有意重新扫描。多条迁移边保持原 generation 不变,并只发布最终目标;中间版本只存在于内存。源 fingerprint 重新检查会在内容变化时重启迁移,排他目标发布只在竞争胜者字节完全相同时接受它。跨进程 append 隔离不在此保证内。
+```text
+JSONL record
+  → released physical row decoder
+  → v0-to-v1 stage
+  → v1-to-v2 stage
+  → current event collector
+```
 
 
-第一条迁移边 `@deepseek-ai/dsh-session-format-v0-to-v1` 有意保持恒等形态:除版本和 v0 已接纳的有限历史归一化外,它保留逻辑 header、事件、序号、引用、时间戳、payload 与已配置的压缩选择。精确的 `session.jsonl[.zstd]` 源保持字节与 inode 相同,当前 writer 则编码新的 `session.v1.jsonl[.zstd]` 后继。这样可在出现改变基数的格式前先验证完整发布生命周期。
+Chain 中不存在 `flatMap`、spread expansion、中间 event array 或 scheduler。只有在每个 migration stage 都已获得直接消费 compact run 的机会后,最终 event collector 才会展开它。
 
 
-投影缓存记录把自己的折叠结果绑定到 Session header 的 `formatVersion`。`session_projcache` v7 reader 可以在结构上载入前代 domain 记录,但缺少格式代的记录不能播种当前 Session;权威日志会重新折叠它,下一次检查点写入完整的当前 identity。这样,任何在有界规范化或基数变化边之前产生的缓存行都不能绕过该迁移。
+### Physical codec 与 packed run
 
 
-## 后果
+每个 released codec 会用显式 `strict` 或 `recoverable` 策略创建 row decoder。Decoder 每次通过不同的 context 方法校验并 emit 一个 event 或 codec-owned `SessionFormatEventRun`。v0-to-v1 与 v1-to-v2 都实现 `transformEvent()` 和 `transformRun()`,因此 packed Assistant chunk 可以直接到达 folding edge,无需先变成数百万个普通事件。
+
+v0-to-v1 除了有限的 released-v0 归一化外,会保留逻辑 header、seq、引用、时间戳与 payload。它转换已移除的 `steering/message` 与 `compact/*` 事件名称,接受出现在对应 `step/end` 之后的已发布 `llm/retry`,按 turn/step/provider/policy chain 为缺失的 `llm/retry.retryId` 确定性补值,并为省略 id 的旧 compaction group 确定性补充同一个 `compactionId`。v1-to-v2 负责 attempt folding 与引用重写,并且只 emit 已结算的 current event。它会把旧的 goal 来源 user message 拆成 `goal/change` 与原本的模型可见 message。它还会为一种有限的已发布 restart 插入 interrupted `turn/end`:一个没有 open step 的 open turn 后出现非空 `next-turn` inbox splice,随后直接开始编号连续的下一轮。
+
+Catalog 为 production、Worker、fixture 与 replay 暴露同一个 `createRestore()`。Recovery policy 与最终 validation policy 在 restore 创建时一次确定。Historical production 使用 recoverable source parsing 与 transformed-current validation;这种策略会在迁移后校验已发布 current 结果,而已经是 current 的输入只接受 codec 校验。Worker 与 fixture verification 使用 strict parsing 与已安装 current 格式的完整 restoration。Migration stage 或 transformed-current validation 的拒绝会保持为 `SessionFormatUnsupportedMigrationError`;物理解码失败仍是 corruption。Test support 只保留 fixture 自身需要的 token 和 envelope materialization。
+
+### JSONL 串联
+
+JSONL provider 只扫描一次 frame boundary,复用一个 Zstandard decoder,增量解析完整 JSONL record,并把 row 直接送入 catalog restore。外层循环按有界 cadence yield;不存在逐 frame `await`、完整 plaintext 或 source-row array。
 
 
-较新 build 读取事件正文时可能持久增加一个更高 generation。精确旧 generation 仍然可用,但 runtime 此后选择最高规范文件名;保留不承诺旧 build 能安全 downgrade,也不保证新 build 在后继损坏时 fallback。只读文件系统会报告可操作的迁移失败,而不会返回与磁盘不一致的内存当前视图。
+Current encode 以单条 record 为单位。Provider 在主线程每个 slice 序列化约 1 MiB plaintext,通过一个会传播 source error 的 Zstandard context 流式压缩,以 4 MiB batch 写入同目录排他创建的临时文件,并在 publication 前 sync。进程级 scheduler 最多允许两个完整 verification Worker 并行,并把释放的 permit 直接交给最早的 waiter。
 
 
-JSONL 发布在 POSIX 上使用硬链接创建与目录同步,在 Windows 上使用 write-through 且不覆盖的 `MoveFileExW`。竞争 writer 已先创建目标时,只有已提交字节完全匹配才接受。每个 Session 只支持一个进程内 writer。未来逐 Session 跨进程锁可以关闭剩余的源检查到发布竞态,而无需改变格式迁移边接口。
+Preparation 会把 cancellation 传给 source read,并在现有的约 500 ms Decode yield 边界观察它。`publish()` 一旦开始,encode、Worker verification 与 publication 不接收 caller cancellation,并运行到终态;write open 会在之后再次检查 caller signal。已经发布的 generation 绝不会回滚。
 
 
-保留的 generation 不是实时流 WAL。未来可选 WAL sidecar 可以在硬崩溃间保留未完成 assistant 流。显式 generation 检查或复制、保留策略工具、压缩转换与流式整产物转换都是独立功能;自动 fallback 与 downgrade compatibility 并非隐含 future work。
+Stage pipeline 终止于一份 prepared current artifact。[历史 Session 只读迁移准备](2026-09-05-read-only-session-migration-preparation.zh.md)定义 read open 如何立即消费该 artifact,以及 write open 如何在返回 append 权限前完成 encode、verification 与 publication。
 
 
-本记录取代 [Session 日志版本机制](2026-08-10-session-log-version-mechanism.zh.md) 中仅在继续时持久化和迁移链仍推迟的规则。原记录继续负责何时递增版本,以及普通同版本 `ignorable` 事件行为。
+### Durable format 与 publication 规则
+
+规范文件名编码 physical format generation:v0 使用 `session.jsonl[.zstd]`,正 generation 使用 `session.vN.jsonl[.zstd]`。Migration 不会移动、覆盖或删除任何 committed generation,并且只写最终 current target;中间版本只存在于 stage state。
+
+POSIX publication 使用 hard-link creation 加目录 sync;Windows 使用 no-overwrite、write-through 的 `MoveFileExW`。已有 target 只有在其已校验 migration prefix 等于 staged bytes 时才会被接受;任何 append tail 都属于 current-generation reader,而不是 migration winner verification。
+
+既有 write handle 继续使用进程内 claim 与内核支持的跨进程 `SessionWriteLease`。仅 header 的 `stat` 与 `list` 可以转换受支持的历史 header,但不打开 body,也不发布 generation。Projection-cache record 会把 fold 绑定到 Session header 的 format version,使 cache row 不能绕过改变 event 基数的 migration。
+
+## 问题与方案对照
+
+| Whole-artifact 问题 | 实现机制 | 结果 |
+|---|---|---|
+| 每个 Zstandard frame 单独异步 Decode | 一个可复用 decoder;外层 500 ms 调度 cadence | 删除 317,540 次异步切换 |
+| 完整 plaintext、string 与 row array | 增量 JSONL parser 与 row decoder | 只保留一条跨 chunk 残行 |
+| 每条 edge 之间都形成完整 event array | Context 直连的有状态 stage | 不保留中间版本 event array |
+| Packed chunk 在 folding 前完整展开 | `SessionFormatEventRun` 与 `transformRun()` | 无需物化 914 万 source events |
+| 每条 edge 都 whole-artifact snapshot/deep freeze | Stage-owned 独占值与最终 validation | 删除重复递归复制与冻结 |
+| One-shot migration function 隐藏状态 | Immutable declaration 创建每次 artifact 独占的 stage class | 状态 ownership 与并发关系显式化 |
+| Bulk current encode 构造完整 string/Buffer | 单条 record encoder、1 MiB input slice、4 MiB write batch | 限制分配与主线程 slice |
+| 主线程重复执行完整 verification | 最多两个 complete-generation Worker | Verification CPU 不占用主线程 |
+| Production 与 fixture 使用不同 migration API | Catalog `createRestore()` 加显式 policy | 只保留一套 decoder/chain 实现 |
 
 
 ## 验证
 ## 验证
 
 
-发布验证针对 `snapshots/`、`packages/` 与 `scripts/snapshots/python-sdk-single-exe/` 下每个带版本、来自持久化或投影的 `session*.jsonl` fixture 运行已提交 Session 格式语料门禁。fixture 专用的缺失信封与 request-header token 会先被实体化,再进入真实静态 catalog;每个 fixture 都会通过当前格式 restore 或历史迁移得到当前 v1 视图。Released-v0 replay 输入保持无后缀,而新鲜 v1 writer 输出对 parent 使用 `session.v1.jsonl`、对 child 使用 `session.<ordinal>.v1.jsonl`。Record 与 refresh 会保留每个已完成 generation,包括后续运行不再产生的 child role generation。Malformed 历史 fixture 在来源处修复,不通过依赖路径的 replay 策略准入。持续运行的门禁会动态发现语料,并拒绝每个 restore failure;独立组装式 JSONL 测试负责精确物理字节迁移。
+### Benchmark 输入与口径
+
+Benchmark 使用 Node v24.18.0 和一份 116,228,655-byte 的 v0 Zstandard 日志,其中包含 317,540 个 frame 与 454,151 个 physical row。老 reader 会恢复 9,143,111 个展开后的 v0 event;migration 会生成 72,784 个 current v2 event,artifact SHA-256 为 `fa16ff9472ca350595a3112c20a3db79655bc2673973469987ecaf2a57ebd17c`。
+
+所有样本均通过 plain Node 运行 build artifact,每个样本使用独立进程,V8 heap limit 为 16 GB。“Retained heap”表示 restored Session 仍存活时强制 GC 后的 heap。除无法得到 handle 的 whole-artifact 失败外,下表使用三次运行中位数。
+
+### Physical Decode
+
+| 数据路径 | Decode 耗时 | 峰值 RSS | 调度 |
+|---|---:|---:|---|
+| Migration 前的高性能 reader | 1.553s | 916MB | 一个 decoder;外层 yield 2–3 次 |
+| Whole-artifact migration | 7.527s | 7,219MB | 317,540 次 async decoder 调用 |
+| Streaming Stage 路径 | 1.467s | 908MB | 一个 decoder;外层 yield 2 次 |
+
+### 历史文件首次冷打开
+
+| 版本 | Session restore 完成 | CPU 时间 | 峰值 RSS | Retained heap | Restore event 数 | 结果 |
+|---|---:|---:|---:|---:|---:|---|
+| Migration 前的高性能 v0 reader | 4.594s | 6.048s | 2.720GB | 2.016GB | 9,143,111 | 读取 v0,不迁移 |
+| Whole-artifact migration | >72.8s | — | Decode 阶段已 ≥7.219GB | — | — | 返回 handle 前 OOM |
+| Streaming Stage migration + 串行 publication | 6.241s | 8.493s | 2.107GB | 477MB | 72,784 | 发布并打开 v2 |
+
+老 reader 的一次性 wall time 更低,因为它不做格式转换和 durable publication;同时它会常驻 914 万 event 的表示。Stage 路径只多支付一次 encode 与 verification,随后保留折叠后的 v2 state。
+
+### Current-format 冷打开
+
+| 版本读取自己的 current format | Session restore 完成 | 峰值 RSS | Retained heap |
+|---|---:|---:|---:|
+| 老 reader 读取 v0 | 4.594s | 2.720GB | 2.016GB |
+| Whole-artifact 时代 reader 读取 v2 | 1.273s | 1.107GB | 476MB |
+| Streaming Stage reader 读取 v2 | 1.284s | 1.109GB | 476MB |
+
+Current-v2 快路径保持性能等价。架构改造不会让 current data 进入 historical stage。
+
+### Streaming 串行 migration 分段
+
+下表记录该 Stage 决策测量的串行 open 流程。当前 preparation-first 调度及其测量由[历史 Session 只读迁移准备](2026-09-05-read-only-session-migration-preparation.zh.md)记录。
+
+| 阶段 | 中位耗时 |
+|---|---:|
+| Source Decode + migration | 2.784s |
+| Encode + write + sync | 0.956s |
+| staged 文件完整 Worker verify | 1.415s |
+| Source recheck + no-overwrite publication | 0.106s |
+| committed-prefix verify + header reopen | 0.046s |
+| Generation ensure-current 总计 | 5.318s |
+| Persistence 观察到的最终 current Decode | 0.620s |
+| Session restore | 0.594s |
+| 端到端 Session restore 完成 | 6.241s |
+
+Generation 分段与端到端数据来自不同 instrumented run,因此四舍五入后的各行不要求精确相加。
+
+Format、catalog、edge、JSONL、fixture、replay 与 built-Worker 测试覆盖两种编码、packed run、仅 header 分类、torn tail、migration refusal、确定性 legacy normalization、source change、target collision、write lease 与 Worker failure。
+
+## 后果
+
+最终 current-event array 仍然不可消除,因为 Session restore 与 Agent 执行需要完整历史。Stage 架构删除完整 source 与中间 target array,但不承诺内存与分页窗口大小成正比。
+
+解码后的单条 `assistant/chunk` 会接受 envelope 校验与最终 target 校验,但其完整冻结 v1 source payload 成员校验仍处于延期状态,因为这项逐事件检查会显著影响已发布日志的 Decode 与 migration 耗时。Packed Assistant run 仍接受严格解码。只有性能证据表明不会破坏该迁移路径的已测表现时,才能恢复单条 chunk 校验。
 
 
-句柄集成验证会一起运行纯格式、catalog、持久化 seam 与 JSONL provider 测试套件:420 个测试覆盖两种编码、不可变发布竞态、仅 header 观察、读写句柄、迁移拒绝、迁移后 append、取消与崩溃尾部行为,并达到逐文件 100% statement、branch、function 与 line coverage。仓库 typecheck 与 lint、含两个已声明 skip 的 113 个无密钥 recorded-session replay,以及 28 个 owner-local expected-output case 也都在合并 master 的 checkpoint 上通过。
+Read-only access 会在 durable publication 前消费 Stage 结果;write open 则复用同一结果,并在 append 前等待 publication。Persistence 调度仍与 format pipeline 相互独立。
 
 
-组装后的 headless profile 测试会暂存 `session.jsonl`,通过随附组合恢复它,在构造 Session 前观察到 v1,验证精确 v0 字节与 inode 保持不变而 `session.v1.jsonl` 出现,并证明下一次 append 以 v1 为目标。JSONL 约定测试覆盖 raw 与 Zstandard 排他发布、撕裂尾部保留、源变化、目标冲突、最高未来版本拒绝、按 revision 复用已解析日志、列表重新扫描、临时文件清理、已提交重开与当前格式直通。
+低 generation 为 operator 检查而保留。Retention 不承诺 downgrade compatibility、automatic fallback,也不保证旧 runtime 能安全理解新 generation。
 
 
 ## 考虑过的替代方案
 ## 考虑过的替代方案
 
 
-- **只在继续时迁移**——让查询、导出、分叉与后缀消费者停留在旧代际,并重复恢复策略。
-- **返回迁移后的内存视图但不持久化**——让进程观察到与最高已提交 generation 不一致的状态,并把失败推迟到后续 writer。
-- **持久化每个中间版本**——消耗空间并产生没有 runtime 消费者的恢复状态;只有源与最终代际应持久。
-- **让已挂载事件 owner 插件注册迁移**——使历史可读性依赖部署;静态 catalog 必须在功能插件挂载前工作。
-- **让每个当前格式复用同一个文件名并迁走前任**——不予采用,因为迁移会移动或覆盖已提交证据,需要冲突与保留规则,并让文件名与存储格式不一致。规范不可变 generation 名让发现流程直接选择最高版本。
+- **只优化 Zstandard Decode**——可以恢复 physical Decode 速度,但 source rows、expanded events、snapshot、intermediate artifact 与 bulk encode 仍会留在内存中。
+- **同步 Generator stage**——每个 yield 都会保留执行帧与 batch。真实日志测量使 migration 更慢,并让 migrate-complete RSS 从约 1.0 GB 增长到约 1.2 GB。
+- **每个 stage 返回数组**——只是给旧的 allocation、遍历与 flattening cost 换了名字。
+- **Stage 内部输出队列**——增加 drain、EOF 与 error ownership,同时仍会保留中间值。
+- **通过 constructor 注入 emit callback**——迫使 chain 反向构建或引入 partially connected lifecycle。操作时传 context 可以让 stage construction 不依赖下游 wiring。
+- **全局复用有状态 codec instance**——会让不同 Session 的 pending attempt、mapping 与 counter 相互污染。
+- **持久化每个中间格式版本**——产生没有 runtime consumer 的 durable state;只需要精确 source 与最终 current generation。
+- **让 mounted plugin 注册 migration**——使历史可读性依赖部署。Static catalog 必须在 feature plugin 挂载前恢复已发布格式。

+ 2 - 2
.agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.md
-2026-09-01-v2-embedded-assistant-streams.md: c9d7a66481d4258de8f9f7abdaa12d25f5217510
-2026-09-01-v2-embedded-assistant-streams.zh.md: ff4caceb68c34e470d2dd81fc9cf214555ad676e
+2026-09-01-v2-embedded-assistant-streams.md: bee4d50fb830caa277bb700f3415e6d7f98ff64b
+2026-09-01-v2-embedded-assistant-streams.zh.md: 9116e94111b78af68d33f6f01dc289ee9f349b7e

+ 9 - 5
.agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.md

@@ -21,17 +21,19 @@ Session format v2 has no top-level `assistant/chunk` event. Each model attempt c
 
 
 `AssistantStreamAccumulator` snapshots each chunk once. Consecutive text, reasoning, or tool-argument deltas for the same block become one compact run with its first timestamp, exact timestamp gaps, and one array member per original delta. Every other chunk remains a timestamped raw record. `expandAssistantStream()` strictly validates and reconstructs the exact timed sequence; compaction never joins delta boundaries.
 `AssistantStreamAccumulator` snapshots each chunk once. Consecutive text, reasoning, or tool-argument deltas for the same block become one compact run with its first timestamp, exact timestamp gaps, and one array member per original delta. Every other chunk remains a timestamped raw record. `expandAssistantStream()` strictly validates and reconstructs the exact timed sequence; compaction never joins delta boundaries.
 
 
-The current v2 validator requires the embedded stream to reproduce a non-empty `assistant/message`'s content, usage, and replay state. An empty stream remains valid for a migrated legacy message that had no source chunks. `assistant/message` cannot carry obsolete chunk `sourceEventSeqs`; ordinary user and tool surface provenance remains available.
+The migration publication verifier and frozen v2 fixture validator require the embedded stream to reproduce a non-empty `assistant/message`'s content, usage, and replay state. An empty stream remains valid for a migrated legacy message that had no source chunks. Ordinary Session restoration validates the settlement fields needed by the runtime without expanding every historical stream; consumers that expand a compact stream validate its records when they read it. `assistant/message` cannot carry obsolete chunk `sourceEventSeqs`; ordinary user and tool surface provenance remains available.
 
 
 ### Live presentation and durable replay
 ### Live presentation and durable replay
 
 
 `agent/assistant-stream` publishes process-local start, transient chunk, and end frames. The loop appends the complete `assistant/message` or `assistant/attempt` before a committed end frame names its type and sequence. An abandoned end has no settlement.
 `agent/assistant-stream` publishes process-local start, transient chunk, and end frames. The loop appends the complete `assistant/message` or `assistant/attempt` before a committed end frame names its type and sequence. An abandoned end has no settlement.
 
 
-The Web follow adapter opts into these process-local frames and adds the last durable sequence observed at each start. It presents chunks as Client-only `assistant/live-chunk` updates between durable cursors, stages only a later matching settlement until the committed end, and reopens follow on a revision gap. A committed end publishes a named settlement delta that removes the attempt's transient matches, adds the durable entry, and replays only affected Conversation Contexts; an abandoned end publishes the same delta without an entry. A reconnect baseline carries the active attempt's durable start cursor and compact prefix. Paged history, replay, telemetry, token accounting, and cold UI assembly read the durable embedded stream rather than the live frames.
+The Web follow adapter opts into these process-local frames and adds the last durable sequence observed at each start. It presents chunks as Client-only `assistant/live-chunk` updates between durable cursors, stages only a later matching settlement until the committed end, and reopens follow on a revision gap. A committed end publishes a named settlement delta that removes the attempt's transient matches, adds the durable entry, and replays only affected Conversation Contexts; an abandoned end publishes the same delta without an entry. A reconnect baseline carries the active attempt's durable start cursor and compact prefix.
+
+The Client event source passes durable settlements through unchanged. The Chat and Trajectory Assistant nodes fold `assistant/live-chunk` while an attempt is active, build settled output directly from `assistant/message`, and do not replay an `assistant/attempt` stream for presentation. Cold settled presentation therefore does not reconstruct per-token timing; other consumers may expand the durable stream when they require its exact evidence.
 
 
 ### Released v1 to v2 migration
 ### Released v1 to v2 migration
 
 
-The adjacent migration validates the complete frozen v1 artifact, groups chunks by turn, step, terminal boundary, and exact message provenance, and then substitutes one settlement per attempt. A successful group's chunks move into its message. An unclaimed group becomes `assistant/attempt` at the last consumed chunk's position. Unrelated interleaved events retain their relative order, and survivors receive dense v2 sequence numbers. The edge compacts, expands, and re-assembles embedded streams through the runtime `AssistantStreamAccumulator`, `expandAssistantStream`, and `BlockAssembler` from `dsh-llm` instead of frozen copies, because that package owns the v2 stream encoding. Target validation re-checks agreement between each migrated `assistant/message` and its embedded stream itself, so a disagreeing v1 log is refused as an unsupported migration with its source artifact retained instead of surfacing as corruption from the installed Session restoration. A later format that changes the stream encoding must freeze copies of these helpers into this edge.
+The adjacent migration validates the complete frozen v1 artifact, groups chunks by turn, step, terminal boundary, and exact message provenance, and then substitutes one settlement per attempt. A successful group's chunks move into its message. An unclaimed group becomes `assistant/attempt` at the last consumed chunk's position. Unrelated interleaved events retain their relative order, and survivors receive dense v2 sequence numbers. The edge compacts embedded streams through the runtime `AssistantStreamAccumulator` from `dsh-llm` instead of a frozen copy, because that package owns the v2 stream encoding. The isolated publication verifier expands and re-assembles the written stream through `expandAssistantStream()` and `BlockAssembler`, then checks each migrated `assistant/message` against it before publication. A later format that changes the stream encoding must freeze copies of these helpers into this edge.
 
 
 The edge remaps the finite declared reference inventory: envelope provenance, surface replacement endpoints, command source events, compaction ranges and shadowed lists, and title message lists. The model-visible text of a validated `session/title-llm-request` remains byte-identical in the source sequence namespace while its `messageSeqs` field moves to the v2 namespace; target validation therefore does not reconstruct that text from remapped sequences. A reference to a consumed chunk refuses migration; it is never redirected to a settlement with different meaning. The edge also refuses an inherited cut that splits an attempt.
 The edge remaps the finite declared reference inventory: envelope provenance, surface replacement endpoints, command source events, compaction ranges and shadowed lists, and title message lists. The model-visible text of a validated `session/title-llm-request` remains byte-identical in the source sequence namespace while its `messageSeqs` field moves to the v2 namespace; target validation therefore does not reconstruct that text from remapped sequences. A reference to a consumed chunk refuses migration; it is never redirected to a settlement with different meaning. The edge also refuses an inherited cut that splits an attempt.
 
 
@@ -49,7 +51,7 @@ The compact-stream tests pin exact accumulation and expansion for text, reasonin
 
 
 The pre-merge performance acceptance measured static catalog-routing overhead against direct released-v2 restoration of the same already parsed physical rows across three runs, 100 warmup pairs, and 600 measured pairs; it did not compare v1 with v2 or time backend I/O. Every pooled median and p95 regression stayed within the 5% budget, with a worst p95 regression of 3.150%.
 The pre-merge performance acceptance measured static catalog-routing overhead against direct released-v2 restoration of the same already parsed physical rows across three runs, 100 warmup pairs, and 600 measured pairs; it did not compare v1 with v2 or time backend I/O. Every pooled median and p95 regression stayed within the 5% budget, with a worst p95 regression of 3.150%.
 
 
-Agent-loop tests pin durable-before-end ordering, interrupted visible prefixes, failed and retry attempts, abandonment, usage, and replay metadata. Session Controller and Conversation tests pin live transient display, reconnect baselines, committed settlement release, history replay, Chat and Trajectory parity, while TypeScript and Python SDK snapshots pin the external event representation.
+Agent-loop tests pin durable-before-end ordering, interrupted visible prefixes, failed and retry attempts, abandonment, usage, and replay metadata. Session Controller and Conversation tests pin live transient display, reconnect baselines, committed settlement release, and history replay. Chat and Trajectory tests pin live partial presentation and direct final-message projection, while TypeScript and Python SDK snapshots pin the external event representation.
 
 
 ## Alternatives considered
 ## Alternatives considered
 
 
@@ -59,13 +61,15 @@ Agent-loop tests pin durable-before-end ordering, interrupted visible prefixes,
 
 
 **Carry packed chunk rows through the history API.** This reduces wire and Client work for v1 but gives the Client a second event vocabulary and keeps transport coupled to token-row cardinality. The current API carries scalar durable settlements plus a separate live transient stream.
 **Carry packed chunk rows through the history API.** This reduces wire and Client work for v1 but gives the Client a second event vocabulary and keeps transport coupled to token-row cardinality. The current API carries scalar durable settlements plus a separate live transient stream.
 
 
+**Strip embedded streams in Session Controller.** This reduces retained Client memory but creates a second durable event type and makes a transport-facing owner decide which evidence presentation consumers need. The measured bottleneck is repeated expansion, so each UI consumer decides whether to inspect the unchanged settlement.
+
 **Store the stream in a sidecar or replay-only fixture.** This splits one attempt's message and evidence across durability owners and cannot give ordinary resumed sessions the same failed-output and timing facts. The settlement is the atomic owner.
 **Store the stream in a sidecar or replay-only fixture.** This splits one attempt's message and evidence across durability owners and cannot give ordinary resumed sessions the same failed-output and timing facts. The settlement is the atomic owner.
 
 
 **Redirect references from consumed chunks to their settlement.** A chunk and an attempt settlement are not interchangeable facts. Refusal prevents a migration from silently changing the meaning of plugin-owned references.
 **Redirect references from consumed chunks to their settlement.** A chunk and an attempt settlement are not interchangeable facts. Refusal prevents a migration from silently changing the meaning of plugin-owned references.
 
 
 ## Consequences
 ## Consequences
 
 
-Current logs, telemetry, history pages, and cold Client assembly scale by model attempts rather than token chunks while retaining exact stream evidence inside each settlement. Live presentation remains incremental and intentionally process-local.
+Current logs, telemetry, and history pages scale by model attempts rather than token chunks while retaining exact stream evidence inside each settlement. The Client event window retains that compact evidence, but the Chat and Trajectory Assistant nodes do not expand settled streams into per-delta objects. Live presentation remains incremental and intentionally process-local.
 
 
 Unlike v1 top-level chunks, which the buffered persistence writer could flush before an attempt ended, v2 has no durable attempt evidence until settlement. A hard process or host loss before settlement discards the complete in-flight stream; `agent/assistant-stream` is not a write-ahead log. This tradeoff avoids a second durability owner for live output.
 Unlike v1 top-level chunks, which the buffered persistence writer could flush before an attempt ended, v2 has no durable attempt evidence until settlement. A hard process or host loss before settlement discards the complete in-flight stream; `agent/assistant-stream` is not a write-ahead log. This tradeoff avoids a second durability owner for live output.
 
 

+ 9 - 5
.agents/notes/implemented/architecture/2026-09-01-v2-embedded-assistant-streams.zh.md

@@ -21,17 +21,19 @@ Session format v2 没有顶层 `assistant/chunk` 事件。每个模型 attempt 
 
 
 `AssistantStreamAccumulator` 对每个 chunk 只快照一次。同一 block 的连续 text、reasoning 或 tool argument delta 会变成一个紧凑 run,包含首个时间戳、精确时间戳间隔和每个原始 delta 对应的一个数组成员。其他 chunk 保留为带时间戳的 raw record。`expandAssistantStream()` 会严格校验并重建精确的带时间序列;压缩绝不会合并 delta 边界。
 `AssistantStreamAccumulator` 对每个 chunk 只快照一次。同一 block 的连续 text、reasoning 或 tool argument delta 会变成一个紧凑 run,包含首个时间戳、精确时间戳间隔和每个原始 delta 对应的一个数组成员。其他 chunk 保留为带时间戳的 raw record。`expandAssistantStream()` 会严格校验并重建精确的带时间序列;压缩绝不会合并 delta 边界。
 
 
-当前 v2 校验器要求嵌入式 stream 能复现非空 `assistant/message` 的 content、usage 与 replay state。对于没有源 chunk 的已迁移旧 message,空 stream 仍然有效。`assistant/message` 不能携带已停用的 chunk `sourceEventSeqs`;普通 user 与 tool surface provenance 保持可用。
+Migration publication verifier 与冻结的 v2 fixture validator 要求嵌入式 stream 能复现非空 `assistant/message` 的 content、usage 与 replay state。对于没有源 chunk 的已迁移旧 message,空 stream 仍然有效。普通 Session restore 只校验 runtime 直接依赖的 settlement 字段,不展开全部历史 stream;需要展开 compact stream 的 consumer 会在读取时校验 record。`assistant/message` 不能携带已停用的 chunk `sourceEventSeqs`;普通 user 与 tool surface provenance 保持可用。
 
 
 ### 实时呈现与持久回放
 ### 实时呈现与持久回放
 
 
 `agent/assistant-stream` 发布进程本地 start、瞬态 chunk 与 end frame。loop 会在 committed end frame 命名其类型和序号前追加完整的 `assistant/message` 或 `assistant/attempt`。abandoned end 没有 settlement。
 `agent/assistant-stream` 发布进程本地 start、瞬态 chunk 与 end frame。loop 会在 committed end frame 命名其类型和序号前追加完整的 `assistant/message` 或 `assistant/attempt`。abandoned end 没有 settlement。
 
 
-Web follow adapter 显式选择接收这些进程本地 frame,并为每个 start 补充当时观察到的最后一个持久序号。它把 chunk 呈现为持久 cursor 之间的 Client-only `assistant/live-chunk` update,只暂存 start 之后匹配的 settlement,并在 revision 缺口时重新打开 follow。committed end 会发布具名 settlement delta,删除该 attempt 的 transient match、加入持久 entry,并只重放受影响的 Conversation Context;abandoned end 会发布不含 entry 的同类 delta。重连 baseline 携带活跃 attempt 的持久起始 cursor 与紧凑前缀。分页历史、replay、遥测、token 记账与冷 UI 组装读取持久嵌入式 stream,而不是 live frame。
+Web follow adapter 显式选择接收这些进程本地 frame,并为每个 start 补充当时观察到的最后一个持久序号。它把 chunk 呈现为持久 cursor 之间的 Client-only `assistant/live-chunk` update,只暂存 start 之后匹配的 settlement,并在 revision 缺口时重新打开 follow。committed end 会发布具名 settlement delta,删除该 attempt 的 transient match、加入持久 entry,并只重放受影响的 Conversation Context;abandoned end 会发布不含 entry 的同类 delta。重连 baseline 携带活跃 attempt 的持久起始 cursor 与紧凑前缀。
+
+Client event source 原样传递持久 settlement。Chat 与 Trajectory 的 Assistant node 在 attempt 活跃期间折叠 `assistant/live-chunk`,直接从 `assistant/message` 构建 settled output,并且不为展示重放 `assistant/attempt` stream。因此冷恢复的 settled presentation 不会重建逐 token timing;其他消费方需要精确证据时仍可展开持久 stream。
 
 
 ### 已发布 v1 到 v2 迁移
 ### 已发布 v1 到 v2 迁移
 
 
-相邻迁移会校验完整的冻结 v1 产物,按 turn、step、terminal boundary 与精确 message provenance 对 chunk 分组,再为每个 attempt 替换一个 settlement。成功分组的 chunk 移入其 message。未被认领的分组会在最后一个被消费 chunk 的位置变成 `assistant/attempt`。无关的交错事件保持相对顺序,存活事件获得密集 v2 序号。该迁移边通过 `dsh-llm` 运行时的 `AssistantStreamAccumulator`、`expandAssistantStream` 与 `BlockAssembler` 压缩、展开并重组嵌入 stream,而不持有冻结副本,因为该包拥有 v2 stream 编码。目标校验会自行复核每个迁移后的 `assistant/message` 与其嵌入 stream 是否一致,因此不一致的 v1 日志会作为 unsupported migration 被拒绝并保留源产物,而不是由 installed Session restoration 报告为损坏。日后若某个格式改变 stream 编码,必须把这些 helper 的冻结副本纳入本迁移边。
+相邻迁移会校验完整的冻结 v1 产物,按 turn、step、terminal boundary 与精确 message provenance 对 chunk 分组,再为每个 attempt 替换一个 settlement。成功分组的 chunk 移入其 message。未被认领的分组会在最后一个被消费 chunk 的位置变成 `assistant/attempt`。无关的交错事件保持相对顺序,存活事件获得密集 v2 序号。该迁移边通过 `dsh-llm` 运行时的 `AssistantStreamAccumulator` 压缩嵌入 stream,而不持有冻结副本,因为该包拥有 v2 stream 编码。隔离的 publication verifier 通过 `expandAssistantStream()` 与 `BlockAssembler` 展开并重组写入后的 stream,并在发布前检查每个迁移后的 `assistant/message` 是否与其一致。日后若某个格式改变 stream 编码,必须把这些 helper 的冻结副本纳入本迁移边。
 
 
 该迁移边会重映射有限的已声明引用清单:信封 provenance、surface replacement 端点、command source event、compaction range 与 shadowed list,以及 title message list。经过校验的 `session/title-llm-request` 模型可见文本会在源序号命名空间中保持逐字节不变,而它的 `messageSeqs` 字段会迁移到 v2 命名空间;因此目标校验不会根据重映射后的序号重建该文本。指向被消费 chunk 的引用会使迁移失败;它绝不会被重定向到含义不同的 settlement。该迁移边也会拒绝切开 attempt 的继承切点。
 该迁移边会重映射有限的已声明引用清单:信封 provenance、surface replacement 端点、command source event、compaction range 与 shadowed list,以及 title message list。经过校验的 `session/title-llm-request` 模型可见文本会在源序号命名空间中保持逐字节不变,而它的 `messageSeqs` 字段会迁移到 v2 命名空间;因此目标校验不会根据重映射后的序号重建该文本。指向被消费 chunk 的引用会使迁移失败;它绝不会被重定向到含义不同的 settlement。该迁移边也会拒绝切开 attempt 的继承切点。
 
 
@@ -49,7 +51,7 @@ Generation 选择与发布遵循[已发布 Session 迁移决策](2026-08-31-rele
 
 
 合并前的 performance acceptance 在三轮、100 组 warmup pair 与 600 组 measured pair 下,针对同一批已经解析的物理 row,把静态 catalog routing 与直接 released-v2 restoration 比较;它不比较 v1 与 v2,也不计入 backend I/O。每个 pooled median 与 p95 regression 都保持在 5% 预算以内,最差 p95 regression 为 3.150%。
 合并前的 performance acceptance 在三轮、100 组 warmup pair 与 600 组 measured pair 下,针对同一批已经解析的物理 row,把静态 catalog routing 与直接 released-v2 restoration 比较;它不比较 v1 与 v2,也不计入 backend I/O。每个 pooled median 与 p95 regression 都保持在 5% 预算以内,最差 p95 regression 为 3.150%。
 
 
-Agent-loop 测试固定先持久后 end 的顺序、中断的可见前缀、失败与重试 attempt、abandonment、usage 与 replay metadata。Session Controller 与 Conversation 测试固定实时瞬态显示、重连 baseline、committed settlement 发布、历史回放以及 Chat 与 Trajectory 一致性;TypeScript 与 Python SDK snapshot 固定外部事件表示。
+Agent-loop 测试固定先持久后 end 的顺序、中断的可见前缀、失败与重试 attempt、abandonment、usage 与 replay metadata。Session Controller 与 Conversation 测试固定实时瞬态显示、重连 baseline、committed settlement 发布与历史回放。Chat 与 Trajectory 测试固定实时 partial 展示和最终 message 的直接投影;TypeScript 与 Python SDK snapshot 固定外部事件表示。
 
 
 ## 备选方案
 ## 备选方案
 
 
@@ -59,13 +61,15 @@ Agent-loop 测试固定先持久后 end 的顺序、中断的可见前缀、失
 
 
 **通过历史 API 传递 packed chunk row。** 这会减少 v1 的 wire 与 Client 工作,却让 Client 拥有第二套事件词汇,并让传输继续与 token-row 基数耦合。当前 API 携带标量持久 settlement,并使用独立的实时瞬态 stream。
 **通过历史 API 传递 packed chunk row。** 这会减少 v1 的 wire 与 Client 工作,却让 Client 拥有第二套事件词汇,并让传输继续与 token-row 基数耦合。当前 API 携带标量持久 settlement,并使用独立的实时瞬态 stream。
 
 
+**在 Session Controller 中删除嵌入式 stream。** 这会减少 Client 保留的内存,却会引入第二种持久事件类型,并让面向传输的 owner 决定展示消费方需要哪些证据。实测瓶颈来自重复展开,因此由各 UI 消费方决定是否检查原样传递的 settlement。
+
 **把 stream 存在 sidecar 或 replay-only fixture 中。** 这会把一个 attempt 的 message 与证据拆给不同持久性 owner,也无法让普通恢复 Session 获得相同的失败输出与时间事实。settlement 是原子 owner。
 **把 stream 存在 sidecar 或 replay-only fixture 中。** 这会把一个 attempt 的 message 与证据拆给不同持久性 owner,也无法让普通恢复 Session 获得相同的失败输出与时间事实。settlement 是原子 owner。
 
 
 **把被消费 chunk 的引用重定向到其 settlement。** Chunk 与 attempt settlement 不是可互换事实。拒绝可以防止迁移悄然改变插件自有引用的含义。
 **把被消费 chunk 的引用重定向到其 settlement。** Chunk 与 attempt settlement 不是可互换事实。拒绝可以防止迁移悄然改变插件自有引用的含义。
 
 
 ## 后果
 ## 后果
 
 
-当前日志、遥测、历史页与冷 Client 组装按模型 attempt 而非 token chunk 扩展,同时在每个 settlement 内保留精确 stream 证据。实时呈现保持增量,并且有意仅存在于进程内。
+当前日志、遥测与历史页按模型 attempt 而非 token chunk 扩展,同时在每个 settlement 内保留精确 stream 证据。Client event window 保留这份紧凑证据,但 Chat 与 Trajectory 的 Assistant node 不会把 settled stream 展开成逐 delta 对象。实时呈现保持增量,并且有意仅存在于进程内。
 
 
 v1 的顶层 chunk 可能在 attempt 结束前由带缓冲的持久化 writer 刷盘;与之不同,v2 在 settlement 之前没有持久 attempt 证据。如果进程或主机在 settlement 前硬中断,完整的 in-flight stream 都会丢失;`agent/assistant-stream` 不是 write-ahead log。这项取舍避免为实时输出增加第二个持久性 owner。
 v1 的顶层 chunk 可能在 attempt 结束前由带缓冲的持久化 writer 刷盘;与之不同,v2 在 settlement 之前没有持久 attempt 证据。如果进程或主机在 settlement 前硬中断,完整的 in-flight stream 都会丢失;`agent/assistant-stream` 不是 write-ahead log。这项取舍避免为实时输出增加第二个持久性 owner。
 
 

+ 6 - 0
.agents/notes/implemented/architecture/2026-09-05-canonical-feedback-log.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-09-05-canonical-feedback-log.md
+2026-09-05-canonical-feedback-log.md: c890064817ac0604fa4cc2b4073850174195e511
+2026-09-05-canonical-feedback-log.zh.md: e7e8563df8ad37d7aba641e58da4fc8872ec10b9

+ 33 - 0
.agents/notes/implemented/architecture/2026-09-05-canonical-feedback-log.md

@@ -0,0 +1,33 @@
+# Agent Note: Canonical feedback log and request delivery
+
+Status: implemented
+
+English | [中文](2026-09-05-canonical-feedback-log.zh.md)
+
+## Problem
+
+Editable message ratings need one durable authority that Session export and request delivery can retain. A separate feedback store makes those consumers incomplete and introduces a second commit relationship with the target message. Recording a human judgment must not change model input or imply that a collector accepted it.
+
+## Decision
+
+The canonical Session log owns feedback. Session-level remarks use `feedback/record`; material message edits and deletions use `feedback/message-put` and `feedback/message-delete`. All are log-only. The service folds current items from events matching the requested `sessionId`, so inherited parent events do not become a fork's current feedback. Deletion removes the current item, not earlier ratings or notes from the log.
+
+Live message-feedback mutations append through the owning Session and await its durability checkpoint; cold mutations hold a persistence write handle across read, comparison, append, and flush without creating a Session or Agent. A matching no-op appends nothing but still awaits persistence. Failures propagate, and a failed live flush can leave an observable in-memory item for retry. Per-item versions prevent unrelated message edits from conflicting; strict stale-write rejection prevents ABA overwrites even when the desired value matches. Target validation binds a judgment to a sent assistant message, and forks keep independent judgments. These choices retain rationale recorded in the [archived sidecar decision](../../archived/architecture/2026-08-10-message-feedback-sidecar.md), whose storage and commit mechanism is superseded.
+
+The existing opt-in [session-log-deepseek contribution](../../../../packages/session/session-log-deepseek/README.md) includes feedback in the ordinary `dsh_session_log` suffix on a subsequent eligible request. It uses the existing DeepSeek destination selection and acceptance watermark. There is no separate `dsh_feedback` uploader, feedback-triggered LLM request, or model-input field. The [explicit-feedback OTel decision](2026-09-05-nonofficial-feedback-otel.md) owns the independent feedback-triggered upload for all users and providers.
+
+The command confirms recording with the Session and anonymous user ids, without depending on telemetry or disclosing its policy. Its append remains unflushed. This supersedes the command-copy decision in the [archived sharing disclosure note](../../archived/feature/2026-08-07-feedback-acknowledgement-sharing-disclosure.md). The [telemetry service's policy API](../../../../packages/session/session-telemetry/README.md#the-sharing-disclosure) remains independently available: a backend discloses its policy, not delivery or retention, and the optional OTel package does not own that vocabulary.
+
+## Alternatives considered
+
+**Keep the sidecar.** It supports destructive local edits, but cannot make feedback part of ordinary canonical-log export and delivery without another join and durability relationship.
+
+**Reuse `feedback/record` for message edits.** A free-text Session remark does not identify an item mutation. Distinct events preserve message identity and deletion semantics; upload policy remains consumer-owned.
+
+**Add a dedicated feedback uploader or immediate LLM request.** The opt-in log contribution carries canonical events on eligible requests. The existing OTel pipeline independently handles explicit-feedback uploads for all providers, without a custom feedback uploader or another model request.
+
+## Consequences
+
+Feedback survives ordinary log export and replay without consuming model-input tokens or changing KV Cache. Current-item deletion is not erasure. The DeepSeek request contribution can leave final feedback local until another eligible request; OTel sends an authorized batch independently under its own policy. The Web controller remains a unary Remote consumer and does not consume feedback log events for cross-tab updates.
+
+[Message-feedback tests](../../../../packages/feedback/message-feedback/tests/message-feedback.spec.ts) cover material events, no-ops, strict versions, fork isolation, and persistence failures. The [request contribution tests](../../../../packages/session/session-log-deepseek/tests) own suffix acceptance and retry; the [command tests](../../../../packages/feedback/command-feedback/tests/command-feedback.spec.ts) pin the plain confirmation.

+ 33 - 0
.agents/notes/implemented/architecture/2026-09-05-canonical-feedback-log.zh.md

@@ -0,0 +1,33 @@
+# Agent Note: 权威反馈日志与请求投递
+
+Status: implemented
+
+[English](2026-09-05-canonical-feedback-log.md) | 中文
+
+## 问题
+
+可编辑的消息评分需要一个能由 Session 导出与请求投递保留的持久权威来源。独立的反馈存储会让这些消费方拿到不完整的数据,并引入与目标消息之间的第二套提交关系。记录人类判断不能改变模型输入,也不能暗示采集端已经接受数据。
+
+## 决策
+
+权威 Session 日志拥有反馈。Session 级备注使用 `feedback/record`;消息的实质编辑与删除使用 `feedback/message-put` 和 `feedback/message-delete`。三者都仅写日志。服务从与请求的 `sessionId` 匹配的事件中归约当前条目,因此继承的父级事件不会成为 fork 的当前反馈。删除会移除当前条目,但不会抹除日志中早先的评分或备注。
+
+live 消息反馈变更通过所属 Session 追加,并等待其持久化检查点;cold 变更在读取、比较、追加和 flush 期间持有持久化写句柄,不创建 Session 或 Agent。匹配版本的无变更操作不追加事件,但仍等待持久化。故障会原样传播,live flush 失败可能留下可观测的内存条目以供重试。逐条版本避免不同消息的编辑互相冲突;严格拒绝陈旧写入避免 ABA 覆盖,即使期望值已经匹配也不例外。目标校验把判断绑定到已发送的 assistant 消息,fork 保持独立判断。这些选择保留[已归档伴随记录决策](../../archived/architecture/2026-08-10-message-feedback-sidecar.md)记载的理由,但其存储与提交机制已被取代。
+
+现有需显式启用的 [session-log-deepseek 贡献](../../../../packages/session/session-log-deepseek/README.zh.md)会在后续符合条件的请求中,把反馈纳入普通 `dsh_session_log` 后缀。它使用现有的 DeepSeek 目标选择和接受水位。没有独立的 `dsh_feedback` 上传器、反馈触发的 LLM 请求或模型输入字段。[显式反馈 OTel 决策](2026-09-05-nonofficial-feedback-otel.zh.md)负责面向所有用户和提供方的独立反馈触发上传。
+
+命令用 Session 与匿名用户 id 确认记录,不依赖遥测,也不披露其策略。其追加仍不执行 flush。这取代[已归档共享披露记录](../../archived/feature/2026-08-07-feedback-acknowledgement-sharing-disclosure.md)中的命令文案决策。[遥测服务的策略 API](../../../../packages/session/session-telemetry/README.zh.md#the-sharing-disclosure) 仍可独立使用:后端披露策略,而不保证投递或保留,可选 OTel 包不拥有这套词汇。
+
+## 考虑过的替代方案
+
+**保留伴随记录。** 它支持破坏性的本地编辑,但若不增加关联读取及持久化关系,就无法让反馈参与普通权威日志导出与投递。
+
+**对消息编辑复用 `feedback/record`。** 自由文本的 Session 备注不能标识条目变更。独立事件保留消息身份和删除语义;上传策略仍由消费方负责。
+
+**增加专用反馈上传器或立即发起 LLM 请求。** 需显式启用的日志贡献在符合条件的请求上传送权威事件。现有 OTel 流水线独立处理所有提供方的显式反馈上传,无需自定义反馈上传器或另一个模型请求。
+
+## 后果
+
+反馈随普通日志导出与回放保留,不消耗模型输入 token,也不改变 KV Cache。删除当前条目不等于抹除历史。DeepSeek 请求贡献可能让最终反馈留在本地,直到下次符合条件的请求;OTel 按自身策略独立发送已授权批次。Web 控制器仍消费一元 Remote,不消费反馈日志事件来更新其他标签页。
+
+[消息反馈测试](../../../../packages/feedback/message-feedback/tests/message-feedback.spec.ts)覆盖实质事件、无变更操作、严格版本、fork 隔离与持久化故障。[请求贡献测试](../../../../packages/session/session-log-deepseek/tests)负责后缀接受与重试;[命令测试](../../../../packages/feedback/command-feedback/tests/command-feedback.spec.ts)固定纯确认文本。

+ 6 - 0
.agents/notes/implemented/architecture/2026-09-05-nonofficial-feedback-otel.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-09-05-nonofficial-feedback-otel.md
+2026-09-05-nonofficial-feedback-otel.md: bbc947455f7a84a41af6651223eb62129fa8d452
+2026-09-05-nonofficial-feedback-otel.zh.md: de9b0322861118cdc1163a0faefb172fc1062776

+ 33 - 0
.agents/notes/implemented/architecture/2026-09-05-nonofficial-feedback-otel.md

@@ -0,0 +1,33 @@
+# Agent Note: Explicit-feedback-only upload through OpenTelemetry
+
+Status: implemented
+
+English | [中文](2026-09-05-nonofficial-feedback-otel.zh.md)
+
+## Problem
+
+Feedback needs the session context it describes and a delivery path independent of the model provider or a later model request. Ordinary activity must not authorize uploads. Inherited feedback must not count as a child Session's consent.
+
+## Decision
+
+The base mounts OTel in `FEEDBACK_ONLY` for all users and providers, including `deepseek-official` and Sessions without a request header. Only new own `feedback/record`, `feedback/message-put`, and `feedback/message-delete` events authorize capture through that exact canonical event. Text feedback, ratings, material note edits, and withdrawals count. Cold `feedback/committed` notifications supply a committed snapshot without publishing a live Session or Agent.
+
+An authorized prefix includes all unhanded canonical context from seq 0 through the feedback, not just its payload. A child needs its own new feedback; its prefix then includes inherited history. Later records wait for the next explicit feedback. Request activity, request headers, Session creation or adoption, restoration, plugin mount, and HMR never authorize capture; stored feedback alone triggers nothing.
+
+The backend uses on-demand capture with complete history and the existing redaction waterfall. `DISABLED` constructs no transport. `FULL` is rejected rather than aliased. Direct `ctx.sessionTelemetry.emit()` calls are no-ops, so callers cannot bypass feedback authorization. SDK scheduled flush and shutdown may finish previously authorized batches but never capture new records. Sending after submission needs no further user interaction or model call.
+
+The [canonical-feedback decision](2026-09-05-canonical-feedback-log.md) owns storage, versions, deletion, and plain command confirmation. The [opt-in DeepSeek contribution](../../../../packages/session/session-log-deepseek/README.md) remains independent, with its existing destination and acceptance behavior.
+
+## Alternatives considered
+
+**Filter by provider or endpoint hostname.** Feedback authorizes the same bounded context for every user; a provider choice, gateway, or missing header does not change that authorization.
+
+**Use only later DeepSeek requests.** Other providers do not carry `dsh_session_log`, and final feedback may have no subsequent request. The existing OTel pipeline sends independently without a custom uploader or model call.
+
+**Keep continuous capture or replay stored feedback on lifecycle events.** Deployment configuration and old feedback do not authorize new capture. Only a new explicit submission does. Parent feedback likewise cannot authorize a child upload.
+
+## Consequences
+
+Handoff is best-effort, not collector acceptance. Same-object cursors suppress repeated capture, but fresh cold snapshots and new feedback after restart can repeat prefixes; receivers deduplicate on `(session.id, session.format_version, event.seq)`. There is no durable OTel outbox, delivery watermark, or harness HTTP retry promise. SDK batching and loss behavior apply after enqueue. OTel and the opt-in DeepSeek path can overlap. Withdrawal exports a deletion event, not remote erasure.
+
+[OTel tests](../../../../packages/session/session-telemetry-otel/tests/otel.spec.ts) cover explicit-feedback capture, provider-independent behavior, lifecycle silence, fork consent, cold commits, and direct-call denial. [Coordinator tests](../../../../packages/session/session-telemetry/tests/telemetry.spec.ts) cover history capture; [base tests](../../../../packages/bundle/base/tests/base.spec.ts) pin the mounted default.

+ 33 - 0
.agents/notes/implemented/architecture/2026-09-05-nonofficial-feedback-otel.zh.md

@@ -0,0 +1,33 @@
+# Agent Note: 仅在显式反馈后通过 OpenTelemetry 上传
+
+Status: implemented
+
+[English](2026-09-05-nonofficial-feedback-otel.md) | 中文
+
+## 问题
+
+反馈需要它所描述的会话上下文,以及不依赖模型提供方或后续模型请求的投递路径。普通活动不得授权上传。继承的反馈不得视为子 Session 的同意。
+
+## 决策
+
+基础配置为所有用户和提供方以 `FEEDBACK_ONLY` 挂载 OTel,包括 `deepseek-official` 和没有请求头的 Session。只有新的自身 `feedback/record`、`feedback/message-put` 和 `feedback/message-delete` 事件授权捕获,且截止该确切的权威事件。文本反馈、评分、实质备注编辑与撤回均算作反馈。冷会话 `feedback/committed` 通知提供已提交快照,不发布存活 Session 或 Agent。
+
+授权前缀包含从 seq 0 到该反馈的所有尚未交接的权威上下文,而非只有反馈载荷。子会话需要新的自身反馈;之后其前缀包含继承历史。后续记录等待下一次显式反馈。请求活动、请求头、Session 创建或接纳、恢复、插件挂载和 HMR(热模块替换)绝不授权捕获;仅有存储的反馈不会触发任何上传。
+
+后端使用包含完整历史的按需捕获与现有脱敏 waterfall(瀑布式事件)。`DISABLED` 不构造传输。`FULL` 被拒绝,不作为别名。直接调用 `ctx.sessionTelemetry.emit()` 是空操作,因此调用方不能绕过反馈授权。SDK 定时刷新和关闭可以完成先前已授权的批次,但绝不捕获新记录。提交后的发送无需进一步用户交互或模型调用。
+
+[权威反馈决策](2026-09-05-canonical-feedback-log.zh.md)负责存储、版本、删除与纯命令确认。[需主动开启的 DeepSeek 贡献](../../../../packages/session/session-log-deepseek/README.zh.md)保持独立,保留现有目标与接受行为。
+
+## 考虑过的替代方案
+
+**按提供方或端点主机名过滤。** 反馈为每位用户授权相同的有界上下文;提供方选择、网关或缺失请求头不改变该授权。
+
+**仅使用后续 DeepSeek 请求。** 其他提供方不携带 `dsh_session_log`,而最终反馈之后可能没有请求。现有 OTel 流水线可独立发送,无需自定义上传器或模型调用。
+
+**保留持续捕获或在生命周期事件上回放存储的反馈。** 部署配置和旧反馈不授权新捕获。只有新的显式提交才授权。父会话反馈同样不能授权子会话上传。
+
+## 后果
+
+交接尽力而为,不代表采集端接受。同对象游标抑制重复捕获,但新冷快照和重启后的新反馈可能重复前缀;接收方按 `(session.id, session.format_version, event.seq)` 去重。没有持久化 OTel outbox、投递水位或 harness HTTP 重试承诺。入队后适用 SDK 批处理与丢失行为。OTel 与需主动开启的 DeepSeek 路径可能重叠。撤回导出删除事件,不是远端擦除。
+
+[OTel 测试](../../../../packages/session/session-telemetry-otel/tests/otel.spec.ts)覆盖显式反馈捕获、提供方无关行为、生命周期静默、fork 同意、冷会话提交与直接调用拒绝。[协调器测试](../../../../packages/session/session-telemetry/tests/telemetry.spec.ts)覆盖历史捕获;[基础配置测试](../../../../packages/bundle/base/tests/base.spec.ts)固定挂载默认值。

+ 6 - 0
.agents/notes/implemented/architecture/2026-09-05-read-only-session-migration-preparation.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-09-05-read-only-session-migration-preparation.md
+2026-09-05-read-only-session-migration-preparation.md: c343ca457184554f4b47a0795dcb33b8b07e9d39
+2026-09-05-read-only-session-migration-preparation.zh.md: 3a283259729eda6a01fa4208cac5399f45978fac

+ 170 - 0
.agents/notes/implemented/architecture/2026-09-05-read-only-session-migration-preparation.md

@@ -0,0 +1,170 @@
+# Agent Note: Historical Session reads prepare before write publication
+
+Status: implemented
+
+English | [中文](2026-09-05-read-only-session-migration-preparation.zh.md)
+
+## Problem
+
+The stateful Stage pipeline makes historical Decode and migration bounded and fast, but a serial persistence open still performs encode, sync, Worker verification, publication, and committed reopen before returning either handle kind. A read-only consumer therefore waits for about 2.2 seconds of work that it does not need and mutates storage merely to display history.
+
+### Serial readiness cost
+
+- History pagination, projection preparation, export, and the opening `session.follow` snapshot need only the validated current logical artifact.
+- A historical read open nevertheless creates and syncs a temporary v2 generation, starts a full verification Worker, rechecks the source, publishes v2, and reopens the target.
+- The migration result already exists in memory before encode, but the serial API returns only the committed physical snapshot. Persistence must Decode current bytes again to reconstruct the same logical events.
+- `session.follow` cannot deliver its opening snapshot until publication completes, even though Agent resume is the first operation that requires append access.
+- Read-only storage cannot serve a logically valid historical Session because read open requires generation publication.
+
+### A naive split would break lifecycle guarantees
+
+- Returning a write handle before verification would route append into an unpublished temporary file and create a second durability state for accepted events.
+- Starting publication automatically after every read would require backend ownership for task failure, shutdown, cleanup, and a later writer joining work it did not request.
+- A shared preparation cannot inherit the first caller's AbortSignal. One cancelled reader must not terminate work still awaited by another.
+- A read handle must initially serve prepared memory but later observe a current file and its appended tail after another caller publishes.
+- Once readers have observed one prepared artifact, source drift cannot silently rerun migration and substitute a different logical history.
+
+## Decision
+
+The JSONL backend separates logical preparation from durable publication. Read open waits only for preparation. Write open reuses a matching preparation and waits for publication before returning a writable handle.
+
+### Prepared generation API
+
+```text
+interface PreparedJsonlMigration {
+  readonly sourceIdentity: JsonlPhysicalIdentity
+  readonly artifact: SessionFormatArtifact
+  publish(): Promise<JsonlPhysicalIdentity>
+}
+```
+
+`prepareJsonlMigration()` reads one stable historical revision, runs the complete Stage chain once, and returns the current artifact without encoding or writing. `publish()` is idempotent: concurrent and later calls share one terminal Promise, including its rejection, and cannot encode the same prepared artifact twice.
+
+`publish()` streams current records into an exclusively created same-directory temporary file, syncs it, awaits the bounded Worker verifier, compares the source identity captured by preparation, and publishes the canonical path without overwrite. The successful publisher reuses the prepared logical artifact instead of decoding its target. A losing publisher verifies that the winner begins with the exact staged migration prefix; append tail validation remains a current-reader responsibility.
+
+Publication runs to settlement after invocation and is not cancelled midway by the write caller. Write open checks its caller signal before and after publication, so an abort can reject the open after the successor commits without leaking the write lease. A source identity change throws `JsonlGenerationSourceChangedError`, removes the temporary file, and does not repeat Decode or migration.
+
+### Preparation ownership and cancellation
+
+The persistence backend keeps one in-flight entry per Session id, selected source path, and stat-derived revision:
+
+```text
+interface MigrationPreparation {
+  sourcePath: string
+  sourceRevision: SessionPersistenceRevision
+  controller: AbortController
+  promise: Promise<PreparedStoredLog>
+  settled: boolean
+  waiters: number
+}
+```
+
+A new read or write open joins the existing entry only when its source path and revision still match. `waitWithAbort()` races each caller's AbortSignal against the shared Promise without forwarding that signal to shared work. The backend-owned controller is aborted only when the last waiter leaves while preparation is still running.
+
+Completed results enter the existing bounded `coldLogMemo`. The `StoredLog` discriminant separates published current state from `PreparedStoredLog`, whose `publication` field binds current logical events to their matching publication operation. A query followed by Agent resume therefore reuses the same Decode and migration result. The in-flight map owns only running work; it is not a second completed-result cache.
+
+`SessionHandle.read()` reports whether its event values are detached or shared-frozen. The JSONL backend deep-freezes each decoded event graph once before memoization and creates the `shared-frozen` result there; later reads and slices preserve that producer-established state even when the slice is empty. `readColdSessionLog()` combines those values with locally owned interrupted-turn closers and passes the `eventState` through `SessionObservationReader`; `Session.fromRestore()` validates and adopts the seed without copying or freezing. Ordinary create and fork seeds keep their defensive snapshot path.
+
+Read-only restoration validates the event and settlement fields required by Session runtime behavior but does not expand every embedded Assistant stream. The publication Worker retains complete stream replay and checks content, usage, and replay-state agreement before a migrated successor is committed. Existing current-v2 files rely on their writer; consumers that expand a compact stream validate its records when they read it.
+
+### Read handle transition
+
+A read open adopts a handle with prepared events in `state.primed` while no current generation exists. Each later `read()` resolves the current path:
+
+```text
+if current generation is absent:
+  return slice of primed events
+else:
+  clear primed events
+  read current generation and enforce non-shrinking history
+```
+
+`resolveCurrentLog()` may therefore return `undefined` for an existing historical Session: it answers whether a current canonical file exists, not whether the Session can be read. Public `stat` and `list` continue to discover the historical header.
+
+### Write-open publication
+
+Write open acquires the process-local claim and kernel-backed cross-process lease before re-resolving the selected generation. If it remains historical, it obtains or reuses the prepared `StoredLog` and awaits `publish()`. Only then does it return a write handle primed with the prepared events.
+
+```text
+write open
+  → claim process-local ownership
+  → acquire SessionWriteLease
+  → re-resolve generation
+  → join or create preparation
+  → encode + sync temp
+  → Worker verify
+  → source identity check
+  → no-overwrite publish
+  → return writable handle
+```
+
+No external caller can append before the handle exists. `append`, `flush`, and `close` therefore retain their ordinary current-generation behavior and never need a “publishing” branch. Service `flush()` continues to flush only already adopted writers; it does not turn a read-only preparation into a write.
+
+### Follow and Agent promotion
+
+`session.follow` opens history through the read path, restores the Session and projections, emits the opening snapshot, and then starts Agent promotion. Agent resume uses write open, so it waits for publication before the Agent accepts a new turn. History visibility and write readiness are separate timing points without introducing an unpublished append state.
+
+## Problem-to-solution mapping
+
+| Serial-flow problem | Implemented mechanism | Guarantee |
+|---|---|---|
+| Read-only callers wait for encode and verify | Read open returns prepared events | First content waits only for Decode and migration |
+| Concurrent historical opens repeat work | Session/source-revision keyed single-flight | One migration per selected revision |
+| First caller owns shared cancellation | Caller-local `waitWithAbort()` plus backend controller | One cancellation does not kill other waiters |
+| Preparation is lost between query and resume | `PreparedStoredLog.publication` in bounded memo | Write open reuses the same artifact |
+| No current path exists for a read handle | Primed in-memory read | Historical data is readable before publication |
+| Read handle must observe later append | Re-resolve and switch from primed data to current file | Existing handles converge after publication |
+| Append before verification is unsafe | Publish inside write open before returning the handle | Returned writer is immediately durable-ready |
+| Automatic background publication has no owner | Only write open invokes `publish()` | No orphan write task from read-only access |
+| Source changes after readers saw the artifact | Fail publication without rerunning migration | Exposed logical history is never silently replaced |
+
+## Verification
+
+The benchmark uses the same 116,228,655-byte v0 Zstandard Session as the Stage decision. The first table compares every relevant implementation; the detailed scheduling comparison then holds the Codec/Stage chain constant between #3585 and preparation-first scheduling.
+
+### First opening of historical data
+
+| Implementation | Session restored | CPU | Peak RSS | Retained heap | Result |
+|---|---:|---:|---:|---:|---|
+| Original high-performance v0 reader | 4.594s | 6.048s | 2.720GB | 2.016GB | Reads about 9.14 million v0 events without migration |
+| Master whole-artifact v0-to-v2 migration | >72.8s | — | Decode stage reached at least 7.219GB | — | OOM before returning a handle |
+| #3585 streaming migration with serial publication | 6.241s | 8.493s | 2.107GB | 477MB | Produces and publishes a 72,784-event v2 Session |
+| #3586 preparation-first scheduling | 2.954s | — | 1.026GB | 463MB | Produces the same v2 Session and defers publication until write open |
+
+Preparation-first restoration is 53% faster than #3585 and 36% faster than the original high-performance reader even though it also migrates the artifact to v2.
+
+### Scheduling observation points
+
+| User-visible point | Serial publication | Preparation-first | Change |
+|---|---:|---:|---:|
+| Read open plus Session restoration | 6.241s | 2.954s | -53% |
+| `session.follow` opening snapshot | 7.587s | 2.912s | -62% |
+| Agent receives writable Session | 6.246s | 5.161s | -17% |
+| Reopen an already-current v2 Session | 1.284s | 0.964s | -25% |
+| Follow opening-snapshot peak RSS | 2.353GB | 1.059GB | -55% |
+
+Preparation spends about 2.61 seconds in Decode and migration. Deferred publication takes about 2.56 seconds: 0.83 seconds for encode/write/sync, 1.72 seconds for strict Worker verification, and about 0.005 seconds for source check and atomic publication. A read-only request performs none of that publication work.
+
+The prepared artifact and Session restoration peak near 1.03 GB RSS. Preparation and Worker verification together peak near 2.19 GB because the parent retains the logical artifact while the Worker independently validates the physical generation.
+
+Tests cover shared-waiter cancellation, all-waiters cancellation, memo handoff, read-handle switching, source drift, winner collision, publication idempotence, write-open ordering, Worker failure, and the plain-Node bundled Worker entry.
+
+## Consequences
+
+Read-only body access does not publish a generation. The first writer pays publication once before append. A configured JSONL root must still be readable and structurally valid, but historical body migration itself does not require a successor write.
+
+The bounded memo retains one migrated event array to bridge read and write opens. This is intentional: avoiding that retained artifact would require a second Decode and migration or would prevent early read availability.
+
+Publication failure rejects Agent resume and other write opens but does not invalidate read results already delivered from the unchanged historical source. Source drift is terminal for that write attempt rather than a trigger to recompute hidden state.
+
+The backend still has a broader pre-existing lifecycle gap: dispose does not own every `create()` or `open()` operation that has not yet returned a handle. This decision does not add migration-specific tracking to `flush()` or solve that general pending-operation problem.
+
+## Alternatives considered
+
+- **Keep serial publication for every open** — is the simplest physical state model but adds about 2.2 seconds to read-only first content and requires writable storage.
+- **Publish automatically in the background after read** — needs backend task ownership, shutdown quiescence, error reporting, and writer joining even when no caller requested a write.
+- **Return a writer before verification** — requires append to an unpublished stage and creates an additional durability and failure state for accepted events.
+- **Give each caller an independent preparation** — repeats the dominant Decode and migration work and multiplies peak memory under concurrent list/follow/resume operations.
+- **Let the first caller's signal cancel shared work** — makes later callers depend on unrelated cancellation timing.
+- **Rerun migration after source drift** — can replace history already shown to readers and makes one logical operation process the same large file more than once.
+- **Always keep read handles on primed memory** — prevents an existing handle from seeing later append and diverges from ordinary persistence refresh behavior.

+ 170 - 0
.agents/notes/implemented/architecture/2026-09-05-read-only-session-migration-preparation.zh.md

@@ -0,0 +1,170 @@
+# Agent Note: 历史 Session 在写入发布前提供只读迁移结果
+
+Status: implemented
+
+[English](2026-09-05-read-only-session-migration-preparation.md) | 中文
+
+## 问题
+
+有状态 Stage pipeline 已经把历史 Decode 与 migration 恢复到有界、高性能的数据流,但串行 persistence open 仍会在返回任一种 handle 前执行 encode、sync、Worker verification、publication 与 committed reopen。只读 consumer 因此需要额外等待约 2.2 秒不需要的工作,而且仅为展示历史就会修改存储。
+
+### 串行 readable 的额外代价
+
+- 历史分页、projection preparation、export 与 `session.follow` 的 opening snapshot 只需要已经校验的 current logical artifact。
+- Historical read open 仍会创建并 sync 临时 v2 generation、启动完整 verification Worker、复查 source、发布 v2 并重新打开 target。
+- Migration 在 encode 前已经得到完整 current artifact,但串行 API 只返回 committed physical snapshot。Persistence 必须再次 Decode current bytes 才能重建相同逻辑事件。
+- `session.follow` 必须等 publication 完成后才能发出 opening snapshot,而 Agent resume 才是第一个真正要求 append 权限的操作。
+- Read-only storage 无法提供逻辑上有效的历史 Session,因为 read open 强制发布 generation。
+
+### 直接拆分会破坏 lifecycle 保证
+
+- Verify 前返回 write handle 会让 append 写入 unpublished temporary file,并为已接纳事件引入第二种 durability state。
+- 每次 read 后自动启动 publication,需要 backend 负责 task failure、shutdown、cleanup,以及后续 writer 加入一个自己没有请求的任务。
+- Shared preparation 不能继承第一个 caller 的 AbortSignal;一个 reader 取消不能终止其他 waiter 仍依赖的工作。
+- Read handle 必须先提供 prepared memory,并在其他 caller 发布后切换到 current file 与其 append tail。
+- Reader 已经观察一个 prepared artifact 后,source drift 不能悄悄重跑 migration 并替换成另一份逻辑历史。
+
+## 决策
+
+JSONL backend 将 logical preparation 与 durable publication 分开。Read open 只等待 preparation;write open 复用 matching preparation,并在返回 writable handle 前等待 publication。
+
+### Prepared generation API
+
+```text
+interface PreparedJsonlMigration {
+  readonly sourceIdentity: JsonlPhysicalIdentity
+  readonly artifact: SessionFormatArtifact
+  publish(): Promise<JsonlPhysicalIdentity>
+}
+```
+
+`prepareJsonlMigration()` 读取一个稳定 historical revision,只执行一次完整 Stage chain,并在不 encode、不写文件的情况下返回 current artifact。`publish()` 是幂等操作:并发与后续调用共享同一个终态 Promise(包括拒绝结果),不会对同一 prepared artifact 重复 encode。
+
+`publish()` 把 current records 流式写入同目录排他创建的 temporary file,执行 sync,等待 bounded Worker verifier,比较 preparation 捕获的 source identity,再通过 no-overwrite 操作发布 canonical path。成功 publisher 复用 prepared logical artifact,不重新 Decode 自己的 target。竞争失败者只验证 winner 以精确 staged migration prefix 开头;append tail validation 仍属于 current reader。
+
+Publication 调用后会运行到 settlement,不会被 write caller 中途取消。Write open 会在 publication 前后检查 caller signal,因此取消可能在后继已经提交后拒绝 open,但不会泄漏 write lease。Source identity 变化会抛出 `JsonlGenerationSourceChangedError`、删除临时文件,并且不会重复 Decode 或 migration。
+
+### Preparation ownership 与取消
+
+Persistence backend 按 Session id、selected source path 与 stat-derived revision 保存一个 in-flight entry:
+
+```text
+interface MigrationPreparation {
+  sourcePath: string
+  sourceRevision: SessionPersistenceRevision
+  controller: AbortController
+  promise: Promise<PreparedStoredLog>
+  settled: boolean
+  waiters: number
+}
+```
+
+新的 read/write open 只有在 source path 与 revision 仍匹配时才加入已有 entry。`waitWithAbort()` 让每个 caller 的 AbortSignal 与 shared Promise 竞争,但不会把 caller signal 传给共享工作。只有最后一个 waiter 在 preparation 仍运行时离开,backend-owned controller 才会 abort。
+
+完成结果进入既有 bounded `coldLogMemo`。`StoredLog` 判别字段把已发布 current state 与 `PreparedStoredLog` 分开,后者的 `publication` 字段把 current logical events 与匹配的 publication operation 绑定,使 query 后紧接的 Agent resume 复用同一次 Decode 与 migration。In-flight map 只拥有运行中的工作,不是第二个 completed-result cache。
+
+`SessionHandle.read()` 会报告 event value 是 detached 还是 shared-frozen。JSONL backend 在 memo 化前只对每个已解码 event graph 深度冻结一次,并在该处构造 `shared-frozen` 结果;后续读取和 slice 即使为空也会保留生产者建立的状态。`readColdSessionLog()` 将这些 event 与本地独占的 interrupted-turn closer 组合,并通过 `SessionObservationReader` 继续传递 `eventState`;`Session.fromRestore()` 只校验和接管 seed,不再复制或冻结。普通 create 与 fork seed 继续使用 defensive snapshot 路径。
+
+Read-only restoration 会校验 Session runtime 直接依赖的 event 与 settlement 字段,但不会展开每一段嵌入式 Assistant stream。Publication Worker 继续执行完整 stream replay,并在提交 migrated successor 前校验 content、usage 与 replay state 一致性。已有 current-v2 文件信任其 writer;需要展开 compact stream 的 consumer 会在读取时校验 record。
+
+### Read handle 切换
+
+Current generation 不存在时,read open 会采用在 `state.primed` 中保存 prepared events 的 handle。后续每次 `read()` 都重新解析 current path:
+
+```text
+if current generation is absent:
+  return slice of primed events
+else:
+  clear primed events
+  read current generation and enforce non-shrinking history
+```
+
+因此,一个已有 historical Session 也可能让 `resolveCurrentLog()` 返回 `undefined`:它回答的是 current canonical file 是否存在,而不是 Session 是否可读。公开 `stat` 与 `list` 继续发现 historical header。
+
+### Write-open publication
+
+Write open 先取得进程内 claim 与内核支持的跨进程 lease,再重新解析 selected generation。如果它仍是 historical,就取得或复用 prepared `StoredLog` 并等待 `publish()`。之后才返回以 prepared events 为 primed state 的 write handle。
+
+```text
+write open
+  → claim process-local ownership
+  → acquire SessionWriteLease
+  → re-resolve generation
+  → join or create preparation
+  → encode + sync temp
+  → Worker verify
+  → source identity check
+  → no-overwrite publish
+  → return writable handle
+```
+
+Handle 返回前,外部 caller 无法 append。因此 `append`、`flush` 与 `close` 保持普通 current-generation 行为,不需要“publishing”分支。Service `flush()` 继续只 flush 已经 adopt 的 writer;它不会把 read-only preparation 转成 write。
+
+### Follow 与 Agent promotion
+
+`session.follow` 通过 read path 打开历史、恢复 Session 与 projections、发出 opening snapshot,然后启动 Agent promotion。Agent resume 使用 write open,因此会在 Agent 接收新一轮对话前等待 publication。历史可见与写入就绪成为两个明确时间点,同时不引入 unpublished append state。
+
+## 问题与方案对照
+
+| 串行流程问题 | 实现机制 | 保证 |
+|---|---|---|
+| Read-only caller 等待 encode 与 verify | Read open 返回 prepared events | 首屏只等待 Decode + migration |
+| 并发 historical open 重复工作 | Session/source-revision keyed single-flight | 每个 selected revision 只迁移一次 |
+| 第一个 caller 拥有共享取消 | Caller-local `waitWithAbort()` + backend controller | 单个取消不终止其他 waiter |
+| Query 与 resume 之间丢失 preparation | Bounded memo 中的 `PreparedStoredLog.publication` | Write open 复用相同 artifact |
+| Read handle 没有 current path | Primed in-memory read | Publication 前 historical data 可读 |
+| Read handle 需要观察后续 append | 重新 resolve,并从 primed data 切到 current file | Publication 后已有 handle 收敛 |
+| Verify 前 append 不安全 | Write open 返回前完成 publication | 返回 writer 立即具备普通 durability |
+| 自动后台 publication 无 owner | 只有 write open 调用 `publish()` | Read-only access 不产生 orphan write task |
+| Reader 已看到 artifact 后 source 改变 | Publication 失败且不重跑 migration | 已暴露逻辑历史不被静默替换 |
+
+## 验证
+
+Benchmark 使用 Stage 决策中的同一份 116,228,655-byte v0 Zstandard Session。第一张表比较 migration 工作涉及的全部实现;后续调度明细则保持 #3585 与 preparation-first 使用同一条 Codec/Stage chain,仅改变 persistence 调度。
+
+### 用户首次打开历史数据
+
+| 实现 | Session restore | CPU | Peak RSS | Retained heap | 结果 |
+|---|---:|---:|---:|---:|---|
+| 原高性能 v0 reader | 4.594s | 6.048s | 2.720GB | 2.016GB | 不迁移,读取约 914 万个 v0 event |
+| Master whole-artifact v0-to-v2 migration | >72.8s | — | Decode 阶段达到至少 7.219GB | — | 返回 handle 前 OOM |
+| #3585 streaming migration + 串行 publication | 6.241s | 8.493s | 2.107GB | 477MB | 生成并发布包含 72,784 个 event 的 v2 Session |
+| #3586 preparation-first 调度 | 2.954s | — | 1.026GB | 463MB | 生成相同 v2 Session,并把 publication 延迟到 write open |
+
+Preparation-first restore 比 #3585 快 53%,也比原高性能 reader 快 36%,同时仍然完成 artifact 到 v2 的 migration。
+
+### 调度观测点
+
+| 用户观测点 | 串行 publication | Preparation-first | 变化 |
+|---|---:|---:|---:|
+| Read open + Session restore | 6.241s | 2.954s | -53% |
+| `session.follow` opening snapshot | 7.587s | 2.912s | -62% |
+| Agent 得到 writable Session | 6.246s | 5.161s | -17% |
+| 已是 current v2 的再次打开 | 1.284s | 0.964s | -25% |
+| Follow opening-snapshot peak RSS | 2.353GB | 1.059GB | -55% |
+
+Preparation 中约 2.61 秒用于 Decode 与 migration。延后的 publication 约为 2.56 秒:encode/write/sync 0.83 秒、严格 Worker verification 1.72 秒、source check 与 atomic publication 约 0.005 秒。Read-only 请求完全不执行这段 publication。
+
+Prepared artifact 与 Session restore 的 peak RSS 约为 1.03 GB。Preparation 与 Worker verification 同时存在时峰值约 2.19 GB,因为 parent 保留 logical artifact,而 Worker 独立校验 physical generation。
+
+测试覆盖 shared-waiter cancellation、all-waiter cancellation、memo handoff、read-handle switching、source drift、winner collision、publication idempotence、write-open ordering、Worker failure 与 plain-Node bundled Worker entry。
+
+## 后果
+
+Read-only body access 不发布 generation。第一个 writer 会在 append 前支付一次 publication。已配置的 JSONL root 仍必须可读且结构有效,但 historical body migration 本身不要求写 successor。
+
+Bounded memo 会保留一份 migrated event array,用于连接 read 与 write open。这是有意的取舍:不保留该 artifact 就必须重复 Decode 与 migration,或者无法提前提供 read。
+
+Publication failure 会拒绝 Agent resume 和其他 write open,但不会使已经从 unchanged historical source 交付的 read result 失效。Source drift 对该 write attempt 是 terminal failure,不会触发 hidden state 重算。
+
+Backend 仍存在一个更广泛的既有 lifecycle 缺口:dispose 不拥有每个尚未返回 handle 的 `create()` 或 `open()` operation。本决策不会向 `flush()` 增加 migration-specific tracking,也不解决通用 pending-operation 问题。
+
+## 考虑过的替代方案
+
+- **每个 open 都保持串行 publication**——physical state 最简单,但让 read-only 首屏多等待约 2.2 秒并要求存储可写。
+- **Read 后自动后台 publish**——需要 backend task ownership、shutdown quiescence、error reporting,以及 writer 加入一个没有 caller 请求的任务。
+- **Verify 前返回 writer**——要求 append 写入 unpublished stage,并为已接纳事件增加一种 durability 与 failure state。
+- **每个 caller 独立 preparation**——重复最重的 Decode 与 migration,并在 list/follow/resume 并发时放大峰值内存。
+- **让第一个 caller signal 取消共享工作**——使后续 caller 依赖无关的 cancellation timing。
+- **Source drift 后重跑 migration**——可能替换已经展示给 reader 的历史,也会让一次逻辑 operation 重复处理同一大文件。
+- **Read handle 永远停留在 primed memory**——无法观察后续 append,并偏离普通 persistence refresh 行为。

+ 6 - 0
.agents/notes/implemented/architecture/2026-09-06-embedded-stream-record-readers.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-09-06-embedded-stream-record-readers.md
+2026-09-06-embedded-stream-record-readers.md: 972fee634833cef5fd7b0a54f69780b9370f0cc3
+2026-09-06-embedded-stream-record-readers.zh.md: faba6e179887a9943926f2c73aa8e42a10cee300

+ 50 - 0
.agents/notes/implemented/architecture/2026-09-06-embedded-stream-record-readers.md

@@ -0,0 +1,50 @@
+# Agent Note: Embedded Assistant stream consumers read compact records
+
+Status: implemented
+
+English | [中文](2026-09-06-embedded-stream-record-readers.zh.md)
+
+## Problem
+
+Session format v2 embeds each model attempt's compact stream (`AssistantStreamRecord[]`: packed `text-chunks`, `reasoning-chunks`, and `tool-call-chunks` runs plus timestamped raw `chunk` records) in `assistant/message` and `assistant/attempt`. Consumers that folded those settlements called `expandAssistantStream()` first; it materializes the complete per-member array, so a consumer that needs one fact (`find` on the first token, the last usage chunk, a joined text, one block-end) paid O(members) allocation and time: about two objects per member on top of the compact form.
+
+After v2 embedded streams settlement widened with the message content and Chat and Trajectory sections settled directly from it, the remaining expand consumers are the Host and client folds: Session Stats reads the first-token time per `assistant/attempt` and `assistant/message` (the projection phase of every Session open), the token meter rebuilds provider content and scans every stream for its last usage chunk (the projection unit still scans to the end), the subagent output fold joins plain text, and the Session Controller image lookup scans for block-end chunks.
+
+## Decision
+
+`@deepseek-ai/dsh-llm` answers consumer questions directly from compact records; every remaining consumer folds records once with early exit.
+
+`packages/llm/llm/src/assistant-stream.ts` exports record-level readers beside the accumulator and `expandAssistantStream`:
+
+- Chunk rules: `isTokenDelta` (non-empty text, reasoning, or Tool-call arguments fragment, or any name-bearing Tool-call delta), `isVisibleChunk` (non-whitespace text or reasoning, or a block start or end of any kind other than text, reasoning, or Tool call), and `chunkHasVisibleText` (non-whitespace text delta or completed text block).
+- Run readers: `runFirstTokenTime` and `runFirstVisibleTime` reconstruct the first qualifying member's time from `time0` and the `dt` gaps and stop scanning there; a name-bearing Tool-call run yields `time0` without reading a fragment.
+- Stream readers: `assistantStreamFirstTokenTime`, `assistantStreamHasVisibleContent`, `assistantStreamHasVisibleText`, `lastAssistantStreamChunk(stream, type)` (backward scan), `assistantStreamChunks(stream, type)`, `joinAssistantStreamText`, and `assembleAssistantStream`, which feeds a `BlockAssembler` one joined delta per run (assembly only concatenates, so blocks, usage, finish, and replay state equal the per-member result). `RawStreamChunkType` excludes the delta types, so a raw-chunk lookup can never silently skip packed members.
+
+Session Stats reads `assistantStreamFirstTokenTime`; the token meter reads `lastAssistantStreamChunk(stream, 'usage')` and assembles provider output through `assembleAssistantStream`; the subagent output fold appends `joinAssistantStreamText`; the Session Controller scans `assistantStreamChunks(stream, 'block-end')` for images.
+
+`expandAssistantStream` keeps its strict validation and its remaining callers, which need every member or validate the stream at a durable boundary: Session restore validation, the v1-to-v2 migration validator and publication Worker replay, the reconnect baseline, and test support.
+
+### Measurements
+
+The repo's synthetic first-open benchmark (200 turns, 127,400 released-v0 events, 500,000 streamed deltas in 1,600 compact records; five samples, median):
+
+| Phase | Before | After |
+|---|---|---|
+| first-open projection | 28.0 ms | 5.9 ms |
+| first-open total | 76.9 ms | 53.8 ms |
+| first-open peak RSS | 137.2 MB | 94.6 MB |
+| reopen projection | 17.8 ms | 6.5 ms |
+
+Open, read, and restore phases are unchanged; the reader keeps the same first-token time by construction (the first qualifying member is the first record's first qualifying fragment, and the deltas stay ordered).
+
+## Alternatives considered
+
+**Memoize `expandAssistantStream` per input array.** Expanding all streams once costs tens of milliseconds, but retaining the expansions costs about ten times the compact stream for the event's lifetime — a permanent version of the transient allocation the change removes. The readers remove the need for retained expansions entirely.
+
+**Keep the per-member fold.** Early-exit `.find` still materializes the whole array first, so the allocation and O(members) time remain.
+
+## Consequences
+
+Host and Client folds of an embedded settlement cost O(records) plus one join per run, and no consumer materializes members unless it validates at a durable boundary or needs every member. The token, visibility, and visible-text rules have one home in `dsh-llm`, so a record reader and the accumulator's packing rules cannot drift apart.
+
+Publication verification (`assertCurrentAssistantStreams`) still replays every settlement at publish time; because it must prove content-by-chunk agreement, converting it to run-aware assembly without member materialization remains open work.

+ 50 - 0
.agents/notes/implemented/architecture/2026-09-06-embedded-stream-record-readers.zh.md

@@ -0,0 +1,50 @@
+# Agent Note: 内嵌 Assistant 流的消费方直接读取紧凑记录
+
+Status: implemented
+
+[English](2026-09-06-embedded-stream-record-readers.md) | 中文
+
+## 问题
+
+Session 格式 v2 将每次模型尝试的紧凑流(`AssistantStreamRecord[]`:打包的 `text-chunks`、`reasoning-chunks`、`tool-call-chunks` run 加上带时间戳的原始 `chunk` 记录)嵌入 `assistant/message` 与 `assistant/attempt`。折叠这些 settlement 的消费方会先调用 `expandAssistantStream()`;它会物化完整的逐成员数组,因此只需一个事实的消费方(find 首个 token、最后一个 usage chunk、拼接文本、一个 block-end)也要付出 O(members) 的分配与时间:在紧凑形式之上每个成员约两个对象。
+
+在 v2 内嵌流 settlement 随消息内容扩展、Chat 与 Trajectory 区块直接由内容结算之后,剩余的 expand 消费方是 Host 与客户端折叠:Session Stats 读取每个 `assistant/attempt` 与 `assistant/message` 的首 token 时间(每次打开 Session 的 projection 阶段)、token 计量重建提供商内容并扫描每个流到最后一个 usage chunk(projection 单元仍扫描到末尾)、子代理输出折叠拼接纯文本、Session Controller 镜像查找扫描 block-end chunk。
+
+## 决策
+
+`@deepseek-ai/dsh-llm` 直接从紧凑记录回答消费方问题;剩余消费方对记录做一次带提前退出的折叠。
+
+`packages/llm/llm/src/assistant-stream.ts` 在累加器与 `expandAssistantStream` 之外导出记录级读取器:
+
+- Chunk 规则:`isTokenDelta`(非空文本、reasoning 或 Tool-call 参数片段,或任何带名称的 Tool-call delta)、`isVisibleChunk`(非空白文本或 reasoning,或 text/reasoning/Tool call 之外的任意块开始或结束)、`chunkHasVisibleText`(非空白文本 delta 或完成的文本块)。
+- Run 读取器:`runFirstTokenTime` 与 `runFirstVisibleTime` 从 `time0` 与 `dt` 间隔重建首个合格成员的时间并停止扫描;带名称的 Tool-call run 直接产出 `time0`,不读片段。
+- 流读取器:`assistantStreamFirstTokenTime`、`assistantStreamHasVisibleContent`、`assistantStreamHasVisibleText`、`lastAssistantStreamChunk(stream, type)`(逆向扫描)、`assistantStreamChunks(stream, type)`、`joinAssistantStreamText` 与 `assembleAssistantStream`(每个 run 向 `BlockAssembler` 喂入一个拼接后的 delta;组装只做拼接,因此 blocks、usage、finish 与 replay state 与逐成员结果一致)。`RawStreamChunkType` 排除 delta 类型,因此原始 chunk 查找不可能静默跳过打包成员。
+
+Session Stats 读取 `assistantStreamFirstTokenTime`;token 计量读取 `lastAssistantStreamChunk(stream, 'usage')` 并通过 `assembleAssistantStream` 组装提供商输出;子代理输出折叠追加 `joinAssistantStreamText`;Session Controller 用 `assistantStreamChunks(stream, 'block-end')` 扫描镜像。
+
+`expandAssistantStream` 保留其严格校验与其余调用方(需要每个成员或在持久边界校验流):Session 恢复校验、v1-to-v2 迁移校验器与发布 Worker 重放、重连基线、测试支撑。
+
+### 测量
+
+仓库的合成 first-open 基准(200 循环、127,400 个 released-v0 事件、1,600 条紧凑记录中的 500,000 个流式 delta;五次采样取中位数):
+
+| 阶段 | 之前 | 之后 |
+|---|---|---|
+| first-open projection | 28.0 ms | 5.9 ms |
+| first-open 总计 | 76.9 ms | 53.8 ms |
+| first-open 峰值 RSS | 137.2 MB | 94.6 MB |
+| reopen projection | 17.8 ms | 6.5 ms |
+
+Open、read、restore 阶段不变;读取器按构造保持相同的首 token 时间(首个合格成员即首条记录的首个合格片段,且 delta 保持有序)。
+
+## 备选方案
+
+**按输入数组记忆化 `expandAssistantStream`。** 展开全部流只需几十毫秒,但保留展开结果在事件生命周期内约花费紧凑流的十倍内存——这是本变更移除的瞬时分配的永久版本。读取器完全消除了对保留展开的需求。
+
+**保留逐成员折叠。** 提前退出的 `.find` 仍然先物化整个数组,因此分配与 O(members) 时间仍在。
+
+## 后果
+
+Host 与客户端折叠一次内嵌结算的代价为 O(records) 加每个 run 一次拼接,且除非在持久边界校验或需要每个成员,消费方不再物化成员。token、可见性与可见文本规则在 `dsh-llm` 中只有一处,因此记录读取器与累加器的打包规则不可能漂移。
+
+发布校验(`assertCurrentAssistantStreams`)仍在发布时重放每个 settlement;因为它必须按 chunk 证明内容一致,将其转为不入成员的 run 感知组装仍是未完成工作。

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-08-31-win32-picker-path-string-read.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-08-31-win32-picker-path-string-read.md
+2026-08-31-win32-picker-path-string-read.md: 63b2a2ea5ee88b5fa3ac79aca98350f58067fd1e
+2026-08-31-win32-picker-path-string-read.zh.md: 8afc6610a0cf1c7c8649f9e1c62c26b3d89f5a03

+ 25 - 0
.agents/notes/implemented/bug-fix/2026-08-31-win32-picker-path-string-read.md

@@ -0,0 +1,25 @@
+# Agent Note: Read the Win32 picker path without a fixed-size unmanaged view
+
+Status: implemented
+
+English | [中文](2026-08-31-win32-picker-path-string-read.zh.md)
+
+## Problem
+
+The Win32 picker needs to decode a NUL-terminated UTF-16 string allocated by `IShellItem::GetDisplayName` and release it through `CoTaskMemFree`. A fixed-length external ArrayBuffer adds a runtime requirement and manual terminator scanning without providing the allocation size.
+
+## Decision
+
+`readUtf16` stores the native address in a pointer-width buffer and passes it to generic `koffi.decode(buffer, 'str16')`. Generic decoding expects a pointer variable, not the string address directly. The slice follows `koffi.sizeof('void *')`; Koffi 3 represents native addresses as BigInt. The allocation must remain valid and NUL-terminated during decoding. Successful conversion leaves the original address available for `CoTaskMemFree`; if decoding throws, the string is not freed.
+
+## Alternatives considered
+
+**External view and manual scan.** This requires external-buffer support and duplicates Koffi's string conversion. Neither a fixed view nor growing chunks establish the native allocation size.
+
+**String-typed out-param.** `_Out_ str16 *` returns text but discards the pointer needed for explicit COM release. Built-in `str16!` disposal uses the CRT allocator rather than the COM allocator.
+
+**Custom disposable type.** A Koffi disposable can call `CoTaskMemFree`, but explicit conversion keeps the address and release at one call site without registering a native type.
+
+## Consequences
+
+Real-Koffi tests exercise the production result-path conversion over live UTF-16 buffers, including U+5F00, surrogate pairs, NUL termination and strings exceeding 32 KiB. Separate four- and eight-byte BigInt cases verify pointer preservation and release of the original address. Test-owned buffers stay live through the synchronous read; pointer bytes are checked before native dereferencing. The earlier scanning decision remains in the [archived note](../../archived/bug-fix/2026-08-23-win32-utf16-nul-truncation.md).

+ 25 - 0
.agents/notes/implemented/bug-fix/2026-08-31-win32-picker-path-string-read.zh.md

@@ -0,0 +1,25 @@
+# Agent Note: 不用固定长度的非托管视图读取 Win32 选择器路径
+
+Status: implemented
+
+[English](2026-08-31-win32-picker-path-string-read.md) | 中文
+
+## Problem
+
+Win32 选择器需要解码 `IShellItem::GetDisplayName` 分配的 NUL 结尾 UTF-16 字符串,并通过 `CoTaskMemFree` 释放它。固定长度的外部 ArrayBuffer 增加了运行时要求和手工终止符扫描,却不能提供实际分配大小。
+
+## Decision
+
+`readUtf16` 将原生地址存入指针宽度的缓冲区,再交给通用 `koffi.decode(buffer, 'str16')`。通用解码需要指针变量,而非直接传入字符串地址。切片长度取自 `koffi.sizeof('void *')`;Koffi 3 用 BigInt 表示原生地址。解码期间分配必须保持有效且以 NUL 结尾。转换成功后,原始地址仍可交给 `CoTaskMemFree`;若解码抛错,字符串不会被释放。
+
+## Alternatives considered
+
+**外部视图加手工扫描。** 这要求运行时支持外部缓冲区,并重复实现 Koffi 的字符串转换。固定视图和递增分块都无法确定原生分配大小。
+
+**字符串类型出参。** `_Out_ str16 *` 返回文本,却丢失显式 COM 释放所需的指针。内置 `str16!` 使用 CRT 分配器释放,而非 COM 分配器。
+
+**自定义可释放类型。** Koffi 可释放类型可以调用 `CoTaskMemFree`,但显式转换无需注册原生类型,就能将地址和释放保留在同一调用点。
+
+## Consequences
+
+真实 Koffi 测试使用存活的 UTF-16 缓冲区执行生产结果路径转换,涵盖 U+5F00、代理对、NUL 终止和超过 32 KiB 的字符串。独立的四字节与八字节 BigInt 用例验证指针保持完整,并释放原始地址。测试持有的缓冲区在同步读取期间保持存活;原生解引用之前会检查指针字节。此前的扫描决策保留在[归档记录](../../archived/bug-fix/2026-08-23-win32-utf16-nul-truncation.md)中。

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-09-05-nested-terminal-cards.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-05-nested-terminal-cards.md
+2026-09-05-nested-terminal-cards.md: b369aef20b0b5fbb338e40affd5fd2800b525bcb
+2026-09-05-nested-terminal-cards.zh.md: bf97fffa2a8d92562e982d498d17b830e1e25f48

+ 35 - 0
.agents/notes/implemented/bug-fix/2026-09-05-nested-terminal-cards.md

@@ -0,0 +1,35 @@
+# Agent Note: Nested terminal cards
+
+Status: implemented
+
+English | [中文](2026-09-05-nested-terminal-cards.zh.md)
+
+## Problem
+
+A shell command dispatched through `run_code` carries the arguments and rendered output needed for a terminal card, but rejecting every block with `parentCallId` hides that presentation solely because the call is nested. The rejection also affects running prompts and selected-child Details.
+
+## Decision
+
+`terminalCardModel` applies the same eligibility checks to root and Code Dispatch calls, without rejecting `parentCallId`. Supported running and settled `bash`, `pwsh`, and `terminal_send` calls use the existing terminal card. Background calls, tool errors, malformed inputs, missing call heads, and unsupported result content retain generic fallback. Persistent shells remain eligible while running and generic when settled; a nonzero process exit remains terminal result data rather than a tool error.
+
+This partially supersedes only the terminal child-card prohibition in [Client-derived tool presentation](../architecture/2026-08-23-client-derived-tool-presentation.md). That note remains active for Client presentation ownership and the diff/read/search/web child restrictions. No Host presenter, event, schema, metadata, call-tree, or model-context change is required. The metadata and execution-local value decisions in [canonical tool output](../architecture/2026-07-20-canonical-tool-output-contract.md) and [PTC typed returns](../feature/2026-07-20-ptc-typed-tool-returns.md) remain intact; metadata omission does not prohibit Client-derived terminal cards.
+
+Shell output ending in a recognized spill-policy notice uses generic output: expandable in `BashRow`, raw fallback in Details. The notice can follow or replace the exit marker, so its absence at the end does not justify a successful terminal status. The browser-safe `@deepseek-ai/dsh-spill-policy/notice` entry owns the text convention: the producer calls `formatSpillNotice(omitted, ref)` and the Client calls `hasSpillNotice(text)`. Both share delimiters, and omission validation reuses `describeOmitted` rather than duplicating its prose. The formatter preserves the persisted spelling byte-for-byte; existing Session result bytes stay untouched, with no Session format change or migration.
+
+## Alternatives considered
+
+**Keep the blanket nested-call rejection.** Rejected because nesting does not remove the raw facts the terminal model already consumes. It hides usable shell output while the same call renders as a terminal at the root.
+
+**Enable every nested structured card.** Rejected because other card models have independent metadata requirements and child restrictions. This fix changes only terminal eligibility.
+
+**Parse exit markers around spill suffixes.** Rejected because truncation can remove the real status; conservative generic output avoids guessing success from an incomplete result.
+
+**Maintain a separate UI notice regex.** Rejected because it duplicates the producer's text convention and can drift from persisted output. The shared browser-safe owner keeps formatting and recognition together without loading the Host plugin in the browser.
+
+## Consequences
+
+Rows and Details share terminal derivation for nested calls without a second renderer or presentation hint. Generic fallback and settled-persistent behavior remain separate from terminal-card eligibility. The parent-child relationship still controls tree placement, not terminal rendering. Text recognition cannot authenticate output: a tool can print the same notice. A match selects conservative generic presentation, not proof of spill provenance or process status.
+
+## Verification
+
+The [terminal card specs](../../../../packages/client/ui-tool/tests/terminal-card.client.spec.tsx) cover root/child eligibility, running and settled Details, and fallback cases. The [assembled Code Dispatch specs](../../../../packages/client/ui-tool/tests/chat-code-subcalls.client.spec.tsx) cover nested terminal rendering through the conversation tree. The [notice specs](../../../../packages/spill/spill-policy/tests/notice.spec.ts) pin the historical spelling with a literal fixture independent of the formatter. The [spill-policy-to-UI specs](../../../../packages/client/ui-tool/tests/spill-policy-terminal.client.spec.ts) exercise actual root and PTC spill production, unchanged full text and programmatic values, byte caps, notice-only output, and terminal fallback. Browser replay owns the visible nested-card change; nonterminal child behavior remains outside this fix.

+ 35 - 0
.agents/notes/implemented/bug-fix/2026-09-05-nested-terminal-cards.zh.md

@@ -0,0 +1,35 @@
+# Agent Note: 嵌套 terminal 卡片
+
+Status: implemented
+
+[English](2026-09-05-nested-terminal-cards.md) | 中文
+
+## 问题
+
+经 `run_code` 分发的 shell 命令携带 terminal 卡片所需的参数与渲染输出,但对所有带有 `parentCallId` 的块一律拒绝,会仅因调用嵌套而隐藏这类展示。该拒绝也影响运行中的命令提示行与选中子调用的 Details。
+
+## 决策
+
+`terminalCardModel` 对根调用与 Code Dispatch 调用应用相同的适用检查,不因 `parentCallId` 拒绝调用。受支持的运行中与已完成的 `bash`、`pwsh` 和 `terminal_send` 调用使用现有 terminal 卡片。后台调用、工具错误、格式错误的输入、缺失的调用头和不受支持的结果内容保留通用回退。持久 shell 在运行中仍可使用 terminal,完成后使用通用展示;非零进程退出仍是 terminal 结果数据,而非工具错误。
+
+本文仅部分取代 [Client 派生工具展示](../architecture/2026-08-23-client-derived-tool-presentation.zh.md)中的 terminal 子调用卡片禁令。该文继续负责 Client 展示所有权及 diff/read/search/web 子调用限制。无需更改 Host 展示转换器、事件、schema、元数据、调用树或模型上下文。[规范工具输出](../architecture/2026-07-20-canonical-tool-output-contract.zh.md)与 [PTC 类型化返回值](../feature/2026-07-20-ptc-typed-tool-returns.zh.md)中的元数据和执行期值决策保持不变;省略元数据不禁止 Client 派生 terminal 卡片。
+
+以已识别的 spill 策略提示结尾的 shell 输出使用通用展示:在 `BashRow` 中可展开,在 Details 中使用原始回退。提示可能位于退出标记之后或取代它,因此末尾缺少退出标记不能作为 terminal 成功状态的依据。浏览器安全入口 `@deepseek-ai/dsh-spill-policy/notice` 负责文本约定:生产方调用 `formatSpillNotice(omitted, ref)`,Client 调用 `hasSpillNotice(text)`。两者共用分隔符,省略信息校验复用 `describeOmitted`,不复制其文案。格式化函数逐字节保留持久化拼写;现有 Session 结果字节保持不变,不更改 Session 格式,也不执行迁移。
+
+## 考虑过的替代方案
+
+**保留对嵌套调用的一律拒绝。** 不予采用,因为嵌套不会移除 terminal model 已消费的原始事实。这会隐藏可用的 shell 输出,而同一调用位于根时却可渲染为 terminal。
+
+**启用所有嵌套结构化卡片。** 不予采用,因为其他 card model 有独立的元数据要求与子调用限制。本修复只改变 terminal 适用性。
+
+**解析 spill 后缀附近的退出标记。** 不予采用,因为截断可能移除真实状态;保守的通用输出避免从不完整结果猜测成功。
+
+**维护独立的 UI 通知正则表达式。** 不予采用,因为它复制生产方的文本约定,可能与持久化输出偏离。共享的浏览器安全模块统一负责格式化与识别,无需在浏览器中加载 Host 插件。
+
+## 后果
+
+行与 Details 对嵌套调用共享 terminal 派生,不增加第二个渲染器或展示提示字段。通用回退与已完成持久 shell 的行为仍独立于 terminal 卡片适用性。父子关系仍控制树中的位置,而非 terminal 渲染。文本识别无法认证输出来源:工具也能打印相同的提示。匹配结果只选择保守的通用展示,不能证明 spill 来源或进程状态。
+
+## 验证
+
+[Terminal 卡片测试](../../../../packages/client/ui-tool/tests/terminal-card.client.spec.tsx)覆盖根/子调用适用性、运行中与已完成的 Details 以及回退情况。[组装后的 Code Dispatch 测试](../../../../packages/client/ui-tool/tests/chat-code-subcalls.client.spec.tsx)覆盖经对话树渲染的嵌套 terminal。[通知测试](../../../../packages/spill/spill-policy/tests/notice.spec.ts)使用独立于格式化函数的字面量 fixture(测试前置数据)固定历史拼写。[spill-policy 到 UI 的测试](../../../../packages/client/ui-tool/tests/spill-policy-terminal.client.spec.ts)覆盖真实的根调用与 PTC spill 生成、保持不变的完整文本和程序化值、字节上限、仅含通知的输出以及 terminal 回退。浏览器回放负责验证可见的嵌套卡片变化;非 terminal 子调用行为不属于本修复。

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-09-05-pi-ai-upgrade-compatibility.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-05-pi-ai-upgrade-compatibility.md
+2026-09-05-pi-ai-upgrade-compatibility.md: ed400ab62d221e1b58025617f05500590bfed5e7
+2026-09-05-pi-ai-upgrade-compatibility.zh.md: 408669ddc963a505254529acaab4c19e9774c6b1

+ 25 - 0
.agents/notes/implemented/bug-fix/2026-09-05-pi-ai-upgrade-compatibility.md

@@ -0,0 +1,25 @@
+# Agent Note: pi-ai upgrade compatibility
+
+Status: implemented
+
+English | [中文](2026-09-05-pi-ai-upgrade-compatibility.zh.md)
+
+## Problem
+
+The pi-ai adapter classifies upstream compatibility fields explicitly and persists only replay metadata needed by later requests. An SDK upgrade can add fields to either set without changing the Harness provider-neutral API. Unclassified configuration fields fail compilation; omitted replay metadata can silently change subsequent provider requests.
+
+## Decision
+
+The adapter follows [pi-ai 0.85.1](https://github.com/earendil-works/pi/blob/v0.85.1/packages/ai/CHANGELOG.md). `thinkingTokenBudgetField`, `vllmPriority`, and `supportsMaxOutputTokens` are opt-in gateway controls; `thinking.budget` joins the existing template placeholders. The SDK owns budget resolution and serialization. `supportsMidConvoEffort` and `allowedFallbackModels` remain catalog-owned because their correctness depends on exact Anthropic transports, model capabilities, and fallback pricing.
+
+Optional `providerThinkingLevel` remains in the adapter replay-v2 response metadata so Anthropic history retains its provider-native effort. Absence remains valid; neither the replay version nor the released Session format changes. Replay provenance retains the requested model while `responseModel` retains an Anthropic alias resolution or fallback. Reconstruction restores that native model so pi-ai still applies its cross-model signature rules. The provider-neutral LLM API stays unchanged; the [provider-routed replay ownership rules](../architecture/2026-07-14-provider-routed-llm-adapters.md) still apply. The 0.84.2 Anthropic adapter [initializes `model` from the request](https://github.com/earendil-works/pi/blob/v0.84.2/packages/ai/src/api/anthropic-messages.ts#L510-L515) and [records only response id and usage at message start](https://github.com/earendil-works/pi/blob/v0.84.2/packages/ai/src/api/anthropic-messages.ts#L589-L605). It never writes `responseModel`, so its replay records retain the requested model without native-model metadata.
+
+## Alternatives considered
+
+**Withhold every new field.** This would misclassify deployment-owned gateway controls as catalog facts: upstream explicitly leaves budget-field selection and vLLM priority out of its generated catalog.
+
+**Expose every new field.** This would let arbitrary gateways claim model-specific Anthropic effort and fallback support without the catalog evidence that makes those features valid.
+
+## Consequences
+
+Compile-time coverage retains explicit field classification. [Compatibility tests](../../../../packages/llm/llm-pi-ai/tests/compat-upgrade.spec.ts) cover schema acceptance, invalid values, protocol applicability, and materialization without changing defaults. [Replay conversion tests](../../../../packages/llm/llm-pi-ai/tests/convert.spec.ts) cover optional effort preservation. Mixed-protocol catalog tests use the installed OpenCode catalog. Provider behavior remains upstream-owned; live provider verification is separate from keyless adapter tests.

+ 25 - 0
.agents/notes/implemented/bug-fix/2026-09-05-pi-ai-upgrade-compatibility.zh.md

@@ -0,0 +1,25 @@
+# Agent Note: pi-ai 升级兼容性
+
+Status: implemented
+
+[English](2026-09-05-pi-ai-upgrade-compatibility.md) | 中文
+
+## Problem
+
+pi-ai 适配器显式分类上游兼容字段,并且只持久化后续请求需要的回放元数据。SDK 升级可能在不改变 Harness 提供方无关 API 的情况下为任一集合新增字段。未分类的配置字段会导致编译失败;遗漏回放元数据则可能静默改变后续提供方请求。
+
+## Decision
+
+适配器遵循 [pi-ai 0.85.1](https://github.com/earendil-works/pi/blob/v0.85.1/packages/ai/CHANGELOG.md)。`thinkingTokenBudgetField`、`vllmPriority` 和 `supportsMaxOutputTokens` 是显式启用的网关控制;`thinking.budget` 加入现有模板占位符。SDK 拥有预算解析和序列化。`supportsMidConvoEffort` 和 `allowedFallbackModels` 仍由目录拥有,因为其正确性依赖确切的 Anthropic 传输、模型能力和回退定价。
+
+可选的 `providerThinkingLevel` 保存在适配器 replay-v2 响应元数据中,让 Anthropic 历史保留提供方原生 effort。缺失仍然有效;回放版本与已发布 Session 格式均不改变。回放来源保留请求模型,`responseModel` 则保留 Anthropic 别名解析或回退后的模型。重建会恢复该原生模型,让 pi-ai 继续应用其跨模型签名规则。提供方无关的 LLM API 保持不变;[按提供方路由的回放归属规则](../architecture/2026-07-14-provider-routed-llm-adapters.zh.md) 仍然适用。0.84.2 Anthropic 适配器[从请求初始化 `model`](https://github.com/earendil-works/pi/blob/v0.84.2/packages/ai/src/api/anthropic-messages.ts#L510-L515),并且[在消息开始时只记录响应 ID 和用量](https://github.com/earendil-works/pi/blob/v0.84.2/packages/ai/src/api/anthropic-messages.ts#L589-L605)。它从不写入 `responseModel`,因此其回放记录保留请求模型,不携带原生模型元数据。
+
+## Alternatives considered
+
+**保留所有新增字段为不可配置。** 这会把部署拥有的网关控制误分类为目录事实:上游明确不在生成目录中设置预算字段选择和 vLLM 优先级。
+
+**开放所有新增字段。** 这会允许任意网关在缺少目录证据时声明支持模型专属的 Anthropic effort 与回退能力,而这些证据正是相关功能有效的依据。
+
+## Consequences
+
+编译期覆盖保留显式字段分类。[兼容性测试](../../../../packages/llm/llm-pi-ai/tests/compat-upgrade.spec.ts) 覆盖 schema 接受、无效值、协议适用范围及物化,且不改变默认值。[回放转换测试](../../../../packages/llm/llm-pi-ai/tests/convert.spec.ts) 覆盖可选 effort 的保留。混合协议目录测试使用已安装的 OpenCode 目录。提供方行为仍由上游拥有;真实提供方验证与无需密钥的适配器测试分开。

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-09-05-session-reference-model-budget.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-05-session-reference-model-budget.md
+2026-09-05-session-reference-model-budget.md: 0654f2b89187727668fddb97ab6e982bacd82dc1
+2026-09-05-session-reference-model-budget.zh.md: 84f4fee76c7fb54f3ec37f0ae3417cc56d772af7

+ 25 - 0
.agents/notes/implemented/bug-fix/2026-09-05-session-reference-model-budget.md

@@ -0,0 +1,25 @@
+# Agent Note: Model-relative session-reference budgets
+
+Status: implemented
+
+English | [中文](2026-09-05-session-reference-model-budget.zh.md)
+
+## Problem
+
+A fixed 64 KiB reference budget discards useful source context on large-context models. The target session header describes a prior request, while agent options seed routing; neither necessarily identifies the model selected for the entering step.
+
+## Decision
+
+[Session-reference](../../../../packages/context/session-reference/README.md) observes the completed `system-prompt/assemble` waterfall with a local prepend listener and stores its provider/model pair in a WeakMap keyed by Agent. Preparation resolves that route through the optional LLM service; direct preparation before any assembly uses agent options. Diagnostics without an Agent do not update the map.
+
+Each source receives `max(65536, floor(contextWindow × 4 × referenceContextFraction))` bytes, with a default fraction of `0.2`. Four bytes per token is a sizing heuristic. Explicit `maxReferenceBytes` bypasses model lookup and remains exact. Missing route, service, adapter, or capacity retains the floor; other lookup failures and cancellation propagate. An absent adapter does not prevent stream middleware from serving the route.
+
+## Alternatives considered
+
+**Read the header or options for every step.** Either can select a stale model after a live switch. The completed assembly exposes the route captured by model selection.
+
+**Reassemble or redispatch request routing during pre-step.** These operations repeat plugin effects and can capture a different selection. A local observer needs neither loop changes nor another public routing API.
+
+## Consequences
+
+The budget grows with model capacity without changing projection, retention, or preview policy. It remains per source, not an aggregate token reservation. The listener is effect-owned and disposable; the map does not retain agents. Focused tests cover the floor, fractional conversion, explicit overrides, live selection, absent metadata, cancellation, lookup errors, and listener removal.

+ 25 - 0
.agents/notes/implemented/bug-fix/2026-09-05-session-reference-model-budget.zh.md

@@ -0,0 +1,25 @@
+# Agent Note: 模型相对会话引用预算
+
+Status: implemented
+
+[English](2026-09-05-session-reference-model-budget.md) | 中文
+
+## Problem
+
+固定的 64 KiB 引用预算会在大上下文模型上丢弃有用的来源上下文。目标会话头描述上一次请求,而 agent options 为路由提供初始值;两者都不一定标识当前进入步骤所选的模型。
+
+## Decision
+
+[Session-reference](../../../../packages/context/session-reference/README.zh.md) 通过本地 prepend 监听器观察已完成的 `system-prompt/assemble` 瀑布,并把 provider/model 对存入以 Agent 为键的 WeakMap。准备阶段通过可选 LLM 服务解析该路由;首次组装前直接准备则使用 agent options。不带 Agent 的诊断不会更新映射。
+
+每个来源获得 `max(65536, floor(contextWindow × 4 × referenceContextFraction))` 字节,默认比例为 `0.2`。每个 token 四字节是容量估算。显式 `maxReferenceBytes` 跳过模型查询并保持精确值。缺少路由、服务、适配器或容量时保留下限;其他查询失败和取消会传播。缺少适配器不妨碍流中间件处理该路由。
+
+## Alternatives considered
+
+**每步读取会话头或 options。** 实时切换后,两者都可能选中旧模型。完成的组装公开模型选择所捕获的路由。
+
+**在 pre-step 中重新组装或重新分派请求路由。** 这些操作会重复插件效果,并可能捕获不同的选择。本地观察器不需要修改循环或增加公共路由 API。
+
+## Consequences
+
+预算随模型容量增长,不改变投影、保留或预览策略。它仍按来源计算,而不是聚合 token 预留。监听器由 effect 持有并可释放;映射不会保留 agent。聚焦测试覆盖下限、比例换算、显式覆盖、实时选择、元数据缺失、取消、查询错误与监听器移除。

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-09-05-session-reference-spill-reuse.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-05-session-reference-spill-reuse.md
+2026-09-05-session-reference-spill-reuse.md: d9f9e3702075417a9b5c742831737e412e75bd00
+2026-09-05-session-reference-spill-reuse.zh.md: 346ed1a39ee65a560cf9e74e519ae45c2ce133e1

+ 41 - 0
.agents/notes/implemented/bug-fix/2026-09-05-session-reference-spill-reuse.md

@@ -0,0 +1,41 @@
+# Agent Note: Reuse spill storage for truncated session references
+
+Status: implemented
+
+English | [中文](2026-09-05-session-reference-spill-reuse.zh.md)
+
+## Problem
+
+A bounded cross-session preview can omit whole messages or most of a retained message. A model that sees only the preview needs an accurate account of the omission and a way to inspect the captured text, without treating another session's instructions as current authority. Rereading the source later would not recover the same observation when the source advances or compacts.
+
+## Decision
+
+[Session-reference preparation](../../../../packages/context/session-reference/README.md) retains its existing preview policy and per-reference JSON byte budget. Each truncated reference attempts `saveText` through optional `ctx.get("spillStore")`; an untruncated reference writes no artifact. The full transcript and bounded preview derive from the same captured user/assistant text projection, including compaction checkpoints but excluding tools, reasoning, and other injected context. No second source read occurs.
+
+The artifact belongs to the target session receiving the context. Its descriptive source is `{ kind: "session-reference", sessionId, label }`, where `sessionId` identifies the referenced session. [Spill storage](../../../../packages/spill/spill/README.md) accepts this minimal alternative alongside the existing tool source; it requires no fabricated tool name or call id. Storage ownership does not authorize retrieval.
+
+A separate omission notice outside the bounded preview JSON records exact `omittedMessages` and `omittedBytes`. It carries the saved locator and backend `retrievalHint`, or an unavailable outcome distinguishing missing storage from a failed save. This notice is model-visible content in the same durable reference message, not metadata-only UI decoration. A tiny preview budget cannot remove it. The saved transcript carries capture metadata, including `capturedFormatVersion`, and the same untrusted-background warning as the preview. Per-message JSON string fragments contain at most 64 Unicode code points per line; decoding and concatenating them restores exact text, including original newlines. This fixed artifact format keeps long single-line middles retrievable with ordinary paged file reads without changing preview retention.
+
+Cancellation after an asynchronous save prevents context publication, even if storage already created the artifact. The consumer does not add rollback or deletion APIs; the existing backend expiry policy governs that artifact. Replay uses the logged preview and notice and never repeats the save or source read.
+
+## Alternatives considered
+
+**Write a separate session-reference file store.** Rejected because private naming, session-scoped ownership, locator guidance, and artifact lifetime already belong to spill storage. A second store would duplicate those policies.
+
+**Reread the source when saving or retrieving.** Rejected because source mutation could make the artifact disagree with the preview and its captured sequence. Saving the original projection preserves the observation.
+
+**Put omission and retrieval data inside the bounded preview JSON.** Rejected because that spends the conversation budget on metadata and can hide the notice precisely when the budget is smallest. Separate durable model-visible text preserves both obligations.
+
+**Use tool provenance for every spill.** Rejected because a session reference has no model-issued tool call. Invented tool ids would misattribute the artifact rather than describe its producer.
+
+## Consequences
+
+The model can inspect text omitted from a preview without increasing the preview budget. Notices add request tokens outside that budget, and retrieval adds the requested transcript text later. Storage is best-effort: an unavailable notice is honest about loss of retrieval while the bounded preview remains usable. A saved locator can expire even while its notice remains in durable history; this feature does not promise permanent archival or recover content already removed by source compaction.
+
+## Verification
+
+The [unit suite](../../../../packages/context/session-reference/tests/session-reference.spec.ts) pins omission counts, full Unicode and control-character recovery, whole-message drops, three-reference isolation, missing and failed storage, source exclusions and mutation isolation, and cancellation before publication. The [Loader composition test](../../../../packages/context/session-reference/tests/loader-composition.spec.ts) exercises the real local store and paged `read` tool against the middle of a giant single-line message, with target-session storage ownership. The [keyless recorded-session scenario](../../../../snapshots/session/session-reference-spill/snapshot.yml) pins the durable model-visible reference context. Nested Windows-locator regressions cover both serialized extraction and normalization without rewriting unrelated backslashes. Replay [normalizes known quoted spill locators](../../../../packages/test-support/session-snapshot/README.md) while preserving saved byte lengths and omission counts.
+
+## Related decisions
+
+The [tool-output spill decision](../architecture/2026-07-08-tool-output-spill-files.md) remains active: its storage/policy separation, failure degradation, provider caps, and retrieval alternatives still constrain tool consumers. This note extends its producer vocabulary without replacing that rationale. [Separate context injection from turn execution](../architecture/2026-07-24-separate-context-injection-from-turn-execution.md) remains the authority for durable message admission, and [producer-declared context forms](../feature/2026-08-05-context-form-vocabulary.md) remains the authority for recall presentation.

+ 41 - 0
.agents/notes/implemented/bug-fix/2026-09-05-session-reference-spill-reuse.zh.md

@@ -0,0 +1,41 @@
+# Agent Note: 为截断的会话引用复用 spill 存储
+
+Status: implemented
+
+[English](2026-09-05-session-reference-spill-reuse.md) | 中文
+
+## 问题
+
+有界的跨会话预览可能省略整条消息,也可能省略保留消息中的大部分文本。只看到预览的模型需要准确了解省略情况,并能检查已捕获的文本,同时不能把其他会话的指令视为当前授权。源会话继续推进或发生压缩后,再次读取无法恢复同一次观察。
+
+## 决策
+
+[会话引用准备](../../../../packages/context/session-reference/README.zh.md)保留既有预览策略和逐引用 JSON 字节预算。每个被截断的引用通过可选的 `ctx.get("spillStore")` 尝试 `saveText`;未截断的引用不写入产物。完整转录与有界预览来自同一份已捕获的 user/assistant 文本投影,包含压缩检查点,但排除工具、推理与其他注入上下文。不发生第二次源读取。
+
+产物归接收上下文的目标会话所有。其描述性来源是 `{ kind: "session-reference", sessionId, label }`,其中 `sessionId` 标识被引用的会话。[spill 存储](../../../../packages/spill/spill/README.zh.md)在工具来源之外接受这一最小分支;不需要伪造工具名称或调用 id。存储归属不授权取回。
+
+有界预览 JSON 之外的独立省略通知记录精确的 `omittedMessages` 与 `omittedBytes`。通知携带保存后的定位信息和后端 `retrievalHint`,或区分未配置存储与保存失败的不可用结果。该通知是同一条持久引用消息中的模型可见内容,而不是只供 UI 使用的元数据装饰。极小的预览预算无法移除它。保存的转录携带包括 `capturedFormatVersion` 在内的捕获元数据,以及与预览相同的不受信任背景警告。每条消息的 JSON 字符串片段每行至多包含 64 个 Unicode 码点;解码并拼接后可恢复精确文本,包括原始换行。这种固定产物格式让普通分页文件读取可以取回很长的单行文本中部,而不改变预览保留策略。
+
+异步保存后的取消会阻止上下文发布,即使存储已经创建了产物。消费方不增加回滚或删除 API;该产物遵循后端既有过期策略。回放使用已记录的预览与通知,不会重复保存或源读取。
+
+## 考虑过的替代方案
+
+**另写一个会话引用文件存储。** 不予采纳,因为私有命名、会话级归属、定位指引与产物生命周期已经由 spill 存储负责。第二套存储会重复这些策略。
+
+**保存或取回时重新读取源。** 不予采纳,因为源变更可能使产物与预览及其捕获序列不一致。保存原始投影可以保留该次观察。
+
+**把省略与取回数据放入有界预览 JSON。** 不予采纳,因为这会让元数据占用对话预算,并可能在预算最小时恰好隐藏通知。独立的持久模型可见文本同时保留两项保证。
+
+**所有 spill 都使用工具来源。** 不予采纳,因为会话引用没有模型发出的工具调用。虚构工具 id 会错误归属产物,而不是描述其生产者。
+
+## 后果
+
+模型可以检查预览省略的文本,而无需增加预览预算。通知在该预算之外增加请求 token,之后的取回再添加所请求的转录文本。存储采用尽力而为策略:不可用通知如实说明无法取回,而有界预览仍可使用。即使通知仍在持久历史中,保存的定位信息也可能过期;此功能不承诺永久归档,也无法恢复源压缩已经移除的内容。
+
+## 验证
+
+[单元测试](../../../../packages/context/session-reference/tests/session-reference.spec.ts)锁定省略计数、完整 Unicode 与控制字符恢复、整条消息丢弃、三个引用的隔离、无存储与保存失败、来源排除与变更隔离,以及发布前取消。[Loader 组合测试](../../../../packages/context/session-reference/tests/loader-composition.spec.ts)使用真实本地存储和分页 `read` 工具,读取巨型单行消息的中部,并检查存储归目标会话所有。[无密钥录制会话场景](../../../../snapshots/session/session-reference-spill/snapshot.yml)锁定持久的模型可见引用上下文。嵌套 Windows 定位信息回归覆盖序列化提取与规范化,且不改写无关反斜杠。回放会[规范化已知的带引号 spill 定位信息](../../../../packages/test-support/session-snapshot/README.zh.md),同时保留保存字节数与省略计数。
+
+## 相关决策
+
+[工具输出 spill 决策](../architecture/2026-07-08-tool-output-spill-files.zh.md)保持活跃:其存储/策略分离、失败降级、提供方上限与取回替代方案仍约束工具消费方。本说明扩展其生产者词汇,而不替代这些理由。[分离上下文注入与轮次执行](../architecture/2026-07-24-separate-context-injection-from-turn-execution.zh.md)仍负责持久消息准入,[生产者声明的上下文形式](../feature/2026-08-05-context-form-vocabulary.zh.md)仍负责 recall 展示。

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md
+2026-09-06-windows-python-console-spawn-wait.md: 92443bcf8a6e4e5609dc469efa4ebd1d82ab127f
+2026-09-06-windows-python-console-spawn-wait.zh.md: dba2f324b29955580fc11e7cea7a0525a8bc8c86

+ 27 - 0
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md

@@ -0,0 +1,27 @@
+# Agent Note: Wait for the Windows Python console runtime
+
+Status: implemented
+
+English | [中文](2026-09-06-windows-python-console-spawn-wait.zh.md)
+
+## Problem
+
+The installed Python `dsh.exe` console command intermittently exits with Windows access violation `0xc0000005` before initializing a profile. Its smoke assertion omitted the process status and reported only empty streams. A [native faulthandler probe](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34030851888) captures the fault in Python 3.10 `os._execvpe`, called by the runtime console entry, rather than in the bundled Node executable. Direct executable controls pass.
+
+## Decision
+
+The [Python console entry](../../../../python/sdk-runtime/src/deepseek_harness_runtime/__init__.py) uses `subprocess.run` on Windows, inherits standard streams and environment, waits for runtime completion, and exits with the runtime status. POSIX retains `os.execvpe` process replacement. Windows CRT exec is not POSIX process replacement; the explicit spawn-and-wait path avoids the observed native exec operation.
+
+The [installed-wheel smoke](../../../../scripts/smoke-python-runtime.py) reports decimal and unsigned 32-bit hexadecimal status alongside captured streams when profile installation fails. This preserves the distinction between ordinary command failure and native process exceptions.
+
+## Alternatives considered
+
+**Disable Node compile caching.** Not selected: cache environment changes correlated with early probes, but cold-cache controls also passed and Python faulthandler locates the actual fault at the native exec call. Cache configuration remains unchanged.
+
+**Retry or bypass the installed console command.** Rejected because either masks the shipped command failure instead of repairing its process launch. The keyless installed-wheel assertion remains required.
+
+## Consequences
+
+Windows keeps a Python parent until the runtime exits; it no longer depends on CRT overlay behavior. The standard synchronous subprocess implementation owns waiting and interruption cleanup. No custom process-tree manager or global host setting is added.
+
+[Runtime-resolution tests](../../../../python/sdk/tests/test_runtime_resolution.py) retain POSIX forwarding and cover Windows argument/environment forwarding, statuses 0/37/513, real child completion, Unicode streams and arguments with spaces. Native Windows owns the wide exit-status case because POSIX truncates process statuses to eight bits. The [native fixed-count comparison](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34031142773) passes all four patched launches with compile caching enabled; all four unpatched controls also pass in that batch, so it is not a same-batch reproduction. Full installed-wheel CI must validate the final artifact separately from local branch-level tests.

+ 27 - 0
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.zh.md

@@ -0,0 +1,27 @@
+# Agent Note: 等待 Windows Python 控制台运行时
+
+Status: implemented
+
+[English](2026-09-06-windows-python-console-spawn-wait.md) | 中文
+
+## 问题
+
+Python 安装的 `dsh.exe` 控制台命令会在初始化 profile 前间歇性地以 Windows 访问冲突 `0xc0000005` 退出。其冒烟断言遗漏进程状态,只报告空标准流。[原生 faulthandler 探测](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34030851888) 将故障定位在运行时控制台入口调用的 Python 3.10 `os._execvpe`,而非打包的 Node 可执行文件。直接启动可执行文件的对照通过。
+
+## 决策
+
+[Python 控制台入口](../../../../python/sdk-runtime/src/deepseek_harness_runtime/__init__.py) 在 Windows 上使用 `subprocess.run`,继承标准流与环境,等待运行时结束,再以运行时状态退出。POSIX 保留 `os.execvpe` 进程替换。Windows CRT exec 并非 POSIX 进程替换;显式启动并等待的路径避开观测到的原生 exec 操作。
+
+[安装后 wheel 冒烟测试](../../../../scripts/smoke-python-runtime.py) 在 profile 安装失败时,同时报告十进制、无符号 32 位十六进制状态与捕获的标准流。这保留普通命令失败和原生进程异常的区别。
+
+## 已考虑的替代方案
+
+**禁用 Node 编译缓存。** 未采用:早期探测中缓存环境变化与结果相关,但冷缓存对照也能通过,且 Python faulthandler 将实际故障定位在原生 exec 调用。缓存配置保持不变。
+
+**重试或绕过已安装的控制台命令。** 拒绝,因为二者都会掩盖已发布命令的失败,而不是修复进程启动。keyless 安装后 wheel 断言仍为必需检查。
+
+## 后果
+
+Windows 保留 Python 父进程直到运行时退出,不再依赖 CRT overlay 行为。标准同步子进程实现负责等待和中断清理。不添加自定义进程树管理器或全局主机设置。
+
+[运行时解析测试](../../../../python/sdk/tests/test_runtime_resolution.py) 保留 POSIX 转发验证,并覆盖 Windows 参数/环境转发、状态 0/37/513、真实子进程完成、Unicode 标准流和带空格的参数。宽退出状态由原生 Windows 验证,因为 POSIX 会将进程状态截断为八位。[原生固定次数对照](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34031142773) 中,启用编译缓存的四次修复后启动全部通过;该批次四次未修复对照也全部通过,因此它不是同批次复现。完整安装后 wheel CI 必须独立于本地分支级测试,验证最终产物。

+ 2 - 2
.agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.md
-2026-07-20-ptc-typed-tool-returns.md: cc784ca8c9b1565135bc59c11ae0af973845fe5f
-2026-07-20-ptc-typed-tool-returns.zh.md: c4dd52d3c136569610beb8a65303809a1fe5fe4d
+2026-07-20-ptc-typed-tool-returns.md: 04ef9f7da4cd59ba07632684541dd6962008b810
+2026-07-20-ptc-typed-tool-returns.zh.md: c197d3131cc0f7fa326a9a47d945b2b7730f01c0

+ 2 - 2
.agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.md

@@ -75,7 +75,7 @@ Temporary Cordis Plugins follow the same rule: `cordis_mount` returns `{ id, plu
 
 
 Nested dispatch logs the sub-call's full rendered `content`/`isError` on `tool/code-dispatch` but does not persist canonical values. `tool/result` continues to persist only rendered content, error, and optional metadata. A successful final content sequence containing an image is also wrapped in a source-attributed user message and deferred through the outer result; the normal session event makes that model-visible input reconstructable. This feature did not itself require a structural Session-format change; the released v0-to-v1 identity edge preserves these records, and replay still cannot recreate intermediate canonical program values.
 Nested dispatch logs the sub-call's full rendered `content`/`isError` on `tool/code-dispatch` but does not persist canonical values. `tool/result` continues to persist only rendered content, error, and optional metadata. A successful final content sequence containing an image is also wrapped in a source-attributed user message and deferred through the outer result; the normal session event makes that model-visible input reconstructable. This feature did not itself require a structural Session-format change; the released v0-to-v1 identity edge preserves these records, and replay still cannot recreate intermediate canonical program values.
 
 
-The opaque `exec.parent` token marks nested calls. Presentation metadata and generic or tool-owned spill projections skip those calls because they have no direct result card and their canonical values never enter context. The outer `run_code` call alone produces one card and may spill its final post-policy presentation; `run_code` intentionally declares neither a result presenter nor presentation metadata, so UI adapters complete the card through their generic raw-content fallback using durable `tool/result.content`.
+The opaque `exec.parent` token marks nested calls. Presentation metadata and generic or tool-owned spill projections skip those calls; their canonical values never enter context. The Client can derive [nested terminal cards](../bug-fix/2026-09-05-nested-terminal-cards.md) from raw dispatch events without metadata. The outer `run_code` call produces the model-facing result and may spill its final post-policy presentation; `run_code` intentionally declares neither a result presenter nor presentation metadata, so UI adapters complete the card through their generic raw-content fallback using durable `tool/result.content`.
 
 
 ## Testing
 ## Testing
 
 
@@ -112,5 +112,5 @@ The worker performs bounded-depth flat-wire transport and lossless validation bu
 - The 64 MiB hard cap applies only to the outer variable payloads, excluding fixed result-envelope syntax and presentation whitespace; spill cannot recover bytes rejected beyond that cap.
 - The 64 MiB hard cap applies only to the outer variable payloads, excluding fixed result-envelope syntax and presentation whitespace; spill cannot recover bytes rejected beyond that cap.
 - Provider or executor acquisition limits may already have discarded source data before a canonical value reaches PTC mode.
 - Provider or executor acquisition limits may already have discarded source data before a canonical value reaches PTC mode.
 - Unsupported MCP output schemas fall back to `JsonValue`; admitted MCP images use the generic deferred projection, while audio and embedded-resource payloads remain diagnostic-only.
 - Unsupported MCP output schemas fall back to `JsonValue`; admitted MCP images use the generic deferred projection, while audio and embedded-resource payloads remain diagnostic-only.
-- There is one result card per outer `run_code`, never per nested call.
+- Nested calls have no separate model-facing result; Client child cards do not add model context.
 - Code failures expose `ToolCallError` message and tool name only, without a programmatic error-code union.
 - Code failures expose `ToolCallError` message and tool name only, without a programmatic error-code union.

+ 2 - 2
.agents/notes/implemented/feature/2026-07-20-ptc-typed-tool-returns.zh.md

@@ -75,7 +75,7 @@ PTC mode 通过运行时请求中的 `{ name: "ToolCallError", memberNamePropert
 
 
 嵌套分发在 `tool/code-dispatch` 上记录子调用完整渲染后的 `content`/`isError`,但不会持久化规范值。`tool/result` 继续只持久化渲染后的内容、错误和可选元数据。包含图片的成功最终内容序列还会包装成带来源归属的用户消息,并经外层结果延后;普通会话事件使该模型可见输入可以重建。该功能本身不要求结构性 Session 格式变更;已发布的 v0-to-v1 恒等边会保留这些记录,回放仍无法重建程序的规范中间值。
 嵌套分发在 `tool/code-dispatch` 上记录子调用完整渲染后的 `content`/`isError`,但不会持久化规范值。`tool/result` 继续只持久化渲染后的内容、错误和可选元数据。包含图片的成功最终内容序列还会包装成带来源归属的用户消息,并经外层结果延后;普通会话事件使该模型可见输入可以重建。该功能本身不要求结构性 Session 格式变更;已发布的 v0-to-v1 恒等边会保留这些记录,回放仍无法重建程序的规范中间值。
 
 
-不透明的 `exec.parent` token 用于标识嵌套调用。由于这些调用没有直接对应的结果卡片,而且其规范值永远不会进入上下文,展示元数据以及通用或工具自有的 spill 投影都会跳过它们。只有外层 `run_code` 调用会生成一张卡片,并且可能对 post-policy 处理后的最终展示执行 spill;`run_code` 有意既不声明结果展示器,也不声明展示元数据,因此 UI 适配器会通过通用的原始内容回退机制,使用持久化的 `tool/result.content` 补全该卡片。
+不透明的 `exec.parent` token 用于标识嵌套调用。展示元数据以及通用或工具自有的 spill 投影都会跳过这些调用;它们的规范值永远不会进入上下文。Client 可以从原始分发事件派生[嵌套 terminal 卡片](../bug-fix/2026-09-05-nested-terminal-cards.zh.md),无需元数据。外层 `run_code` 调用产生面向模型的结果,并且可能对 post-policy 处理后的最终展示执行 spill;`run_code` 有意既不声明结果展示器,也不声明展示元数据,因此 UI 适配器会通过通用的原始内容回退机制,使用持久化的 `tool/result.content` 补全该卡片。
 
 
 ## 测试
 ## 测试
 
 
@@ -112,5 +112,5 @@ worker 会以嵌套深度有界的扁平协议格式传输数据并执行无损
 - 64 MiB 硬上限只适用于外层可变负载,不计固定的结果封装语法与展示空白;spill 无法恢复超出该上限后被拒绝的字节。
 - 64 MiB 硬上限只适用于外层可变负载,不计固定的结果封装语法与展示空白;spill 无法恢复超出该上限后被拒绝的字节。
 - 提供方或执行器的采集上限可能在规范值到达 PTC mode 前就已丢弃部分源数据。
 - 提供方或执行器的采集上限可能在规范值到达 PTC mode 前就已丢弃部分源数据。
 - 不支持的 MCP 输出 schema 会回退为 `JsonValue`;已准入的 MCP 图片使用通用延后投影,而音频和嵌入资源载荷仍只提供诊断。
 - 不支持的 MCP 输出 schema 会回退为 `JsonValue`;已准入的 MCP 图片使用通用延后投影,而音频和嵌入资源载荷仍只提供诊断。
-- 每个外层 `run_code` 只有一张结果卡片,嵌套调用不会各自生成卡片。
+- 嵌套调用没有独立的面向模型的结果;Client 子调用卡片不会增加模型上下文。
 - PTC mode 失败只暴露 `ToolCallError` 的消息与工具名,不提供程序可用的错误代码联合。
 - PTC mode 失败只暴露 `ToolCallError` 的消息与工具名,不提供程序可用的错误代码联合。

+ 2 - 2
.agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.md
-2026-07-23-session-telemetry-otel-revival.md: f0dca55a3bf705f650a7a8fc998e81bba04649d4
-2026-07-23-session-telemetry-otel-revival.zh.md: 488716f9723238f1841f79d2cf55f52e269bb3db
+2026-07-23-session-telemetry-otel-revival.md: dc52ee87ef78107d674e49af3c20c255266f0414
+2026-07-23-session-telemetry-otel-revival.zh.md: f647ea289684eca648fe797011daad858612546f

+ 3 - 3
.agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.md

@@ -14,7 +14,7 @@ Every deployment that wants harness sessions in an observability stack must hand
 
 
 - **`@deepseek-ai/dsh-session-telemetry`** — the seam. `SessionTelemetrySink` (`emit`/`flush?`/`shutdown`), the service-registered `SessionTelemetryBackend` form, and `SessionTelemetryCoordinator` own lifecycle-local capture: each new Session object starts immediately before `firstLiveSeq`, then the per-append firehose deep-copies, redacts, and hands off every event with zero I/O; re-adopting the same object resumes after its module-scope cursor. Buffer-free on-demand capture uses the same one-record-per-event mapping through an optional inclusive boundary. Ledger identity includes `session.id`, `session.format_version`, and `event.seq`; live capture also relays `agent/error` and creates dispose-time `shutdown` records.
 - **`@deepseek-ai/dsh-session-telemetry`** — the seam. `SessionTelemetrySink` (`emit`/`flush?`/`shutdown`), the service-registered `SessionTelemetryBackend` form, and `SessionTelemetryCoordinator` own lifecycle-local capture: each new Session object starts immediately before `firstLiveSeq`, then the per-append firehose deep-copies, redacts, and hands off every event with zero I/O; re-adopting the same object resumes after its module-scope cursor. Buffer-free on-demand capture uses the same one-record-per-event mapping through an optional inclusive boundary. Ledger identity includes `session.id`, `session.format_version`, and `event.seq`; live capture also relays `agent/error` and creates dispose-time `shutdown` records.
 - **The `session-telemetry/record` waterfall** — the delta over the branch version and the seam's redaction extension point. Every record passes it before reaching any backend; the seam ships NO rules of its own — the innermost `next()` is a pass-through, deployments mount their rules as listeners (stacking by transforming `next()`'s return value), and a throwing rule withholds the record fail-closed. Redaction applies to the exported copy only; the canonical log is never rewritten.
 - **The `session-telemetry/record` waterfall** — the delta over the branch version and the seam's redaction extension point. Every record passes it before reaching any backend; the seam ships NO rules of its own — the innermost `next()` is a pass-through, deployments mount their rules as listeners (stacking by transforming `next()`'s return value), and a throwing rule withholds the record fail-closed. Redaction applies to the exported copy only; the canonical log is never rewritten.
-- **`@deepseek-ai/dsh-session-telemetry-otel`** — the reference backend: OTel JS SDK log pipeline (`LoggerProvider` → `BatchLogRecordProcessor` → OTLP/HTTP exporter), configured verbatim through `exporter`/`processor` passthroughs. `DISABLED` is the default and constructs no transport; the [feedback-gated telemetry decision](../../archived/feature/2026-08-05-feedback-gated-session-telemetry.md) defines the explicit `FULL` and `FEEDBACK_ONLY` delivery modes, which require `exporter.url`, without moving the redaction or backend boundary. [Buffer-free feedback replay](../../archived/simplification/2026-08-06-buffer-free-feedback-telemetry.md) avoids a second in-memory copy of the session prefix.
+- **`@deepseek-ai/dsh-session-telemetry-otel`** — the reference backend: OTel JS SDK log pipeline (`LoggerProvider` → `BatchLogRecordProcessor` → OTLP/HTTP exporter), configured verbatim through `exporter`/`processor` passthroughs. `DISABLED` is the plugin default and constructs no transport. `FEEDBACK_ONLY` requires `exporter.url`; the [explicit-feedback policy](../architecture/2026-09-05-nonofficial-feedback-otel.md) requires a new submission for every bounded capture, for all providers. [Buffer-free feedback replay](../../archived/simplification/2026-08-06-buffer-free-feedback-telemetry.md) avoids a second in-memory copy of the session prefix.
 
 
 The boundary axiom holds: the harness's aspect ends at `emit()`. Batching, retry, queueing, and loss policy are the reporting SDK's, configured through passthroughs. Delivery is best-effort: a crash can lose queued records. Backend retry or losing the same live object's module-scope cursor can duplicate lifecycle-local rows, so receivers deduplicate on `(session.id, session.format_version, event.seq)`.
 The boundary axiom holds: the harness's aspect ends at `emit()`. Batching, retry, queueing, and loss policy are the reporting SDK's, configured through passthroughs. Delivery is best-effort: a crash can lose queued records. Backend retry or losing the same live object's module-scope cursor can duplicate lifecycle-local rows, so receivers deduplicate on `(session.id, session.format_version, event.seq)`.
 
 
@@ -28,10 +28,10 @@ The boundary axiom holds: the harness's aspect ends at `emit()`. Batching, retry
 
 
 **Map onto OTel spans (GenAI semantic conventions) instead of logs.** Rejected for this revival: the branch implementation's log mapping is reviewed and shipped-shaped; the span model is lossy for forkable, interruptible sessions and belongs to a future consumer with real span queries to serve.
 **Map onto OTel spans (GenAI semantic conventions) instead of logs.** Rejected for this revival: the branch implementation's log mapping is reviewed and shipped-shaped; the span model is lossy for forkable, interruptible sessions and belongs to a future consumer with real span queries to serve.
 
 
-**Replay the complete constructor seed for every new Session object.** Rejected because a fork's inherited events and a resumed or migrated log belong to another lifecycle and may predate the current sharing act. Replaying them under the new object can re-release already shared data, attribute inherited events to a child Session, and make feedback acknowledgement text understate what is handed off. A new object therefore starts immediately before `firstLiveSeq`: fresh Sessions still begin at seq 0, while seeded Sessions begin with their lifecycle boundary. Re-adopting the same object resumes after its cursor, so HMR does not duplicate the settled lifecycle suffix. This rule does not provide crash backfill; a deployment requiring guaranteed historical delivery needs the deferred durable outbox and an explicit disclosure for that broader scope.
+**Replay every constructor seed automatically.** A fork's inherited events and a resumed log may predate the current sharing act. The coordinator defaults to `firstLiveSeq` for a new object; `includeHistory` explicitly permits a backend to capture stored context. OTel selects complete history only through new own feedback, never on construction, restoration, or HMR. Same-object cursors reduce duplicate handoff but do not guarantee historical delivery or replace the deferred durable outbox.
 
 
 **Forwarding the seam's turn-boundary `flush()` hint to the OTel provider's `forceFlush()`.** Shipped in the first revival round, then removed: three distinct silent-loss paths shared the wrapper state — a dispose racing an in-flight flush (the SDK's concurrent-flush guard makes shutdown's internal drain skip), overlapping hints displacing the retained promise, and the provider's fixed 30-second flush timeout rejecting while the processor still drains. Every path exists only because the forwarding made this backend the process's second flusher against undocumented SDK internals from the upstream experimental tree; with no `flush()` implemented, the batch processor is the only flusher, its `scheduledDelayMillis` (already deployment-tunable through the `processor` passthrough) governs export cadence, and `shutdown()`'s drain is complete by construction. Reinstate only if a deployment states a turn-boundary latency requirement `scheduledDelayMillis` cannot meet — and then by calling the retained `BatchLogRecordProcessor`'s own `forceFlush()`, never the provider's timeout-wrapped one.
 **Forwarding the seam's turn-boundary `flush()` hint to the OTel provider's `forceFlush()`.** Shipped in the first revival round, then removed: three distinct silent-loss paths shared the wrapper state — a dispose racing an in-flight flush (the SDK's concurrent-flush guard makes shutdown's internal drain skip), overlapping hints displacing the retained promise, and the provider's fixed 30-second flush timeout rejecting while the processor still drains. Every path exists only because the forwarding made this backend the process's second flusher against undocumented SDK internals from the upstream experimental tree; with no `flush()` implemented, the batch processor is the only flusher, its `scheduledDelayMillis` (already deployment-tunable through the `processor` passthrough) governs export cadence, and `shutdown()`'s drain is complete by construction. Reinstate only if a deployment states a turn-boundary latency requirement `scheduledDelayMillis` cannot meet — and then by calling the retained `BatchLogRecordProcessor`'s own `forceFlush()`, never the provider's timeout-wrapped one.
 
 
 ## Consequences
 ## Consequences
 
 
-A deployment adds one `cordis.yml` entry with an OTLP endpoint and explicitly selects `FULL` to stream lifecycle-local canonical events into an OTel-compatible stack or `FEEDBACK_ONLY` to replay a lifecycle-local canonical-log prefix when feedback is recorded. `DISABLED` is the default ([`dsh-session-telemetry-otel` README](../../../../packages/session/session-telemetry-otel/README.md)) and constructs no reporting pipeline; removing the entry remains a silent opt-out, while the disabled mode keeps the local feedback warning. A rule-free deployment exports each event body exactly as captured — including every assistant chunk and any credentials embedded in file contents or command output — so a deployment crossing a trust boundary must mount `session-telemetry/record` listeners, and both READMEs state this plainly. Where rules are mounted, exported bodies can differ from canonical log bytes, so receivers must not treat telemetry as a byte-exact replica; the log remains the source of truth. Same-object replay after lost cursor state can duplicate lifecycle-local ledger rows, and crash durability remains out of scope until the outbox decision above is revisited.
+A deployment configures an OTLP endpoint and selects `FEEDBACK_ONLY` to release a complete canonical prefix at new explicit feedback. `DISABLED` is the plugin default ([`dsh-session-telemetry-otel` README](../../../../packages/session/session-telemetry-otel/README.md)) and constructs no reporting pipeline; removing the entry is a silent opt-out, while disabled mode keeps the local feedback warning. SDK scheduled flush and shutdown only drain authorized batches, without capturing later records. A rule-free deployment exports each captured event body, including compact assistant streams and credentials embedded in file contents or command output, so deployments crossing a trust boundary must mount `session-telemetry/record` listeners. Redacted bodies can differ from canonical log bytes; the log remains the source of truth. Lost handoff cursors can cause duplicates on a later authorized capture, and crash durability remains out of scope until the outbox decision above is revisited.

+ 3 - 4
.agents/notes/implemented/feature/2026-07-23-session-telemetry-otel-revival.zh.md

@@ -14,8 +14,7 @@ Status: implemented
 
 
 - **`@deepseek-ai/dsh-session-telemetry`** —— seam 本体。`SessionTelemetrySink`(`emit`/`flush?`/`shutdown`)、服务注册形态的 `SessionTelemetryBackend` 与 `SessionTelemetryCoordinator` 共同拥有生命周期本地捕获:每个新的 Session 对象从 `firstLiveSeq` 之前开始,随后逐 append firehose 以零 I/O 深拷贝、脱敏并交接每个事件;重新收养同一对象时从模块作用域游标之后继续。无缓冲按需捕获使用同样的一事件一记录映射,直到可选的包含式边界。Ledger 身份包含 `session.id`、`session.format_version` 与 `event.seq`;实时捕获还会转发 `agent/error`,并创建 dispose(资源释放)时的 `shutdown` 记录。
 - **`@deepseek-ai/dsh-session-telemetry`** —— seam 本体。`SessionTelemetrySink`(`emit`/`flush?`/`shutdown`)、服务注册形态的 `SessionTelemetryBackend` 与 `SessionTelemetryCoordinator` 共同拥有生命周期本地捕获:每个新的 Session 对象从 `firstLiveSeq` 之前开始,随后逐 append firehose 以零 I/O 深拷贝、脱敏并交接每个事件;重新收养同一对象时从模块作用域游标之后继续。无缓冲按需捕获使用同样的一事件一记录映射,直到可选的包含式边界。Ledger 身份包含 `session.id`、`session.format_version` 与 `event.seq`;实时捕获还会转发 `agent/error`,并创建 dispose(资源释放)时的 `shutdown` 记录。
 - **`session-telemetry/record` waterfall(瀑布式事件)** —— 相对分支版本的增量,也是该 seam 的脱敏扩展点。每条记录抵达任何后端前必经此处;seam 自身不带任何规则——最内层 `next()` 原样透传,部署方以监听器挂载自己的规则(通过变换 `next()` 的返回值堆叠),抛异常的规则将该记录 fail-closed 扣下。脱敏只作用于导出副本;canonical log 永不改写。
 - **`session-telemetry/record` waterfall(瀑布式事件)** —— 相对分支版本的增量,也是该 seam 的脱敏扩展点。每条记录抵达任何后端前必经此处;seam 自身不带任何规则——最内层 `next()` 原样透传,部署方以监听器挂载自己的规则(通过变换 `next()` 的返回值堆叠),抛异常的规则将该记录 fail-closed 扣下。脱敏只作用于导出副本;canonical log 永不改写。
-- **`@deepseek-ai/dsh-session-telemetry-otel`** —— 参考后端:OTel JS SDK 日志流水线(`LoggerProvider` → `BatchLogRecordProcessor` → OTLP/HTTP exporter),经 `exporter`/`processor` passthrough 原样配置。`DISABLED` 是默认值,且不构造任何传输;[反馈门控遥测决策](../../archived/feature/2026-08-05-feedback-gated-session-telemetry.md)定义了需显式启用的 `FULL` 与 `FEEDBACK_ONLY` 投递模式,这两种模式要求 `exporter.url`,且不移动脱敏或后端边界。[无缓冲反馈回放](../../archived/simplification/2026-08-06-buffer-free-feedback-telemetry.md)避免在内存中创建会话前缀的第二份副本。
-
+- **`@deepseek-ai/dsh-session-telemetry-otel`** —— 参考后端:OTel JS SDK 日志流水线(`LoggerProvider` → `BatchLogRecordProcessor` → OTLP/HTTP exporter),经 `exporter`/`processor` passthrough 原样配置。`DISABLED` 是插件默认值,且不构造任何传输。`FEEDBACK_ONLY` 要求 `exporter.url`;[显式反馈策略](../architecture/2026-09-05-nonofficial-feedback-otel.zh.md)要求每次有界捕获都由新提交触发,适用于所有提供方。[无缓冲反馈回放](../../archived/simplification/2026-08-06-buffer-free-feedback-telemetry.md)避免在内存中创建会话前缀的第二份副本。
 
 
 边界公理保持不变:harness 的职责止于 `emit()`。批处理、重试、排队与丢失策略属于 reporting SDK,并经 passthrough 配置。投递是尽力而为:崩溃可能丢失已排队记录。后端重试或丢失同一 live 对象的模块作用域游标可能重复生命周期本地 row,因此接收端基于 `(session.id, session.format_version, event.seq)` 去重。
 边界公理保持不变:harness 的职责止于 `emit()`。批处理、重试、排队与丢失策略属于 reporting SDK,并经 passthrough 配置。投递是尽力而为:崩溃可能丢失已排队记录。后端重试或丢失同一 live 对象的模块作用域游标可能重复生命周期本地 row,因此接收端基于 `(session.id, session.format_version, event.seq)` 去重。
 
 
@@ -29,10 +28,10 @@ Status: implemented
 
 
 **映射到 OTel span(GenAI 语义约定)而非日志。** 本次复活否决:分支实现的日志映射已经过评审、形态可交付;span 模型对可 fork、可中断的会话有损,留给将来真正有 span 查询需求的消费方。
 **映射到 OTel span(GenAI 语义约定)而非日志。** 本次复活否决:分支实现的日志映射已经过评审、形态可交付;span 模型对可 fork、可中断的会话有损,留给将来真正有 span 查询需求的消费方。
 
 
-**为每个新 Session 对象回放完整 constructor seed。** 不予采用,因为 fork 的继承事件以及 resume 或迁移日志属于其他生命周期,可能早于当前共享动作。在新对象下回放这些内容会再次释放已共享数据、把继承事件归因给 child Session,并使反馈确认文本低估实际交接范围。因此,新对象从 `firstLiveSeq` 之前开始:全新 Session 仍从 seq 0 开始,seeded Session 则从其生命周期边界开始。重新收养同一对象时仍从游标之后继续,因此 HMR 不会重复稳定的生命周期后缀。该规则不提供崩溃回填;要求保证历史投递的部署需要延后的 durable outbox,并对更广范围作出显式披露。
+**自动回放每个 constructor seed。** fork 的继承事件和恢复的日志可能早于当前共享动作。协调器对新对象默认从 `firstLiveSeq` 开始;`includeHistory` 显式允许后端捕获存储的上下文。OTel 仅在新的自身反馈时选择完整历史,绝不在构造、恢复或 HMR 时捕获。同对象游标减少重复交接,但不保证历史投递,也不替代延后的持久化 outbox。
 
 
 **将 seam 的轮次边界 `flush()` 提示转发到 OTel 提供方的 `forceFlush()`。** 首轮复活曾交付此转发,其后移除:三条不同的静默丢失路径共用同一份包装层状态——dispose 与进行中的 flush 之间的竞态(SDK 的并发 flush 防护会令 shutdown 的内部排空被跳过)、相互重叠的提示顶掉留存的 promise、以及提供方固定的 30 秒 flush 超时在批处理器仍在排空时便 reject。这些路径存在的唯一原因,是该转发让这个后端成为进程内第二个执行 flush 的组件,面对的还是上游实验性(experimental)源码树中未见诸文档的 SDK 内部行为;不实现 `flush()` 时,批处理器就是唯一执行 flush 的组件,其 `scheduledDelayMillis`(已可由部署方经 `processor` passthrough 调优)决定导出节奏,`shutdown()` 的排空从构造上就是完整的。仅当某个部署提出 `scheduledDelayMillis` 无法满足的轮次边界延迟要求时才恢复此转发——且届时应调用留存的 `BatchLogRecordProcessor` 自身的 `forceFlush()`,绝不调用提供方那个带超时包装的版本。
 **将 seam 的轮次边界 `flush()` 提示转发到 OTel 提供方的 `forceFlush()`。** 首轮复活曾交付此转发,其后移除:三条不同的静默丢失路径共用同一份包装层状态——dispose 与进行中的 flush 之间的竞态(SDK 的并发 flush 防护会令 shutdown 的内部排空被跳过)、相互重叠的提示顶掉留存的 promise、以及提供方固定的 30 秒 flush 超时在批处理器仍在排空时便 reject。这些路径存在的唯一原因,是该转发让这个后端成为进程内第二个执行 flush 的组件,面对的还是上游实验性(experimental)源码树中未见诸文档的 SDK 内部行为;不实现 `flush()` 时,批处理器就是唯一执行 flush 的组件,其 `scheduledDelayMillis`(已可由部署方经 `processor` passthrough 调优)决定导出节奏,`shutdown()` 的排空从构造上就是完整的。仅当某个部署提出 `scheduledDelayMillis` 无法满足的轮次边界延迟要求时才恢复此转发——且届时应调用留存的 `BatchLogRecordProcessor` 自身的 `forceFlush()`,绝不调用提供方那个带超时包装的版本。
 
 
 ## 后果
 ## 后果
 
 
-部署方在 `cordis.yml` 加一个带 OTLP endpoint 的 Cordis 配置项,并显式选择 `FULL`,即可把生命周期本地权威事件流接入任何 OTel 兼容体系;选择 `FEEDBACK_ONLY` 则会在记录反馈时回放生命周期本地权威日志前缀。`DISABLED` 是默认值([`dsh-session-telemetry-otel` README](../../../../packages/session/session-telemetry-otel/README.zh.md)),且不构造上报流水线;删除该配置项仍是静默退出方式,而禁用模式会保留本地反馈警告。未挂载规则的部署会按捕获原样导出每个事件 body,包括每条 assistant chunk,以及文件内容与命令输出中内嵌的任何凭据。因此,跨信任边界的部署必须挂载 `session-telemetry/record` 监听器,两份 README 对此如实陈述。挂载规则后,导出的 body 可能与 canonical log 字节不同,接收端不得把遥测当作字节精确副本;日志仍是真源。同一对象丢失游标状态后的回放可能重复生命周期本地 ledger 行,崩溃持久性则在上述 outbox 决定重新审议前继续不在范围内。
+部署方配置 OTLP endpoint 并选择 `FEEDBACK_ONLY`,即可在新的显式反馈时释放完整权威日志前缀。`DISABLED` 是插件默认值([`dsh-session-telemetry-otel` README](../../../../packages/session/session-telemetry-otel/README.zh.md)),且不构造上报流水线;删除配置项是静默退出方式,而禁用模式保留本地反馈警告。SDK 定时刷新和关闭只排空已授权批次,不捕获后续记录。未挂载规则的部署会导出每条已捕获事件的正文,包括紧凑 assistant stream 以及文件内容或命令输出中内嵌的凭据,因此跨信任边界的部署必须挂载 `session-telemetry/record` 监听器。脱敏后的正文可能与权威日志字节不同;日志仍是真源。丢失交接游标可能使后续已授权捕获产生重复,崩溃持久性则在上述 outbox 决定重新审议前继续不在范围内。

+ 2 - 2
.agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.md
-2026-08-25-feedback-gated-telemetry-default.md: 19ed75859b54044d38c473f9bd5e8840c5cd2844
-2026-08-25-feedback-gated-telemetry-default.zh.md: d82776b4250e1c37b821df5f387f9df05e0fbfd0
+2026-08-25-feedback-gated-telemetry-default.md: 1a5ecd6b34415650db5e5068d96f9d2dc5f30e42
+2026-08-25-feedback-gated-telemetry-default.zh.md: 55936aa4045214b627ecaad39d898b1898a50fc9

+ 10 - 8
.agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.md

@@ -10,20 +10,22 @@ Diagnosing a `/feedback` report needs the session data the report describes. Wit
 
 
 ## Decision
 ## Decision
 
 
-The shared dsh base resolves an unset or empty `DSH_TELEMETRY_MODE` to `FEEDBACK_ONLY` instead of `DISABLED`. Nothing is uploaded before the user records `/feedback`. On a Session object already captured, each feedback uploads the suffix after the last handoff through that exact event. A new object starts at its constructor boundary: a fresh Session begins at seq 0, while a forked, resumed, or migrated Session excludes its constructor seed and begins with this lifecycle's `session/end-seed`. The acknowledgement's sharing disclosure therefore matches the released lifecycle-local prefix. `FULL` and `DISABLED` remain explicit `DSH_TELEMETRY_MODE` overrides, any non-empty `DSH_TELEMETRY_DISABLED` remains the authoritative pre-load hard opt-out, and the plugin's own omitted-`mode` default stays `DISABLED`: the default changes only in the shared base's config expression, where deployments already override it.
+The [explicit-feedback OTel decision](../architecture/2026-09-05-nonofficial-feedback-otel.md) owns upload authorization for all users, including `deepseek-official`. This note retains the rationale for the base default: feedback-gated release rather than continuous export.
 
 
-This supersedes the session-backend default of the [default-off decision](../../archived/feature/2026-08-10-telemetry-default-off.md), accepting the user's explicit feedback action as the release authorization that note required a deployment setting for. That note's hard opt-out and its launcher-feed history remain current, and the [default-mount decision](../../archived/feature/2026-07-31-web-telemetry-default-mount.md) continues to own the endpoint, batching cadence, and exit-drain settings.
+The shared base resolves an unset or empty `DSH_TELEMETRY_MODE` to `FEEDBACK_ONLY`. The plugin's own omitted-`mode` default is `DISABLED`; `FULL` rejects, and non-empty `DSH_TELEMETRY_DISABLED` is the pre-load hard opt-out. New own text feedback, message-rating edits, and withdrawals release the unhanded canonical prefix through that event, including stored context. Inherited parent feedback does not authorize a child export.
+
+Feedback-gated release lets a reporter share the Session that exhibited the problem without reproducing it. It trades continuous export for an explicit feedback trigger. The [archived default-off](../../archived/feature/2026-08-10-telemetry-default-off.md) and [default-mount](../../archived/feature/2026-07-31-web-telemetry-default-mount.md) notes record the earlier composition; the [base patch](../../../../packages/bundle/base/cordis.patch.yml) and [OTel README](../../../../packages/session/session-telemetry-otel/README.md) own current configuration.
 
 
 ## Alternatives considered
 ## Alternatives considered
 
 
-**Keep `DISABLED` and instruct reporters to re-run with `DSH_TELEMETRY_MODE=FEEDBACK_ONLY`.** Rejected: the session that exhibited the problem is the one worth uploading, and re-running loses it.
+**Require reporters to re-run after enabling telemetry.** Rejected as the feedback-gated workflow: the Session that exhibited the problem is the useful evidence, and re-running loses it.
 
 
-**Default to `FULL`.** Rejected: continuous export without any user action is exactly what the default-off decision forbids, and nothing in a fresh installation authorizes it.
+**Permit continuous export.** Rejected: deployment configuration does not authorize capture without explicit feedback.
 
 
-**Gate the official DeepSeek `dsh_session_log` request contribution on feedback instead of reviving the OTel default.** Not taken here: that contribution uploads through subsequent LLM requests rather than at the feedback boundary, so a session's final feedback would never be delivered; a feedback-triggered flush on that path is a larger design than a default flip.
+**Use only subsequent DeepSeek requests for delivery.** The independent opt-in contribution can carry canonical feedback, but a final feedback entry may have no later request. OTel releases it for every provider without initiating another LLM request.
 
 
 ## Consequences
 ## Consequences
 
 
-- A fresh installation uploads the not-yet-shared session-log records to the production collector when — and only when — the user records `/feedback`; no other trigger uploads.
-- Released exports remain the raw captured copy: the shipped base mounts no `session-telemetry/record` redaction rule, so they can contain message text, tool arguments and results, and workspace paths.
-- The sharing disclosure is part of the `/feedback` acknowledgement, so the user reads it after the release has been triggered. A deployment that requires prior informed consent must override the default to `DISABLED` or add a pre-upload confirmation before this default is defensible there.
+- The shipped base releases a bounded prefix only at new explicit feedback. Ordinary requests, lifecycle events, and stored feedback do not trigger capture. Later records wait for the next explicit feedback.
+- On-demand capture copies and redacts the canonical log at feedback time. Without a deployment redaction rule, exported data can include message text, tool arguments and results, and workspace paths.
+- The command acknowledgement confirms recording, not sharing or delivery. A deployment requiring prior informed consent must provide it before enabling uploads; OTel handoff remains subject to the SDK's batching, retry, and loss policy.

+ 10 - 8
.agents/notes/implemented/feature/2026-08-25-feedback-gated-telemetry-default.zh.md

@@ -10,20 +10,22 @@ Status: implemented
 
 
 ## 决定
 ## 决定
 
 
-共享 dsh 基础配置把未设置或为空的 `DSH_TELEMETRY_MODE` 解析为 `FEEDBACK_ONLY` 而不是 `DISABLED`。用户记录 `/feedback` 之前不上传任何数据。对于已经捕获的 Session 对象,每条反馈会上传从上次交接之后至该事件的后缀。新对象从 constructor boundary 开始:全新 Session 从 seq 0 开始,而 fork、resume 或迁移 Session 排除 constructor seed,从本生命周期的 `session/end-seed` 开始。因此,确认信息中的共享声明与所释放的生命周期本地前缀一致。`FULL` 和 `DISABLED` 仍是显式的 `DSH_TELEMETRY_MODE` 覆盖值,任何非空的 `DSH_TELEMETRY_DISABLED` 仍是加载前的强制关闭开关,插件自身省略 `mode` 的默认值仍是 `DISABLED`:默认值只在共享基础配置的配置表达式中改变,部署本来就在那里覆盖它。
+[显式反馈 OTel 决策](../architecture/2026-09-05-nonofficial-feedback-otel.zh.md)负责所有用户的上传授权,包括 `deepseek-official`。本记录保留基础默认值的理由:选择反馈门控释放而非持续导出。
 
 
-本决定取代[默认关闭决定](../../archived/feature/2026-08-10-telemetry-default-off.md)中会话后端的默认值,把用户显式的反馈动作接受为该决定原本要求由部署设置提供的释放授权。该决定的强制关闭开关和 launcher 上报历史仍然有效,端点、批处理节奏和退出排空设置仍由[默认挂载决定](../../archived/feature/2026-07-31-web-telemetry-default-mount.md)持有。
+共享基础配置把未设置或为空的 `DSH_TELEMETRY_MODE` 解析为 `FEEDBACK_ONLY`。插件自身省略 `mode` 的默认值是 `DISABLED`;`FULL` 被拒绝,非空 `DSH_TELEMETRY_DISABLED` 是加载前的强制关闭开关。新的自身文本反馈、消息评分编辑和撤回释放尚未交接的权威前缀,截止该事件,包含存储的上下文。继承的父会话反馈不授权子会话导出。
+
+反馈门控释放让报告者无需复现问题就能共享出问题的 Session。它用显式反馈触发取代持续导出。[已归档默认关闭](../../archived/feature/2026-08-10-telemetry-default-off.md)与[默认挂载](../../archived/feature/2026-07-31-web-telemetry-default-mount.md)记录记载早期组合;当前配置由[基础补丁](../../../../packages/bundle/base/cordis.patch.yml)与 [OTel README](../../../../packages/session/session-telemetry-otel/README.zh.md) 持有。
 
 
 ## 考虑过的替代方案
 ## 考虑过的替代方案
 
 
-**保持 `DISABLED`,让报告者带着 `DSH_TELEMETRY_MODE=FEEDBACK_ONLY` 重跑。** 否决:值得上传的正是出现问题的那个会话,重跑会丢掉它。
+**要求报告者启用遥测后重跑。** 不作为反馈门控工作流:值得保留的证据是出问题的那个 Session,重跑会丢掉它。
 
 
-**默认 `FULL`。** 否决:没有任何用户动作的持续导出正是默认关闭决定所禁止的,全新安装中没有任何东西授权它。
+**允许持续导出。** 否决:部署配置不授权没有显式反馈的捕获。
 
 
-**改为在反馈时门控官方 DeepSeek `dsh_session_log` 请求贡献,而不是恢复 OTel 默认值。** 此处未采用:该贡献通过后续 LLM 请求上传,而不是在反馈边界上传,会话的最后一条反馈永远不会被交付;在那条路径上做反馈触发的冲刷是比翻转默认值更大的设计。
+**仅通过后续 DeepSeek 请求投递。** 独立的需显式启用的贡献可传送权威反馈,但最终反馈之后可能没有请求。OTel 为每个提供方释放反馈,无需发起另一个 LLM 请求。
 
 
 ## 后果
 ## 后果
 
 
-- 全新安装只在用户记录 `/feedback` 时把尚未共享的会话日志记录上传到生产 collector;没有其他触发上传的途径。
-- 释放的导出仍是未加工的原始副本:随附基础配置没有挂载 `session-telemetry/record` 脱敏规则,导出可能包含消息文本、工具参数和结果,以及 workspace 路径。
-- 共享声明是 `/feedback` 确认信息的一部分,用户读到它时释放已被触发。要求事先知情同意的部署必须把默认值覆盖为 `DISABLED`,或在上传前增加确认步骤,此默认值在那类部署中才站得住。
+- 随附基础配置仅在新的显式反馈时释放有界前缀。普通请求、生命周期事件和已存储反馈不触发捕获。后续记录等待下一次显式反馈。
+- 按需捕获在反馈时复制权威日志并脱敏。部署未挂载脱敏规则时,导出数据可能包含消息文本、工具参数和结果,以及 workspace 路径。
+- 命令确认文本确认记录,而非共享或投递。要求事先知情同意的部署必须在启用上传前提供该步骤;OTel 交接仍受 SDK 的批处理、重试与丢失策略约束。

+ 2 - 2
.agents/notes/implemented/process/2026-07-26-ci-failover-runbook.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.md
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.md
-2026-07-26-ci-failover-runbook.md: b24996a4ba4dfaa4b26f88519a61f45c81efb5b5
-2026-07-26-ci-failover-runbook.zh.md: ee7339d70e4796c367f97490ea57468687464b3f
+2026-07-26-ci-failover-runbook.md: d5c12492671941c45cf3ccab255dd76fb53773bd
+2026-07-26-ci-failover-runbook.zh.md: b42a14dc9a3e4f40c59b27633752f6969173777e

+ 12 - 8
.agents/notes/implemented/process/2026-07-26-ci-failover-runbook.md

@@ -6,11 +6,11 @@ English | [中文](2026-07-26-ci-failover-runbook.zh.md)
 
 
 ## Problem
 ## Problem
 
 
-The three required Linux worker jobs in [CI](../../../../.github/workflows/ci.yml) (`node 24 / static`, `node 24 / coverage`, `node 24 / snapshots and artifacts`) run on the hosted enterprise 32-core pools; the required verdict job that aggregates them (`all checks passed`) runs on standard `ubuntu-latest`; the independent native Windows job (`windows node 24 / native complete`) runs on the hosted `dsh-windows-2025-16core` larger runner. When the enterprise pools degrade — jobs queue indefinitely or the enterprise labels vanish — every open pull request becomes unmergeable, and the ordinary recovery of merging a fix is itself deadlocked behind the very required checks that cannot run. **Scope: two independent switches, one per platform.** `DSH_CI_FAILOVER_LINUX` recovers an enterprise Linux-pool outage (the three required Linux workers plus the `all checks passed` verdict); `DSH_CI_FAILOVER_WINDOWS` recovers a hosted Windows-pool outage (the native Windows job). A Linux-pool outage need not retarget the native Windows job and vice versa. The verdict's other required dependencies (`node-compat`, `python-sdk`, `windows`) stay on standard hosted runners by design (the portable boundary); in a broader GitHub-hosted capacity failure that also takes out the standard pools, those dependencies still block `all checks passed`. An outage therefore needs a switch any responder with repository write access can throw without merging anything.
+The three required Linux worker jobs in [CI](../../../../.github/workflows/ci.yml) (`node 24 / static`, `node 24 / coverage`, `node 24 / snapshots and artifacts`) run on the hosted enterprise 32-core pools; the required verdict job that aggregates them (`all checks passed`) runs on standard `ubuntu-latest`; the [native Windows jobs](2026-08-08-native-windows-pull-request-ci.md) run on the hosted `dsh-windows-2025-16core` larger runner. When the enterprise pools degrade — jobs queue indefinitely or the enterprise labels vanish — every open pull request becomes unmergeable, and the ordinary recovery of merging a fix is itself deadlocked behind the very required checks that cannot run. **Scope: two independent switches, one per platform.** `DSH_CI_FAILOVER_LINUX` recovers an enterprise Linux-pool outage (the three required Linux workers plus the `all checks passed` verdict); `DSH_CI_FAILOVER_WINDOWS` recovers a hosted Windows-pool outage (the native Windows jobs). A Linux-pool outage need not retarget Windows jobs and vice versa. The verdict's other required dependencies (`node-24-bench`, `node-compat`, `python-sdk`, `windows`) stay on standard hosted runners by design (the portable boundary); in a broader GitHub-hosted capacity failure that also takes out the standard pools, those dependencies still block `all checks passed`. An outage therefore needs a switch any responder with repository write access can throw without merging anything.
 
 
 ## Decision
 ## Decision
 
 
-Each of the three required Linux worker jobs, the independent native Windows job, and the `all checks passed` verdict job — which would otherwise stay queued on the failed pool even after every worker passed — resolves its runner pool through a repository variable, and the switch is split by platform so an outage on one platform does not retarget the other. The three Linux workers and the `all checks passed` verdict (whose `needs` are the required Linux workers and which runs on the `vm-backup` pool) resolve through `DSH_CI_FAILOVER_LINUX`; the native Windows job resolves through `DSH_CI_FAILOVER_WINDOWS`. Unset (normal), they run on the hosted enterprise pools. Set to `selfhosted` by any repository writer, the corresponding jobs retarget onto the in-house self-hosted pool: under `DSH_CI_FAILOVER_LINUX`, the Linux jobs and verdict move onto the `vm-backup` pool, snapshot concurrency drops to the shared-VM bound, and the hosted-path pnpm cache restores are skipped; under `DSH_CI_FAILOVER_WINDOWS`, the native Windows job moves onto the `dsh-win-ci` pool. Each switch is writer-manageable repository state, not a merge, so it works while every check is red. The in-house pools' readiness is continuously re-proven by the `serial / linux (self-hosted standby)` and `serial / windows (self-hosted standby)` lanes, which run the complete unsharded aggregates on every master push.
+Each of the three required Linux worker jobs, the native Windows jobs, and the `all checks passed` verdict job — which would otherwise stay queued on the failed pool even after every worker passed — resolves its runner pool through a repository variable, and the switch is split by platform so an outage on one platform does not retarget the other. The three Linux workers and the `all checks passed` verdict (whose `needs` are the required Linux workers and which runs on the `vm-backup` pool) resolve through `DSH_CI_FAILOVER_LINUX`; the native Windows jobs resolve through `DSH_CI_FAILOVER_WINDOWS`. Unset, they default to their hosted pools; selecting `selfhosted` is an explicit operator choice. Set to `selfhosted` by any repository writer, the corresponding jobs retarget onto the in-house self-hosted pool: under `DSH_CI_FAILOVER_LINUX`, the Linux jobs and verdict move onto the `vm-backup` pool, snapshot concurrency drops to the shared-VM bound, and the hosted-path pnpm cache restores are skipped; under `DSH_CI_FAILOVER_WINDOWS`, the native Windows jobs move onto the `dsh-win-ci` pool. Each switch is writer-manageable repository state, not a merge, so it works while every check is red. The in-house pools' readiness is continuously re-proven by the `serial / linux (self-hosted standby)` and `serial / windows (self-hosted standby)` lanes, which run the complete unsharded aggregates on every master push.
 
 
 `ci-master.yml` exempts exactly one event from `cancel-in-progress` (`${{ github.event_name != 'push' }}`), so one master push does not cancel the drill still running from the previous one. Each drill runs its complete unsharded aggregate with one gate worker, which takes longer than the interval between master merges; under unconditional cancellation a drill is superseded before reaching a verdict and the lane yields no readiness evidence for a responder to check.
 `ci-master.yml` exempts exactly one event from `cancel-in-progress` (`${{ github.event_name != 'push' }}`), so one master push does not cancel the drill still running from the previous one. Each drill runs its complete unsharded aggregate with one gate worker, which takes longer than the interval between master merges; under unconditional cancellation a drill is superseded before reaching a verdict and the lane yields no readiness evidence for a responder to check.
 
 
@@ -18,13 +18,17 @@ The exemption is narrower than "a drill always finishes", in two ways. GitHub ke
 
 
 The decision belongs at workflow level because cancellation applies to the whole superseded run: a job-level `concurrency` group does not exempt its job. The negated form is load-bearing rather than cosmetic: naming `pull_request` alone would also stop cancelling `workflow_dispatch`, and each runner benchmark fans out to twelve larger runners for up to fifteen minutes inside this same group on master, so a re-dispatch would queue ahead of a drill instead of replacing a stale measurement. What bounds the cost is that a master push in `ci-master.yml` carries only `wine-apt-cache` and these two drills; the pull-request jobs live in the separate `ci.yml` (which does not see `push`), and the benchmarks are `workflow_dispatch`-gated within `ci-master.yml`. `scripts/ci-workflow.spec.ts` pins that push-reachable set — classifying by exact condition, since a negated event test mentions the event it excludes — so a new push-reachable job cannot quietly start accumulating uncancelled runs.
 The decision belongs at workflow level because cancellation applies to the whole superseded run: a job-level `concurrency` group does not exempt its job. The negated form is load-bearing rather than cosmetic: naming `pull_request` alone would also stop cancelling `workflow_dispatch`, and each runner benchmark fans out to twelve larger runners for up to fifteen minutes inside this same group on master, so a re-dispatch would queue ahead of a drill instead of replacing a stale measurement. What bounds the cost is that a master push in `ci-master.yml` carries only `wine-apt-cache` and these two drills; the pull-request jobs live in the separate `ci.yml` (which does not see `push`), and the benchmarks are `workflow_dispatch`-gated within `ci-master.yml`. `scripts/ci-workflow.spec.ts` pins that push-reachable set — classifying by exact condition, since a negated event test mentions the event it excludes — so a new push-reachable job cannot quietly start accumulating uncancelled runs.
 
 
+### Release rehearsals share the Linux switch
+
+`DSH_CI_FAILOVER_LINUX=selfhosted` also routes the credential-free dependency-layout job and both dsh/vendor pack jobs onto `vm-backup` for eligible same-repository PRs and master pushes. Their [release rehearsal decision](2026-09-06-release-rehearsal-selfhosted.md) owns the stricter event eligibility and hosted manual dispatch. This coupling is intentional: keeping the variable set to save release minutes also keeps the eligible main-CI Linux jobs self-hosted. Clearing it returns both workloads to their hosted targets for subsequent runs; publication stays hosted regardless.
+
 ### What the in-house pool is
 ### What the in-house pool is
 
 
-`vm-backup`: one 64-core VM, six always-on systemd-managed runner instances. Its image must preinstall Playwright Chromium's Linux system packages; CI downloads the lockfile-selected browser but never runs `apt` on this persistent shared host. Check the latest `serial / linux (self-hosted standby)` run before switching: its aggregate includes browser replay, so a green standby verifies both ordinary capacity and this browser prerequisite.
+`vm-backup`: one shared VM with multiple always-on systemd-managed runner instances. Registrations share its CPU, memory, and disk; their count is not a count of independent machines. Its image must preinstall Playwright Chromium's Linux system packages; CI downloads the lockfile-selected browser but never runs `apt` on this persistent shared host. Check the latest `serial / linux (self-hosted standby)` run before switching: its aggregate includes browser replay, so a green standby verifies both ordinary capacity and this browser prerequisite.
 
 
 #### Windows pool
 #### Windows pool
 
 
-`dsh-win-ci`: 32 always-on runner instances (scheduled tasks `GH-Runner-01`…`GH-Runner-32`) on the in-house Windows CI server (one 96-core / 580 GB machine). Labels: `[self-hosted, dsh-win-ci, windows]`. The image must preinstall Node 24, pnpm, Git (with Git Bash on `PATH`, i.e. `C:\Program Files\Git\bin` — the `bash` tool spawns `bash` by name), PowerShell 7, and enable Developer Mode for symlink support. The workspaces and the pnpm store must both live on a ReFS volume (`F:`): the Windows installs pass `--package-import-method=clone` on ReFS, which needs that volume layout and the `@reflink/reflink` native module that the system corepack pnpm carries (see [the Windows ReFS store note](../../archived/process/2026-08-30-windows-refs-store-block-clone-install.md)); a rebuilt runner without this layout fails the Windows build gates with TS6231. Check the latest `serial / windows (self-hosted standby)` run before switching: a green standby verifies the pool can execute `check:ci:windows-complete` end-to-end.
+`dsh-win-ci`: 32 always-on runner instances (scheduled tasks `GH-Runner-01`…`GH-Runner-32`) on the in-house Windows CI server (one 96-core / 580 GB machine). Labels: `[self-hosted, dsh-win-ci, windows]`. The image must preinstall Node 24, pnpm, Git (with Git Bash on `PATH`, i.e. `C:\Program Files\Git\bin` — the `bash` tool spawns `bash` by name), PowerShell 7, and enable Developer Mode for symlink support. The general-purpose Windows workspaces and pnpm store must both live on a ReFS volume (`F:`): those installs pass `--package-import-method=clone` on ReFS, which needs that volume layout and the `@reflink/reflink` native module that the system corepack pnpm carries (see [the Windows ReFS store note](../../archived/process/2026-08-30-windows-refs-store-block-clone-install.md)); a rebuilt runner without this layout fails the Windows build gates with TS6231. Check the latest `serial / windows (self-hosted standby)` run before switching: a green standby verifies the pool can execute `check:ci:windows-complete` end-to-end.
 
 
 ### Switch (any repository writer, ~1 minute, no merge)
 ### Switch (any repository writer, ~1 minute, no merge)
 
 
@@ -32,7 +36,7 @@ The two switches are independent: flip only the one whose platform is degraded.
 
 
 1. Repository **Settings → Secrets and variables → Actions → Variables → New repository variable**: name `DSH_CI_FAILOVER_LINUX` (Linux pool outage) or `DSH_CI_FAILOVER_WINDOWS` (Windows pool outage), value `selfhosted`.
 1. Repository **Settings → Secrets and variables → Actions → Variables → New repository variable**: name `DSH_CI_FAILOVER_LINUX` (Linux pool outage) or `DSH_CI_FAILOVER_WINDOWS` (Windows pool outage), value `selfhosted`.
 2. Retrigger the required jobs so they re-resolve their pool. Jobs already **queued** for the hosted labels do not retarget and cannot be re-run in place, so for the documented indefinite-queue outage, cancel the stuck run and re-run all jobs, or push a new commit; "Re-run failed jobs" only helps once a job has actually failed rather than queued.
 2. Retrigger the required jobs so they re-resolve their pool. Jobs already **queued** for the hosted labels do not retarget and cannot be re-run in place, so for the documented indefinite-queue outage, cancel the stuck run and re-run all jobs, or push a new commit; "Re-run failed jobs" only helps once a job has actually failed rather than queued.
-3. That is the entire switch. Under Linux failover the workflow also drops `DSH_SNAPSHOT_MAX_CONCURRENCY` to 12 for the shared VM and skips the hosted-path pnpm cache restores because the VM's persistent store serves warm installs. Coverage uses the same four single-worker instrumented partitions and two exempt workers on both Linux pools. The Windows switch has no concurrency or cache branches; it only retargets the native Windows job's pool.
+3. That is the entire switch. Under Linux failover the workflow also drops `DSH_SNAPSHOT_MAX_CONCURRENCY` to 12 for the shared VM and skips the hosted-path pnpm cache restores because the VM's persistent store serves warm installs. Coverage uses the same four single-worker instrumented partitions and two exempt workers on both Linux pools. The Windows switch has no concurrency or cache branches; it only retargets the native Windows jobs' pool.
 
 
 #**Dependabot exception.** Both switches' selectors deliberately exclude `dependabot[bot]`: under failover, Dependabot PRs stay queued for the hosted pool rather than executing dependency-supplied code on the persistent VMs. A Dependabot PR that remains queued during an outage is expected behavior, not a failed switch; it completes when the hosted pool recovers.
 #**Dependabot exception.** Both switches' selectors deliberately exclude `dependabot[bot]`: under failover, Dependabot PRs stay queued for the hosted pool rather than executing dependency-supplied code on the persistent VMs. A Dependabot PR that remains queued during an outage is expected behavior, not a failed switch; it completes when the hosted pool recovers.
 
 
@@ -40,12 +44,12 @@ The two switches are independent: flip only the one whose platform is degraded.
 
 
 ## Capacity during failover
 ## Capacity during failover
 
 
-Six always-on instances absorb normal PR traffic (the pool's steady-state load is one serial standby job per master push, so failover capacity is effectively the full pool). If queues still build, register additional instances with an org registration token (org Settings → Actions → Runners → New runner). Clone an existing runner directory **excluding its identity files** — `rsync -a --exclude '.runner*' --exclude '.credentials*' --exclude '_diag' --exclude '_work' <src>/ <dst>/` (the globs also catch `.runner_migrated`/`.credentials_migrated`, which GitHub writes on migrated runners and which equally trigger the already-configured refusal) — then run `config.sh` (copying `.runner`/`.credentials` verbatim makes it refuse with "already configured"), and **start the listener**: `sudo ./svc.sh install ubuntu && sudo ./svc.sh start`. Registration alone leaves the runner offline; only a started service adds capacity. About a minute per instance.
+Capacity includes the master standby, main-CI jobs, and three release-rehearsal jobs for each eligible PR or master push while the Linux switch is set. The release workflows do not cancel running rehearsals when another run arrives, so overlapping refs can add sustained build, pack, and install load. Check current CPU, memory, disk, and queue pressure before extending self-hosted operation; extra registrations on this VM add scheduling slots, not machine resources. Do not infer spare capacity from the standby alone. When host resources permit extra registrations, use an org registration token (org Settings → Actions → Runners → New runner). Clone an existing runner directory **excluding its identity files** — `rsync -a --exclude '.runner*' --exclude '.credentials*' --exclude '_diag' --exclude '_work' <src>/ <dst>/` (the globs also catch `.runner_migrated`/`.credentials_migrated`, which GitHub writes on migrated runners and which equally trigger the already-configured refusal) — then run `config.sh` (copying `.runner`/`.credentials` verbatim makes it refuse with "already configured"), and **start the listener**: `sudo ./svc.sh install ubuntu && sudo ./svc.sh start`. Registration alone leaves the runner offline; a started service adds a scheduling slot, not CPU or memory.
 
 
 
 
 ### Switch back
 ### Switch back
 
 
-Delete the `DSH_CI_FAILOVER_LINUX` or `DSH_CI_FAILOVER_WINDOWS` variable (or set it to anything other than `selfhosted`). New runs resolve back to the hosted enterprise pools. Remove any extra instances that were registered during the incident.
+Delete the `DSH_CI_FAILOVER_LINUX` or `DSH_CI_FAILOVER_WINDOWS` variable (or set it to anything other than `selfhosted`). New runs resolve back to their hosted pools. Remove any extra instances that were registered during the incident.
 
 
 ### Trust boundary
 ### Trust boundary
 
 
@@ -55,7 +59,7 @@ The variables are writer-manageable repository state; a pull request event itsel
 
 
 **Merge a workflow change to switch pools.** Rejected because the outage that motivates the switch is exactly the state in which no PR can merge: the required checks are the ones failing. A repository variable is writer-manageable state that takes effect on re-run without a merge.
 **Merge a workflow change to switch pools.** Rejected because the outage that motivates the switch is exactly the state in which no PR can merge: the required checks are the ones failing. A repository variable is writer-manageable state that takes effect on re-run without a merge.
 
 
-**Keep the self-hosted pool always in the required path.** Rejected because it trades hosted-pool availability for the in-house VM's, moving a single point of failure rather than adding a fallback. The variables keep the hosted pools primary and the self-hosted pools proven, one-action standbys; splitting them by platform means an outage on one platform does not retarget the other.
+**Keep the self-hosted pool always in the required path.** Rejected because it trades hosted-pool availability for the in-house VM's, moving a single point of failure rather than adding a fallback. The unset defaults retain hosted targets and the switches provide a reversible, operator-selected self-hosted path; splitting them by platform means an outage on one platform does not retarget the other.
 
 
 ## Consequences
 ## Consequences
 
 

+ 11 - 7
.agents/notes/implemented/process/2026-07-26-ci-failover-runbook.zh.md

@@ -6,11 +6,11 @@ Status: implemented
 
 
 ## 问题
 ## 问题
 
 
-[CI](../../../../.github/workflows/ci.yml) 中三个必需的 Linux 工作作业(`node 24 / static`、`node 24 / coverage`、`node 24 / snapshots and artifacts`)运行在托管的企业级 32 核池上;聚合它们的必需判定作业(`all checks passed`)运行在标准 `ubuntu-latest` 上;独立的原生 Windows 作业(`windows node 24 / native complete`)运行在托管的 `dsh-windows-2025-16core` 大型运行器上。当企业池发生故障——作业无限排队或企业标签消失——所有开启的拉取请求都无法合并,而"合并一个修复"这一常规恢复手段本身正被那些无法运行的必需检查死锁。**适用范围:两个独立开关,每个平台一个。**`DSH_CI_FAILOVER_LINUX` 恢复企业级 Linux 池故障(三个必需的 Linux 工作作业加 `all checks passed` 判定作业);`DSH_CI_FAILOVER_WINDOWS` 恢复托管 Windows 池故障(原生 Windows 作业)。Linux 池故障无需重定向原生 Windows 作业,反之亦然。判定作业的其余必需依赖(`node-compat`、`python-sdk`、`windows`)按设计留在标准托管运行器上(可移植边界);若更大范围的 GitHub 托管容量故障连标准池一并击倒,这些依赖仍会阻塞 `all checks passed`。因此故障需要一个任何具备仓库写权限的响应者都能在不合并任何代码的情况下触发的开关。
+[CI](../../../../.github/workflows/ci.yml) 中三个必需的 Linux 工作作业(`node 24 / static`、`node 24 / coverage`、`node 24 / snapshots and artifacts`)运行在托管的企业级 32 核池上;聚合它们的必需判定作业(`all checks passed`)运行在标准 `ubuntu-latest` 上;[原生 Windows 作业](2026-08-08-native-windows-pull-request-ci.zh.md)运行在托管的 `dsh-windows-2025-16core` 大型运行器上。当企业池发生故障——作业无限排队或企业标签消失——所有开启的拉取请求都无法合并,而"合并一个修复"这一常规恢复手段本身正被那些无法运行的必需检查死锁。**适用范围:两个独立开关,每个平台一个。**`DSH_CI_FAILOVER_LINUX` 恢复企业级 Linux 池故障(三个必需的 Linux 工作作业加 `all checks passed` 判定作业);`DSH_CI_FAILOVER_WINDOWS` 恢复托管 Windows 池故障(原生 Windows 作业)。Linux 池故障无需重定向 Windows 作业,反之亦然。判定作业的其余必需依赖(`node-24-bench`、`node-compat`、`python-sdk`、`windows`)按设计留在标准托管运行器上(可移植边界);若更大范围的 GitHub 托管容量故障连标准池一并击倒,这些依赖仍会阻塞 `all checks passed`。因此故障需要一个任何具备仓库写权限的响应者都能在不合并任何代码的情况下触发的开关。
 
 
 ## 决策
 ## 决策
 
 
-三个必需的 Linux 工作作业、独立的原生 Windows 作业,以及 `all checks passed` 判定作业(若不随切换,即使全部工作作业通过,它仍会滞留在故障池的队列中)——各自通过仓库变量解析运行器池,且开关按平台拆分,使一个平台的故障不会重定向另一个平台。三个 Linux 工作作业与 `all checks passed` 判定作业(其 `needs` 是必需的 Linux 工作作业,且运行在 `vm-backup` 池上)通过 `DSH_CI_FAILOVER_LINUX` 解析;原生 Windows 作业通过 `DSH_CI_FAILOVER_WINDOWS` 解析。变量不存在(正常)时它们运行在托管企业池上;由任何具备写权限的协作者设为 `selfhosted` 时,对应作业切换到公司自有的自托管池:`DSH_CI_FAILOVER_LINUX` 下,Linux 作业与判定作业切到 `vm-backup` 池,快照并发降到共享虚拟机上限,并跳过托管路径的 pnpm 缓存恢复;`DSH_CI_FAILOVER_WINDOWS` 下,原生 Windows 作业切到 `dsh-win-ci` 池。每个开关都是写者可管理的仓库状态而非一次合并,因此在所有检查都是红色时仍然有效。自有池的就绪状态由 `serial / linux (self-hosted standby)` 与 `serial / windows (self-hosted standby)` 通道持续验证——每次 master 推送都在其上运行完整的未分片聚合流程。
+三个必需的 Linux 工作作业、原生 Windows 作业,以及 `all checks passed` 判定作业(若不随切换,即使全部工作作业通过,它仍会滞留在故障池的队列中)——各自通过仓库变量解析运行器池,且开关按平台拆分,使一个平台的故障不会重定向另一个平台。三个 Linux 工作作业与 `all checks passed` 判定作业(其 `needs` 是必需的 Linux 工作作业,且运行在 `vm-backup` 池上)通过 `DSH_CI_FAILOVER_LINUX` 解析;原生 Windows 作业通过 `DSH_CI_FAILOVER_WINDOWS` 解析。未设置变量时默认使用各自的托管池;选择 `selfhosted` 是运维人员的明确操作;由任何具备写权限的协作者设为 `selfhosted` 时,对应作业切换到公司自有的自托管池:`DSH_CI_FAILOVER_LINUX` 下,Linux 作业与判定作业切到 `vm-backup` 池,快照并发降到共享虚拟机上限,并跳过托管路径的 pnpm 缓存恢复;`DSH_CI_FAILOVER_WINDOWS` 下,原生 Windows 作业切到 `dsh-win-ci` 池。每个开关都是写者可管理的仓库状态而非一次合并,因此在所有检查都是红色时仍然有效。自有池的就绪状态由 `serial / linux (self-hosted standby)` 与 `serial / windows (self-hosted standby)` 通道持续验证——每次 master 推送都在其上运行完整的未分片聚合流程。
 
 
 `ci-master.yml` 只豁免一个事件不做取消(`${{ github.event_name != 'push' }}`),因此一次 master 推送不会取消上一次推送留下的、仍在运行的演练。每次演练以单门禁工作进程执行完整的未分片聚合流程,耗时长于 master 合并的间隔;在无条件取消下,演练会在得出结论前被后续运行取代,该通道无法产出供响应者查看的就绪证据。
 `ci-master.yml` 只豁免一个事件不做取消(`${{ github.event_name != 'push' }}`),因此一次 master 推送不会取消上一次推送留下的、仍在运行的演练。每次演练以单门禁工作进程执行完整的未分片聚合流程,耗时长于 master 合并的间隔;在无条件取消下,演练会在得出结论前被后续运行取代,该通道无法产出供响应者查看的就绪证据。
 
 
@@ -18,13 +18,17 @@ Status: implemented
 
 
 这个决定必须放在工作流级:取消作用于被取代的整个运行,作业级 `concurrency` 组并不能豁免其所属作业。采用否定式写法而非仅指名 `pull_request`,是有实质作用的:后者会连 `workflow_dispatch` 一起停止取消,而每次运行器基准测试会在 master 上的同一并发组内同时占用 12 台大规格运行器、最长 15 分钟,届时重复派发会排在演练之前,而不是替换掉已过时的测量。成本之所以可控,是因为 `ci-master.yml` 中一次 master 推送只承载 `wine-apt-cache` 和这两条演练;拉取请求作业位于独立的 `ci.yml`(不监听 `push`),而基准测试在 `ci-master.yml` 内受 `workflow_dispatch` 门控。`scripts/ci-workflow.spec.ts` 会锁定这个推送可达集合——按条件精确匹配,因为否定式事件判断会包含它所排除的事件名——使新的推送可达作业无法悄悄开始累积未取消的运行。
 这个决定必须放在工作流级:取消作用于被取代的整个运行,作业级 `concurrency` 组并不能豁免其所属作业。采用否定式写法而非仅指名 `pull_request`,是有实质作用的:后者会连 `workflow_dispatch` 一起停止取消,而每次运行器基准测试会在 master 上的同一并发组内同时占用 12 台大规格运行器、最长 15 分钟,届时重复派发会排在演练之前,而不是替换掉已过时的测量。成本之所以可控,是因为 `ci-master.yml` 中一次 master 推送只承载 `wine-apt-cache` 和这两条演练;拉取请求作业位于独立的 `ci.yml`(不监听 `push`),而基准测试在 `ci-master.yml` 内受 `workflow_dispatch` 门控。`scripts/ci-workflow.spec.ts` 会锁定这个推送可达集合——按条件精确匹配,因为否定式事件判断会包含它所排除的事件名——使新的推送可达作业无法悄悄开始累积未取消的运行。
 
 
+### 发布演练共用 Linux 开关
+
+`DSH_CI_FAILOVER_LINUX=selfhosted` 还会将符合条件的同仓库 PR 和 master 推送中的无凭据依赖布局作业与 dsh/vendor 两个打包作业路由到 `vm-backup`。[发布演练决策](2026-09-06-release-rehearsal-selfhosted.zh.md) 负责更严格的事件准入规则及保留托管的手动触发。这种耦合是有意的:持续设置变量来节省发布分钟,也会让符合条件的主 CI Linux 作业持续使用自托管。清除变量会让两类负载的后续运行返回各自的托管目标;发布操作始终保留托管。
+
 ### 自有池是什么
 ### 自有池是什么
 
 
-`vm-backup`:一台 64 核虚拟机,6 个常驻 systemd 管理的运行器实例。其镜像必须预装 Playwright Chromium 的 Linux 系统软件包;CI 会下载锁文件选定的浏览器,但绝不在这台持久化共享主机上运行 `apt`。切换前先看 `serial / linux (self-hosted standby)` 最近一次运行:其聚合流程包含浏览器回放,因此绿色热备同时验证常规容量和这项浏览器先决条件。
+`vm-backup`:一台共享虚拟机,运行多个常驻 systemd 管理的运行器实例。注册实例共享 CPU、内存和磁盘;实例数量不代表独立机器数量。其镜像必须预装 Playwright Chromium 的 Linux 系统软件包;CI 会下载锁文件选定的浏览器,但绝不在这台持久化共享主机上运行 `apt`。切换前先看 `serial / linux (self-hosted standby)` 最近一次运行:其聚合流程包含浏览器回放,因此绿色热备同时验证常规容量和这项浏览器先决条件。
 
 
 #### Windows 池
 #### Windows 池
 
 
-`dsh-win-ci`:公司内部 Windows CI 服务器(一台 96 核 / 580 GB 机器)上 32 个常驻运行器实例(计划任务 `GH-Runner-01`…`GH-Runner-32`)。标签:`[self-hosted, dsh-win-ci, windows]`。镜像必须预装 Node 24、pnpm、Git(Git Bash 在 `PATH` 上,即 `C:\Program Files\Git\bin`——`bash` 工具按名称 spawn `bash`)、PowerShell 7,并为符号链接支持启用开发人员模式。工作区与 pnpm store 必须都位于 ReFS 卷(`F:`)上:Windows 安装步骤在 ReFS 上传递 `--package-import-method=clone`,这需要该卷布局以及系统 corepack pnpm 携带的 `@reflink/reflink` 原生模块(见 [Windows ReFS store note](../../archived/process/2026-08-30-windows-refs-store-block-clone-install.md));没有此布局的重建运行器会在 Windows 构建门禁阶段以 TS6231 失败。切换前先看 `serial / windows (self-hosted standby)` 最近一次运行:绿色热备验证该池能端到端执行 `check:ci:windows-complete`。
+`dsh-win-ci`:公司内部 Windows CI 服务器(一台 96 核 / 580 GB 机器)上 32 个常驻运行器实例(计划任务 `GH-Runner-01`…`GH-Runner-32`)。标签:`[self-hosted, dsh-win-ci, windows]`。镜像必须预装 Node 24、pnpm、Git(Git Bash 在 `PATH` 上,即 `C:\Program Files\Git\bin`——`bash` 工具按名称 spawn `bash`)、PowerShell 7,并为符号链接支持启用开发人员模式。通用 Windows 通道的工作区与 pnpm store 必须都位于 ReFS 卷(`F:`)上:这些安装步骤在 ReFS 上传递 `--package-import-method=clone`,这需要该卷布局以及系统 corepack pnpm 携带的 `@reflink/reflink` 原生模块(见 [Windows ReFS store note](../../archived/process/2026-08-30-windows-refs-store-block-clone-install.md));没有此布局的重建运行器会在 Windows 构建门禁阶段以 TS6231 失败。切换前先看 `serial / windows (self-hosted standby)` 最近一次运行:绿色热备验证该池能端到端执行 `check:ci:windows-complete`。
 
 
 ### 切换步骤(任何具备写权限的协作者,约 1 分钟,无需合并)
 ### 切换步骤(任何具备写权限的协作者,约 1 分钟,无需合并)
 
 
@@ -40,12 +44,12 @@ Status: implemented
 
 
 ## 切换期间的容量
 ## 切换期间的容量
 
 
-6 个常驻实例可承接正常 PR 流量(该池平时唯一的稳态负载是每次 master 推送一个串行热备作业,故障切换时几乎全池可用)。若仍出现排队,用组织级注册 token(组织 Settings → Actions → Runners → New runner)追加注册实例。复制现有 runner 目录时**必须排除身份文件**——`rsync -a --exclude '.runner*' --exclude '.credentials*' --exclude '_diag' --exclude '_work' <src>/ <dst>/`(通配同时排除 `.runner_migrated`/`.credentials_migrated`——GitHub 会在迁移过的运行器上写入这些文件,它们同样会触发 already-configured 拒绝)——再跑 `config.sh`(原样拷贝 `.runner`/`.credentials` 会使其以 "already configured" 拒绝),然后**启动监听器**:`sudo ./svc.sh install ubuntu && sudo ./svc.sh start`。仅注册不会上线;只有启动了服务的 runner 才会增加容量。每个约一分钟。
+Linux 开关启用期间,容量需覆盖 master 热备、主 CI 作业,以及每个符合条件的 PR 或 master 推送的三个发布演练作业。发布工作流不会因为新运行到来而取消正在执行的演练,因此不同引用的重叠运行会增加持续的构建、打包和安装负载。延长自托管运行前,检查当前 CPU、内存、磁盘和队列压力;同一虚拟机上新增注册只增加调度槽位,不增加机器资源。不能只依据热备负载推断空闲容量。主机资源允许增加注册实例时,使用组织级注册 token(组织 Settings → Actions → Runners → New runner)。复制现有 runner 目录时**必须排除身份文件**——`rsync -a --exclude '.runner*' --exclude '.credentials*' --exclude '_diag' --exclude '_work' <src>/ <dst>/`(通配同时排除 `.runner_migrated`/`.credentials_migrated`——GitHub 会在迁移过的运行器上写入这些文件,它们同样会触发 already-configured 拒绝)——再跑 `config.sh`(原样拷贝 `.runner`/`.credentials` 会使其以 "already configured" 拒绝),然后**启动监听器**:`sudo ./svc.sh install ubuntu && sudo ./svc.sh start`。仅注册不会上线;启动服务增加的是调度槽位,而非 CPU 或内存。
 
 
 
 
 ### 切回
 ### 切回
 
 
-删除 `DSH_CI_FAILOVER_LINUX` 或 `DSH_CI_FAILOVER_WINDOWS` 变量(或改为 `selfhosted` 以外的任何值),新的运行即解析回托管企业池。若故障期间追加注册过实例,将其移除。
+删除 `DSH_CI_FAILOVER_LINUX` 或 `DSH_CI_FAILOVER_WINDOWS` 变量(或改为 `selfhosted` 以外的任何值),新的运行即解析回各自的托管池。若故障期间追加注册过实例,将其移除。
 
 
 ### 信任边界
 ### 信任边界
 
 
@@ -55,7 +59,7 @@ Status: implemented
 
 
 **通过合并一次工作流改动来切换池。** 否决,因为触发切换的故障状态恰恰是任何 PR 都无法合并的状态:必需检查正是失败的那些。仓库变量是写者可管理的状态,重跑即生效,无需合并。
 **通过合并一次工作流改动来切换池。** 否决,因为触发切换的故障状态恰恰是任何 PR 都无法合并的状态:必需检查正是失败的那些。仓库变量是写者可管理的状态,重跑即生效,无需合并。
 
 
-**让自托管池长期处于必需路径中。** 否决,因为这是拿托管池的可用性去换自有虚拟机的可用性,只是搬移了单点故障而非增加回退。这些变量让托管池保持主路径,自托管池作为一个经过验证、一步即可启用的热备;按平台拆分意味着一个平台的故障不会重定向另一个平台。
+**让自托管池长期处于必需路径中。** 否决,因为这是拿托管池的可用性去换自有虚拟机的可用性,只是搬移了单点故障而非增加回退。未设置变量时默认保留托管目标,开关提供由运维人员选择、可逆的自托管路径;按平台拆分意味着一个平台的故障不会重定向另一个平台。
 
 
 ## 后果
 ## 后果
 
 

+ 6 - 0
.agents/notes/implemented/process/2026-09-03-semantic-issue-templates-and-policy.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-03-semantic-issue-templates-and-policy.md
+2026-09-03-semantic-issue-templates-and-policy.md: 96b1981b21e901841d88b09d68a0e7e31ca140a2
+2026-09-03-semantic-issue-templates-and-policy.zh.md: 706cd59047fe5121964183bbc27efc8961a493ed

+ 39 - 0
.agents/notes/implemented/process/2026-09-03-semantic-issue-templates-and-policy.md

@@ -0,0 +1,39 @@
+# Agent Note: Semantic Issue templates and presentation-neutral policy
+
+Status: implemented
+
+English | [中文](2026-09-03-semantic-issue-templates-and-policy.zh.md)
+
+## Problem
+
+Issue and pull-request templates mixed intake questions with review evidence and hid their complete contents in `details` elements. Unused frontmatter and separate Idea and Research templates added choices without changing how the repository planned the work.
+
+Issue policy also treated Markdown presentation as repository metadata. Requirements for `details` elements, a 50-unit visible body, Chinese titles, title metadata prefixes, and an `Owner:` body line produced failures without identifying a missing semantic decision.
+
+## Decision
+
+Issue templates cover Bug, Feature, and Task. Bug asks for a summary, reproduction, current behavior, expected behavior, and environment. Feature asks for motivation and behavior. Task asks for a summary and deliverables. Idea and Research belong in Task unless a future decision gives them distinct lifecycle behavior.
+
+Issue-template frontmatter contains only `name`, `about`, and `type`. Markdown headings define the hierarchy, and HTML comments explain what belongs under each heading.
+
+The pull-request template contains `Motivation`; a `Changes` section with adjacent placeholders for public-interface and behavior changes; and `Testing` entries that show each method directly and place its proof in a local `details` element.
+
+Issue policy does not inspect `details` presentation, visible-body length, title language, title prefixes, or body ownership lines. Every other policy check, warning, lifecycle operation, workflow trigger, and pull-request enforcement exemption retains its existing behavior. This decision adds no metadata repair, Issue classification, data migration, or workflow capability.
+
+## Verification
+
+[Issue-management tests](../../../../.github/issue-management/policy.test.mjs) pin the template inventory and headings, the pull-request testing structure, and acceptance of titles, bodies, and assignee states that differ only in presentation.
+
+## Alternatives considered
+
+**Keep Idea and Research templates.** Their forms did not establish lifecycle or policy behavior distinct from Task, so separate entry points increased choice without preserving a meaningful type distinction.
+
+**Keep presentation rules as warnings.** These rules could fail otherwise actionable Issues and could not establish whether the requested work, expected behavior, or deliverables were clear.
+
+**Add automatic metadata repair or Issue classification.** Those behaviors require new mutation rules, permissions, failure handling, and operational evidence. They remain separate decisions rather than accompanying a policy simplification.
+
+## Consequences
+
+Contributors see shorter forms whose headings match the information needed at Issue intake and pull-request review. Policy failures remain focused on the existing semantic metadata checks.
+
+The policy does not rewrite legacy labels, choose missing Issue Types, synchronize Priority, or migrate existing repository data. Any future automation for those operations needs its own decision and review scope.

+ 39 - 0
.agents/notes/implemented/process/2026-09-03-semantic-issue-templates-and-policy.zh.md

@@ -0,0 +1,39 @@
+# Agent Note: 语义化 Issue template 与不检查展示形式的 policy
+
+Status: implemented
+
+[English](2026-09-03-semantic-issue-templates-and-policy.md) | 中文
+
+## 问题
+
+Issue 与拉取请求(Pull Request,PR)template 把信息收集问题与评审证据混在一起,并用 `details` 元素折叠全部内容。未使用的 frontmatter 以及独立的 Idea 和 Research template 增加了选择,却没有改变仓库规划这些工作的方式。
+
+Issue policy 还把 Markdown 展示方式当成仓库 metadata。对 `details` 元素、50 单位可见正文、中文标题、标题 metadata 前缀和正文 `Owner:` 行的要求会产生失败,却不能指出缺少了哪项语义决策。
+
+## 决策
+
+Issue template 只覆盖 Bug、Feature 和 Task。Bug 收集摘要、复现方式、当前行为、预期行为和环境。Feature 收集动机和行为。Task 收集摘要和交付物。除非未来的决策为 Idea 与 Research 定义不同的生命周期行为,否则它们属于 Task。
+
+Issue template frontmatter 只含 `name`、`about` 和 `type`。Markdown 标题定义信息层级,HTML 注释说明每个标题下应填写的内容。
+
+PR template 包含 `Motivation`、在 `Changes` 中相邻排列的公共接口与行为变化占位说明,以及直接展示每项测试方法并在局部 `details` 元素中放置对应证明的 `Testing` 条目。
+
+Issue policy 不检查 `details` 展示形式、可见正文长度、标题语言、标题前缀或正文 ownership 行。其他所有 policy 检查、警告、生命周期操作、workflow 触发条件和 PR 强制范围豁免均保持既有行为。这项决策不增加 metadata 自动修复、Issue 分类、数据迁移或 workflow 能力。
+
+## 验证
+
+[Issue 管理测试](../../../../.github/issue-management/policy.test.mjs)固定 template 清单与标题、PR 测试结构,并验证仅展示形式不同的标题、正文和 assignee 状态可以通过。
+
+## 考虑过的替代方案
+
+**保留 Idea 和 Research template。** 它们的表单没有建立区别于 Task 的生命周期或 policy 行为,因此独立入口只会增加选择,而不能保留有意义的 Type 区别。
+
+**把展示规则保留为警告。** 这些规则会让原本可执行的 Issue 失败,也不能确定需求工作、预期行为或交付物是否清晰。
+
+**增加 metadata 自动修复或 Issue 分类。** 这些行为需要新的修改规则、权限、失败处理和运行证据。它们应作为独立决策,而不随 policy 简化一同引入。
+
+## 后果
+
+贡献者会看到更短的表单,其标题分别对应 Issue 信息收集和 PR 评审所需的信息。Policy 失败会继续聚焦现有的语义 metadata 检查。
+
+Policy 不会改写旧标签、选择缺失的 Issue Type、同步 Priority 或迁移现有仓库数据。这些操作未来如需自动化,必须单独决策和评审。

+ 6 - 0
.agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.md
+2026-09-06-evidence-driven-performance-skill.md: 5b15cce1adbd7ff47e5668f7332cba8d1b59e5fe
+2026-09-06-evidence-driven-performance-skill.zh.md: c1fbbd76740badd87ed0a95218e2c4082f2c0e8b

+ 48 - 0
.agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.md

@@ -0,0 +1,48 @@
+# Agent Note: Evidence-driven performance optimization workflow
+
+Status: implemented
+
+English | [中文](2026-09-06-evidence-driven-performance-skill.zh.md)
+
+## Problem
+
+Performance work can improve an isolated phase while moving cost into another phase, retaining more data, or skipping required behavior. Historical PR descriptions also retain abandoned implementations and estimates, so copying their apparent solution can restore a rejected design instead of addressing a current bottleneck.
+
+## Decision
+
+The [dsh-speed-up-perf skill](../../../skills/dsh-speed-up-perf/SKILL.md) guides broad surveys toward bounded, measured user paths. It combines focused attribution with independently timed backend and browser endpoints, synthetic workload distributions, comparable cold/warm and retained-memory conditions, and negative controls for tightened budgets. The historical evidence below distinguishes merged implementations, superseded proposals, author-reported measurements, and estimates.
+
+The workflow requires behavior evidence independently of timing: model-visible logs, durable generation and publication rules, stream ordering, cancellation, and disposal remain obligations. Authorized private corpus inspection yields only aggregate workload inspiration; committed inputs and published artifacts contain synthetic material. Optimization PRs carry their tighter budgets, while a preceding benchmark layer can protect the measured baseline and remain independently mergeable.
+
+The [Session-opening performance-gate decision](../testing/2026-09-04-session-open-performance-gate.md) retains ownership of lane mechanics and calibration. The [simplification skill](../../../skills/dsh-find-simplifications/SKILL.md) retains ownership of deletion-oriented surveys. Neither is superseded: this workflow adds performance-specific candidate selection, measurement comparability, and stopping criteria rather than replacing their decisions.
+
+## Historical evidence
+
+These are author-reported historical measurements, not benchmarks rerun for this workflow. Final merged diffs and owning source take precedence over original PR descriptions. The rejected intermediate proposal is retained only to explain why identity registries are not a general prescription.
+
+| Evidence | Measured path and result | Reusable lesson |
+|---|---|---|
+| [#3535](https://github.com/deepseek-harness/deepseek-harness/pull/3535), merged | The [final benchmark design](https://github.com/deepseek-harness/deepseek-harness/pull/3535#issuecomment-5552779119) reports a 4,394 ms first-open negative control against 550 ms, first-history 4,452 against 550, resume 4,333 against 450, and 128 MB heap failures. Client fold: 123.9 ms / 10.84× against 40 ms / 3.125×. | Built-JS user-path gates and positive/negative controls matter more than an earlier PR-body design. |
+| [#3536](https://github.com/deepseek-harness/deepseek-harness/pull/3536), closed unmerged | Repeated snapshot/freeze work occupied about 70% of profiled CPU; synthetic open improved from 4,734–4,921 to 707–823 ms. | Streaming migration superseded this identity-registry proposal. Do not revive it without current ownership evidence. |
+| [#3585](https://github.com/deepseek-harness/deepseek-harness/pull/3585), merged | Historical physical decode: 7.527 s / 7,219 MB peak RSS to 1.467 s / 908 MB; streaming migration with serial publication: 6.241 s, 2.107 GB peak, 477 MB retained. Settled 500,000-delta Client fold: 3.2 ms. | Keep representations compact across consumers; bound intermediate state. Attribution estimates overlap and cannot be added. |
+| [#3586](https://github.com/deepseek-harness/deepseek-harness/pull/3586), merged | Current-v2 opening snapshot: 2,011.4→1,027.9 ms; restore: 598.5→16 ms; retained heap: 1,025.3→478.7 MB. | Separate read-only preparation from awaited write publication; share immutable ownership with revision-keyed preparation and caller-local cancellation. |
+| [#3537](https://github.com/deepseek-harness/deepseek-harness/pull/3537), merged | Synthetic 200-turn projection: 28→5.4 ms; total: 76.9→50 ms; peak RSS: 137.2→94.9 MB. | Read stats, usage, text and image references per compact record. Expanded-stream caching retains unnecessary representation cost. Chat/Trajectory belong to the preceding migration change. |
+| [#2587](https://github.com/deepseek-harness/deepseek-harness/pull/2587), merged | Historical 416,756 events represented by 696 records: client history 4,682→276 ms; sampled additional V8 peak 612.5→199.4 MB. | Preserve compactness through validation and folding; [baseline review](https://github.com/deepseek-harness/deepseek-harness/pull/2587#discussion_r3803082730) requires equal validation and retained output, not parse-and-discard. |
+| [#3331](https://github.com/deepseek-harness/deepseek-harness/pull/3331), merged | 10,000 collapsed tool rows: 22.5→7.5 ms, retained 12.2→1.6 MiB; inactive Trajectory flushes: 4,082→15.5 ms. | Defer unused parsing and materialization; first activation and retained Context still cost work. |
+| [#3391](https://github.com/deepseek-harness/deepseek-harness/pull/3391) and [#3383](https://github.com/deepseek-harness/deepseek-harness/pull/3383), merged | Narrow subscriptions, stable identities, batched publication, and viewport-triggered highlighting. The 10,000-node timing table is estimated, not browser measurement. | Deferral is not virtualization: visited token DOM remains retained. |
+| [#3292](https://github.com/deepseek-harness/deepseek-harness/pull/3292), merged | Two-million-item FIFO drain: 9.656 ms median, excluding enqueue. | A deque removes shift copying, not queue admission or backpressure obligations. |
+| [#1161](https://github.com/deepseek-harness/deepseek-harness/pull/1161), merged | Keyless 100,000-chunk browser stress at 128 chunks per 16 ms. | [Producer catch-up](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970161) and [final heartbeat stalls](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970162) can distort measurements; scheduled events are not trusted keyboard/pointer input. |
+
+The [cancellation review](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940578092), [source-revision review](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940569241), and [typed-reader review](https://github.com/deepseek-harness/deepseek-harness/pull/3537#discussion_r3942974015) illustrate why removing repeated work does not authorize deleting validation or publication obligations. A [standby-runner review](https://github.com/deepseek-harness/deepseek-harness/pull/3535#discussion_r3927945561) distinguishes a dedicated job from an isolated physical host.
+
+## Alternatives considered
+
+**Optimize suspicious code before measuring.** Rejected because local complexity does not identify dominant user cost and cannot establish improvement or regression protection.
+
+**Treat historical speedups as reusable prescriptions.** Rejected because representation, ownership, and lifecycle requirements change. Historical evidence generates hypotheses; current production paths and fresh measurements decide whether a change applies.
+
+**Use only microbenchmarks or only end-to-end timing.** Rejected because isolated phases can omit moved work, while aggregate timing alone cannot locate its cause. Both are required at the scope appropriate to the selected problem.
+
+## Consequences
+
+The skill adds no runtime behavior, benchmark implementation, or new CI policy. Its validation is document/link consistency and skill metadata; each future optimization supplies executable measurements and functional evidence at its owner. The finite scenario/fix scope prevents a broad performance request from becoming an unrelated architectural rewrite.

+ 48 - 0
.agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.zh.md

@@ -0,0 +1,48 @@
+# Agent Note: 以证据驱动的性能优化工作流
+
+Status: implemented
+
+[English](2026-09-06-evidence-driven-performance-skill.md) | 中文
+
+## 问题
+
+性能工作可能改善某个独立阶段,却把成本转移到另一阶段、保留更多数据,或跳过必要行为。历史 PR(Pull Request)描述也可能保留已放弃的实现和估计值,因此照搬其表面方案可能恢复已否决的设计,而不是解决当前瓶颈。
+
+## 决定
+
+[dsh-speed-up-perf skill](../../../skills/dsh-speed-up-perf/SKILL.md)(技能)引导广泛调查收敛到范围明确、可测量的用户路径。它结合聚焦的成本归因与独立计时的后端和浏览器端点、合成负载分布、可比较的冷态/热态与保留内存条件,以及收紧预算的负向对照。下方历史证据区分已合并实现、已被替代的提案、作者报告的测量值和估计值。
+
+该工作流要求独立于计时的行为证据:模型可见日志、持久化代际和发布规则、流顺序、取消及 dispose(资源释放)仍是必须满足的要求。获授权的私有语料检查仅提供聚合负载启发;提交的输入和发布的产物包含合成材料。优化 PR 携带收紧后的预算,而前置基准测试层可以保护已测基线并保持独立可合并。
+
+[会话打开性能门禁决策](../testing/2026-09-04-session-open-performance-gate.zh.md)继续负责测试通道机制与校准。[简化 skill](../../../skills/dsh-find-simplifications/SKILL.md)继续负责以删除为目标的调查。两者均未被替代:本工作流增加面向性能的候选选择、测量可比性和停止条件,而不替换它们的决策。
+
+## 历史证据
+
+这些是作者报告的历史测量,并非为本工作流重新运行的基准测试。最终合并差异与所属源码优先于最初 PR 描述。保留已否决的中间提案,仅用于解释为何身份注册表不是通用处方。
+
+| 证据 | 测量路径与结果 | 可复用经验 |
+|---|---|---|
+| [#3535](https://github.com/deepseek-harness/deepseek-harness/pull/3535),已合并 | [最终基准设计](https://github.com/deepseek-harness/deepseek-harness/pull/3535#issuecomment-5552779119)报告首次打开负向对照 4,394 ms,预算 550 ms;首屏历史 4,452,预算 550;恢复 4,333,预算 450;128 MB 堆检查失败。Client fold:123.9 ms / 10.84×,预算 40 ms / 3.125×。 | built-JS 用户路径门禁与正/负向对照比早期 PR 正文设计更重要。 |
+| [#3536](https://github.com/deepseek-harness/deepseek-harness/pull/3536),关闭未合并 | 重复 snapshot/freeze 工作占采样 CPU 的约 70%;合成打开从 4,734–4,921 改善为 707–823 ms。 | 流式迁移替代了该身份注册表提案。没有当前所有权证据时,不恢复它。 |
+| [#3585](https://github.com/deepseek-harness/deepseek-harness/pull/3585),已合并 | 历史物理解码:7.527 s / 7,219 MB 峰值 RSS 降至 1.467 s / 908 MB;流式迁移加串行发布:6.241 s,2.107 GB 峰值,477 MB 保留。已结算的 500,000-delta Client fold:3.2 ms。 | 跨消费者保持紧凑表示;限制中间状态。归因估计重叠,不能相加。 |
+| [#3586](https://github.com/deepseek-harness/deepseek-harness/pull/3586),已合并 | 当前 v2 打开快照:2,011.4→1,027.9 ms;恢复:598.5→16 ms;保留堆:1,025.3→478.7 MB。 | 分离只读准备与必须等待的写发布;通过按修订号共享准备和调用方局部取消共享不可变所有权。 |
+| [#3537](https://github.com/deepseek-harness/deepseek-harness/pull/3537),已合并 | 合成 200 轮投影:28→5.4 ms;总计:76.9→50 ms;峰值 RSS:137.2→94.9 MB。 | 按紧凑记录读取统计、usage、文本和图像引用。展开流缓存保留不必要的表示成本。Chat/Trajectory 属于前置迁移改动。 |
+| [#2587](https://github.com/deepseek-harness/deepseek-harness/pull/2587),已合并 | 历史 416,756 事件由 696 记录表示:Client 历史 4,682→276 ms;采样额外 V8 峰值 612.5→199.4 MB。 | 验证和折叠过程保持紧凑;[基线审查](https://github.com/deepseek-harness/deepseek-harness/pull/2587#discussion_r3803082730)要求相同验证与保留输出,而不是解析后丢弃。 |
+| [#3331](https://github.com/deepseek-harness/deepseek-harness/pull/3331),已合并 | 10,000 个折叠工具行:22.5→7.5 ms,保留 12.2→1.6 MiB;非活动 Trajectory 刷新:4,082→15.5 ms。 | 延迟未使用的解析和实体化;首次激活与保留 Context 仍有成本。 |
+| [#3391](https://github.com/deepseek-harness/deepseek-harness/pull/3391) 和 [#3383](https://github.com/deepseek-harness/deepseek-harness/pull/3383),已合并 | 缩小订阅范围、稳定身份、批量发布和视口触发高亮。10,000 节点计时表是估计,不是浏览器测量。 | 延迟不等于虚拟化:访问过的 token DOM 仍被保留。 |
+| [#3292](https://github.com/deepseek-harness/deepseek-harness/pull/3292),已合并 | 两百万条 FIFO 排空:中位数 9.656 ms,不含入队。 | deque 删除 shift 复制,不删除队列准入或背压义务。 |
+| [#1161](https://github.com/deepseek-harness/deepseek-harness/pull/1161),已合并 | 无密钥的 100,000-chunk 浏览器压力测试,每 16 ms 推送 128 个 chunk。 | [生产者追赶](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970161)和[最后一次心跳停顿](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970162)可能扭曲测量;定时派发事件不是真实键盘/指针输入。 |
+
+[取消审查](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940578092)、[源修订审查](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940569241)和[类型化读取器审查](https://github.com/deepseek-harness/deepseek-harness/pull/3537#discussion_r3942974015)说明删除重复工作不等于允许删除验证或发布义务。[备用 runner 审查](https://github.com/deepseek-harness/deepseek-harness/pull/3535#discussion_r3927945561)区分独立 job 与隔离的物理主机。
+
+## 考虑过的替代方案
+
+**先优化可疑代码,再测量。** 否决,因为局部复杂度不能确定主要用户成本,也无法证明改善或防止回归。
+
+**把历史加速方案当作可复用处方。** 否决,因为表示方式、所有权和生命周期要求会变化。历史证据用于产生假设;当前生产路径与新的测量决定改动是否适用。
+
+**只使用微基准测试,或只使用端到端计时。** 否决,因为独立阶段可能遗漏被转移的工作,而总计时无法定位原因。两者都需要在与所选问题相符的范围内使用。
+
+## 后果
+
+该 skill 不增加运行时行为、基准测试实现或新的 CI 策略。其验证涵盖文档/链接一致性和 skill 元数据;后续每项优化在其所属位置提供可执行测量与功能证据。有限的场景/修复范围防止广泛性能请求演变成无关的架构重写。

+ 6 - 0
.agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.md
+2026-09-06-preview-hosted-runner-sizing.md: 87298e94f11aa7e483afde31e0963a56f523febc
+2026-09-06-preview-hosted-runner-sizing.zh.md: 285b21d60d755e76db582e4e55c5cde913f7fc2e

+ 46 - 0
.agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.md

@@ -0,0 +1,46 @@
+# Agent Note: Measured GitHub-hosted PR preview sizing
+
+Status: implemented
+
+English | [中文](2026-09-06-preview-hosted-runner-sizing.zh.md)
+
+## Problem
+
+PR previews build the full workspace and browser-worker VFS image. A lower per-minute runner price does not guarantee lower job cost because GitHub rounds each job upward to whole minutes. Moving previews to persistent self-hosted machines also changes isolation and is outside this decision.
+
+## Decision
+
+The [preview workflow](../../../../.github/workflows/build-preview-cloudflare.yml) uses standard GitHub-hosted `ubuntu-24.04`. Build, cache, deployment, protected-image verification, and comment semantics remain unchanged. The [sizing reference](../../../../.github/preview-sizing/README.md) owns comparison requirements. The separate CI [failover runbook](2026-07-26-ci-failover-runbook.md) retains its independent runner-switch decision; previews do not use those switches.
+
+### Measurements
+
+[Experiment 34012729982](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34012729982) succeeds for all eight size/cache combinations plus one cache seed. Every measured job checks out SHA `9149d7e7ef945b5601711badd3cf63d58ab384f5`, uses Node 24.19.0 and pnpm 11.7.0, and executes immutable install, full workspace build, preview/VFS packing, and local upload shaping with gzip integrity verification. Warm jobs restore one exact run-private pnpm cache; cold jobs skip restoration but contain pnpm bootstrap files. No compiled outputs are restored.
+
+| Runner | Cold / warm job seconds | Rounded minutes each | USD each | Workspace seconds cold / warm | Preview seconds cold / warm |
+|---|---:|---:|---:|---:|---:|
+| standard, 2 vCPU | 202 / 203 | 4 | 0.024 | 138.92 / 147.21 | 12.65 / 12.94 |
+| larger, 4 vCPU | 177 / 162 | 3 | 0.036 | 124.21 / 114.86 | 10.77 / 9.88 |
+| larger, 8 vCPU | 154 / 154 | 3 | 0.066 | 110.51 / 110.77 | 9.20 / 9.21 |
+| larger, 16 vCPU | 124 / 125 | 3 | 0.126 | 90.57 / 84.99 | 7.62 / 7.33 |
+
+Using [published rates](https://docs.github.com/en/billing/reference/actions-runner-pricing), measured jobs total $0.504; the 60-second standard seed adds $0.006. The $0.510 gross compute estimate includes setup, restoration, measurement upload, and cleanup, but excludes storage and account discounts. Standard costs 80.95% less than 16-core and 33.33% less than 4-core in each sampled cache state. It adds 78 seconds against the corresponding 16-core job.
+
+Standard jobs expose two vCPUs and 7.75 GiB RAM. Workspace maximum process RSS is 2.86 / 2.76 GiB; preview maximum process RSS is 0.76 / 0.74 GiB. Both complete without an OOM or timeout. GNU time RSS is not simultaneous process-tree memory. These samples establish successful execution, not a permanent memory guarantee.
+
+The comparison fixes source, lockfile, commands, and runtime versions, not physical CPUs or image release: standard and 4-core use image 20260831.293.1; 8-core and 16-core use 20260823.283.1. CPUs vary among AMD EPYC 9V74/7763 and Intel Xeon 8370C/8573C. One sample per cache state measures the offered labels, not isolated CPU scaling or statistical repeatability.
+
+The experiment does not deploy or access Cloudflare credentials. Measurement upload takes zero to one second; warm-cache restore takes six to ten seconds. For context, [production job 101428009994](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34011495156/job/101428009994) spends 14 seconds uploading, one second verifying, and two seconds commenting on a different SHA. Adding that overhead to this experiment is a projection, not a measured standard-runner publication result. The actual PR preview workflow owns deployment confirmation.
+
+## Alternatives considered
+
+**Keep 16-core.** It provides the shortest measured job, but costs $0.102 more per sample for a 78-second improvement. Preview builds do not justify that premium for this cost-focused decision.
+
+**Select 4-core or 8-core.** Both succeed and shorten builds, but their rounded sample costs exceed standard Ubuntu. Four-core retains more RAM and disk headroom if future workloads exhaust standard capacity; such a change requires new measurements.
+
+**Move to self-hosted.** Rejected by scope: previews remain on GitHub CI. The existing Linux and Windows registrations can share persistent hosts; their dependency, store-volume, and cleanup assumptions do not apply to fresh hosted VMs. No failover or trust condition changes.
+
+## Consequences
+
+Previews trade approximately 78 seconds of sampled build-job latency for lower compute cost. Production Cloudflare latency, image rollout variance, future build growth, and broader success rates remain observable limitations. No hourly or monthly savings are extrapolated from this single experiment. The temporary benchmark workflow and its safety test are absent from the final tree; the experiment commits and linked run preserve the method and evidence.
+
+The executed [focused regression](../../../../scripts/preview-workflow.spec.ts) pins hosted routing, PR triggers and permissions, immutable full builds, restore-only caching, publication shaping, protected-image checks, and idempotent comments. A physical self-hosted routing mutation fails its routing assertion; restoration passes all three tests. No model-visible runtime behavior changes, so no Session snapshot changes are required.

+ 46 - 0
.agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.zh.md

@@ -0,0 +1,46 @@
+# Agent Note: 基于测量的 GitHub 托管 PR 预览规格
+
+Status: implemented
+
+[English](2026-09-06-preview-hosted-runner-sizing.md) | 中文
+
+## 问题
+
+PR(Pull Request)预览构建完整工作区及浏览器 worker VFS 镜像。较低的每分钟运行器价格不能保证较低的作业成本,因为 GitHub 将每个作业向上取整至整分钟。将预览移至持久化自托管机器还会改变隔离方式,不属于本决策范围。
+
+## 决策
+
+[预览工作流](../../../../.github/workflows/build-preview-cloudflare.yml) 使用标准 GitHub 托管 `ubuntu-24.04`。构建、缓存、部署、受保护镜像验证及评论语义保持不变。[规格参考](../../../../.github/preview-sizing/README.zh.md) 负责比较要求。独立的 CI [故障切换手册](2026-07-26-ci-failover-runbook.zh.md) 保留其运行器切换决策;预览不使用这些开关。
+
+### 测量
+
+[实验 34012729982](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34012729982) 的八种规格/缓存组合及一个缓存预热作业均成功。每个测量作业检出 SHA `9149d7e7ef945b5601711badd3cf63d58ab384f5`,使用 Node 24.19.0 与 pnpm 11.7.0,并执行不可变安装、完整工作区构建、预览/VFS 打包,以及含 gzip 完整性验证的本地上传内容整理。热作业恢复同一个运行私有精确 pnpm 缓存;冷作业跳过恢复,但包含 pnpm 引导安装文件。不恢复编译产物。
+
+| 运行器 | 冷 / 热作业秒数 | 各自取整分钟数 | 各自美元费用 | 冷 / 热工作区秒数 | 冷 / 热预览秒数 |
+|---|---:|---:|---:|---:|---:|
+| 标准,2 vCPU | 202 / 203 | 4 | 0.024 | 138.92 / 147.21 | 12.65 / 12.94 |
+| 大型,4 vCPU | 177 / 162 | 3 | 0.036 | 124.21 / 114.86 | 10.77 / 9.88 |
+| 大型,8 vCPU | 154 / 154 | 3 | 0.066 | 110.51 / 110.77 | 9.20 / 9.21 |
+| 大型,16 vCPU | 124 / 125 | 3 | 0.126 | 90.57 / 84.99 | 7.62 / 7.33 |
+
+按[公开费率](https://docs.github.com/en/billing/reference/actions-runner-pricing),测量作业合计 $0.504;60 秒标准预热作业增加 $0.006。$0.510 总计算费用估算包含设置、恢复、测量上传及清理,但不含存储和账户折扣。在每种采样缓存状态下,标准运行器比 16 核低 80.95%,比 4 核低 33.33%。相比对应的 16 核作业增加 78 秒。
+
+标准作业提供两个 vCPU 与 7.75 GiB 内存。工作区最大进程 RSS 为 2.86 / 2.76 GiB;预览最大进程 RSS 为 0.76 / 0.74 GiB。两者均未发生 OOM 或超时并完成。GNU time RSS 不是进程树同时占用的内存总量。这些样本证明成功执行,而非永久内存保证。
+
+比较固定源代码、锁文件、命令和运行时版本,但不固定物理 CPU 或镜像版本:标准与 4 核使用镜像 20260831.293.1;8 核与 16 核使用 20260823.283.1。CPU 包括 AMD EPYC 9V74/7763 与 Intel Xeon 8370C/8573C。每种缓存状态的单个样本测量所提供的标签,而非独立 CPU 扩展性或统计可重复性。
+
+实验不部署,也不访问 Cloudflare 凭据。测量上传耗时零至一秒;热缓存恢复耗时六至十秒。作为背景,[生产作业 101428009994](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34011495156/job/101428009994) 在不同 SHA 上上传耗时 14 秒、验证一秒、评论两秒。将该开销加至本实验属于推算,而非已测量的标准运行器发布结果。实际 PR 预览工作流负责部署确认。
+
+## 考虑过的替代方案
+
+**保留 16 核。** 它提供最短的测量作业,但为 78 秒改善使每个样本增加 $0.102。对于本次以成本为重点的决策,预览构建不值得这项溢价。
+
+**选择 4 核或 8 核。** 两者均成功并缩短构建,但取整后的样本费用高于标准 Ubuntu。若未来工作负载耗尽标准容量,4 核可保留更多内存与磁盘余量;这样的变更需要新测量。
+
+**移至自托管。** 因范围限制而拒绝:预览保留在 GitHub CI。现有 Linux 与 Windows 注册实例可能共享持久化主机;其依赖、store 卷及清理假设不适用于全新的托管 VM。不改变故障切换或信任条件。
+
+## 影响
+
+预览以约 78 秒采样构建作业延迟换取更低的计算费用。生产 Cloudflare 延迟、镜像发布差异、未来构建增长及更广泛的成功率仍是可观测限制。不从本次单一实验外推每小时或每月节省。最终文件树不包含临时基准工作流及其安全测试;实验提交与链接的运行保留方法和证据。
+
+已执行的[针对性回归](../../../../scripts/preview-workflow.spec.ts) 固定托管路由、PR 触发器与权限、不可变完整构建、只恢复缓存、发布内容整理、受保护镜像检查及幂等评论。实际修改为自托管路由会使路由断言失败;恢复后全部三个测试通过。不改变模型可见运行时行为,因此不需要修改 Session 快照。

+ 6 - 0
.agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.md
+2026-09-06-python-runtime-windows-hosted.md: ca2f02e8bac8a90be2b10bd6d7ae0b68215152ae
+2026-09-06-python-runtime-windows-hosted.zh.md: e1d2ca1a65de19a6604f0848de23fe5cc100e87f

+ 25 - 0
.agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.md

@@ -0,0 +1,25 @@
+# Agent Note: Windows Python runtime CI stays on GitHub-hosted Windows
+
+Status: implemented
+
+English | [中文](2026-09-06-python-runtime-windows-hosted.zh.md)
+
+## Problem
+
+The Windows x64 target in [build-exe-for-python-sdk.yml](../../../../.github/workflows/build-exe-for-python-sdk.yml) started resolving through `DSH_CI_FAILOVER_WINDOWS=selfhosted` for trusted pull-request CI when #3629 added the failover selector and the job-private Windows toolchain. The shared `dsh-win-ci` pool did not make the lane more reliable. On 2026-09-06 the installed-wheel smoke passed at 09:12 on `dsh-win-ci-16` for [an earlier commit of the same pull request](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384), then failed at 10:06 on `dsh-win-ci-21` for [another pull request](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34026500701) and at 10:46 on `dsh-win-ci-04` for [the same pull request](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34028339888/job/101473395734), where `smoke_sdk_profile_plugin`'s packaged `dsh plugin add` child exited without output while the Linux and macOS cells of that run passed. The migration proposal ([#3629](https://github.com/deepseek-harness/deepseek-harness/pull/3629)) remained `proposed` because its throughput and shared-load acceptance criteria were never measured.
+
+## Decision
+
+The Windows x64 target always uses its hosted `matrix.runner` — `windows-2025` for pull-request CI — with the standard setup-python toolchain, the pnpm cache restore, and the pkg cache. The failover selector, the job-private Python setup step, the self-hosted dependency install and post-step cleanup, the private setup script, and the routing spec from #3629 are removed. `DSH_CI_FAILOVER_WINDOWS=selfhosted` again retargets only the native Windows jobs in [ci.yml](../../../../.github/workflows/ci.yml); the [failover runbook](2026-07-26-ci-failover-runbook.md) and [python/development.md](../../../../python/development.md) describe hosted-only runtime builds. The migration's UTF-8 mode exports existed because the persistent host used a GBK default code page; hosted images provide the locale the lane previously ran under.
+
+## Alternatives considered
+
+**Keep the failover routing.** Rejected: the shared pool reproduced the same silent installed-wheel child death twice in one day while the migrated inventory's throughput acceptance stayed open, and routing a correctness lane through failover state couples it to an unrelated pool-outage switch.
+
+**Fix the shared pool instead.** Left to pool operators: the observed failures are subprocesses dying without output, not a missing image prerequisite, and the same image serves the native Windows failover jobs.
+
+**Retain the job-private toolchain on hosted images.** Rejected: the private uv/Python download exists to avoid mutating a persistent shared host; disposable hosted images already provide the registered Python 3.10 toolchain the pre-migration lane used.
+
+## Consequences
+
+Every qualifying pull request again pays GitHub-hosted Windows capacity for the runtime build, and the job-private setup and cleanup machinery — including the bounded filesystem retries — is gone with the lane. In exchange each build runs on a disposable host with the proven toolchain and hosted caches, and the Windows failover switch covers only the native Windows jobs as documented before the migration. A future self-hosted attempt must re-validate throughput and failure reproducibility on the actual pool before any routing change.

+ 25 - 0
.agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.zh.md

@@ -0,0 +1,25 @@
+# Agent Note: Windows Python runtime CI 保留在 GitHub 托管 Windows 上
+
+Status: implemented
+
+[English](2026-09-06-python-runtime-windows-hosted.md) | 中文
+
+## 问题
+
+当 #3629 加入故障切换选择器与作业私有的 Windows 工具链后,[build-exe-for-python-sdk.yml](../../../../.github/workflows/build-exe-for-python-sdk.yml) 中的 Windows x64 目标开始对受信任的 PR CI 通过 `DSH_CI_FAILOVER_WINDOWS=selfhosted` 解析运行器。共享的 `dsh-win-ci` 池并未让该通道更可靠。2026-09-06,安装后 wheel 冒烟测试在 09:12 于 `dsh-win-ci-16` 上为[同一拉取请求的较早提交](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384)通过,随后 10:06 在 `dsh-win-ci-21` 上为[另一个拉取请求](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34026500701)失败,10:46 在 `dsh-win-ci-04` 上为[同一拉取请求](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34028339888/job/101473395734)失败——`smoke_sdk_profile_plugin` 打包的 `dsh plugin add` 子进程无输出即退出,而该次运行的 Linux 与 macOS 单元均通过。迁移提案([#3629](https://github.com/deepseek-harness/deepseek-harness/pull/3629))保持 `proposed`,因为其吞吐量与共享负载验收标准从未实测。
+
+## 决策
+
+Windows x64 目标始终使用托管的 `matrix.runner`——PR CI 为 `windows-2025`——配以标准 setup-python 工具链、pnpm 缓存恢复与 pkg 缓存。来自 #3629 的故障切换选择器、作业私有 Python 准备步骤、自托管依赖安装与后置清理、私有准备脚本及路由测试均被移除。`DSH_CI_FAILOVER_WINDOWS=selfhosted` 再次只重定向 [ci.yml](../../../../.github/workflows/ci.yml) 中的原生 Windows 作业;[故障切换手册](2026-07-26-ci-failover-runbook.zh.md)与 [python/development.zh.md](../../../../python/development.zh.md) 描述仅托管的 runtime 构建。迁移中的 UTF-8 模式导出之所以存在,是因为持久主机使用 GBK 默认代码页;托管镜像提供该通道此前运行的区域设置。
+
+## 已考虑的替代方案
+
+**保留故障切换路由。** 不采用:共享池同一天两次复现相同的安装后 wheel 子进程无声死亡,而迁移清单的吞吐量验收仍然悬置;并且把正确性通道路由进故障切换状态,会使其耦合到无关的池故障开关。
+
+**改为修复共享池。** 交由池运维者处理:观测到的失败是无输出即退出的子进程,而非镜像前置条件缺失;同一镜像还服务原生 Windows 故障切换作业。
+
+**在托管镜像上保留作业私有工具链。** 不采用:私有 uv/Python 下载的存在理由是不修改持久共享主机;一次性托管镜像已提供迁移前通道使用的已注册 Python 3.10 工具链。
+
+## 后果
+
+每个符合条件的拉取请求再次为 runtime 构建支付 GitHub 托管 Windows 容量,作业私有准备与清理机制(包括有界文件系统重试)随通道一同移除。交换来的是每次构建运行在带标准工具链与托管缓存的一次性主机上,且 Windows 故障切换开关只覆盖迁移前文档所述的原生 Windows 作业。未来的自托管尝试必须在任何路由变更前,对实际池重新验证吞吐量与失败可复现性。

+ 6 - 0
.agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.md
+2026-09-06-release-rehearsal-selfhosted.md: 415ae4716e9bc0ae9b165afc807f6f41e8a57e04
+2026-09-06-release-rehearsal-selfhosted.zh.md: a6fa441d01e66cea998d77a9b1be588053ac60a5

+ 27 - 0
.agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.md

@@ -0,0 +1,27 @@
+# Agent Note: trusted release rehearsals on persistent Linux runners
+
+Status: implemented
+
+English | [中文](2026-09-06-release-rehearsal-selfhosted.zh.md)
+
+## Problem
+
+Dependency-layout and release-pack rehearsals consume hosted Linux minutes without requiring npm or API credentials. Moving arbitrary pull-request code or credentialed publication onto a persistent shared host would weaken isolation; reusing a checkout without cleaning would also weaken the packed-payload proof.
+
+## Decision
+
+The two jobs in [release.yml](../../../../.github/workflows/release.yml) and the pack job in [release-vendor.yml](../../../../.github/workflows/release-vendor.yml) select the existing self-hosted Linux pool only with the writer-controlled `DSH_CI_FAILOVER_LINUX` repository variable set to `selfhosted`. The selector requires the canonical repository and a non-Dependabot actor, then admits only master pushes or same-repository, non-fork PRs whose author is not Dependabot. Manual dispatch always selects `ubuntu-24.04`, as do all other rejected contexts. The [failover runbook](2026-07-26-ci-failover-runbook.md) owns the platform switches and standby operation. Release rehearsals intentionally share the Linux switch with main CI: enabling or disabling it routes both workloads, not releases independently. Unset remains the hosted default; hosted-minute savings occur only while an operator selects `selfhosted`, whether for an outage or a longer-running cost choice.
+
+The runner labels are `[self-hosted, linux, x64, vm-backup]`. Runner registrations share one VM, not independent machine capacity. Each job uses its runner-private temporary volume for Node compile cache and node-gyp headers before pnpm setup, and a pnpm setup destination qualified by run, attempt, and job. `TMPDIR` also points to `runner.temp`, so temporary npm consumers stay outside the checkout but inside runner cleanup even when a killed process cannot execute `finally`. The persistent pnpm store stays outside checkout cleanup; only GitHub-hosted runners restore the remote store cache. Neither rehearsal workflow saves remote caches.
+
+Checkout explicitly cleans ignored and untracked output before immutable installation and the existing builds. Full tag history, pack concurrency, dependency checks, tarball verification, and artifact retention remain unchanged. The packed-install verifier creates a fresh consumer outside the checkout, installs tarballs with npm, removes inherited Node resolution hooks, and deletes the consumer in `finally`; a warm pnpm store cannot substitute workspace links or stale build output for a tarball payload. The [npm release decision](2026-08-10-npm-release-sequences.md) still owns release families and publication. Both manual publish workflows remain entirely hosted and gain no credentials or registry changes here.
+
+## Alternatives considered
+
+Always-hosted rehearsals avoid persistent-host risk but retain all hosted minutes. Always-self-hosted rehearsals remove the portable fallback. A scheduling job or reusable workflow adds another logical job and hides the three short setup sequences. Allowing manual dispatch on arbitrary refs gives a maintainer action broader persistent-host access than the explicit event trust rule.
+
+## Consequences
+
+Unsetting the variable or changing it away from `selfhosted` routes subsequent eligible jobs to hosted Ubuntu. This is an operator-selected fallback, not automatic runner-health detection or failover for already queued jobs. The shared VM can still contend with other trusted jobs, and repository writers remain responsible for code admitted to its persistent trust domain. No workflow provisions host packages or changes global host configuration.
+
+[scripts/tests/ci-release-selfhosted.spec.ts](../../../../scripts/tests/ci-release-selfhosted.spec.ts) evaluates the committed selectors with trusted events and negative controls for forks, Dependabot, other repositories, non-master pushes, dispatches, missing PR data, and disabled switches. It pins setup ordering, checkout cleanup, hosted-only remote cache access, publication isolation, and the retained commands. Real release-build and packed-install execution remains the PR CI verification owner; selector tests do not claim to reproduce those builds.

+ 27 - 0
.agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.zh.md

@@ -0,0 +1,27 @@
+# Agent Note: 在持久化 Linux 运行器上执行受信任的发布演练
+
+Status: implemented
+
+[English](2026-09-06-release-rehearsal-selfhosted.md) | 中文
+
+## Problem
+
+依赖布局检查和发布打包演练消耗托管 Linux 分钟,但不需要 npm 或 API 凭据。将任意拉取请求代码或携带凭据的发布任务放到持久化共享主机会削弱隔离;复用未经清理的检出目录也会削弱打包载荷验证。
+
+## Decision
+
+[release.yml](../../../../.github/workflows/release.yml) 的两个作业和 [release-vendor.yml](../../../../.github/workflows/release-vendor.yml) 的打包作业仅在写权限维护者控制的仓库变量 `DSH_CI_FAILOVER_LINUX` 设为 `selfhosted` 时选择现有 Linux 自托管池。选择器要求当前仓库为正式仓库且触发者不是 Dependabot,然后只接纳 master 推送,或作者不是 Dependabot 的同仓库、非 fork PR(Pull Request)。手动触发始终选择 `ubuntu-24.04`,其他不满足条件的上下文也一样。[故障切换手册](2026-07-26-ci-failover-runbook.zh.md) 负责按平台划分的开关与热备操作。发布演练有意与主 CI 共用 Linux 开关:启用或禁用会同时路由两类负载,不能独立切换发布演练。未设置时仍默认使用托管池;只有运维人员选择 `selfhosted` 期间才节省托管分钟,无论该选择用于故障恢复还是持续的成本控制。
+
+运行器标签为 `[self-hosted, linux, x64, vm-backup]`。运行器注册共享一台虚拟机,不代表独立机器容量。每个作业在 pnpm 设置前将 Node 编译缓存与 node-gyp 头文件放在运行器私有临时卷上,pnpm 设置目标路径包含运行、重试次数和作业标识。`TMPDIR` 也指向 `runner.temp`,因此临时 npm 消费目录既在检出目录之外,也在运行器清理范围之内,即使进程被强杀而无法执行 `finally` 也一样。持久化 pnpm 存储位于检出清理范围之外;只有 GitHub 托管运行器恢复远端存储缓存。两个演练工作流都不保存远端缓存。
+
+检出操作显式清理被忽略和未跟踪的输出,再执行锁定依赖安装与现有构建。完整标签历史、打包并发、依赖检查、压缩包验证和产物保留期均保持不变。打包安装验证器在检出目录外创建全新的消费目录,用 npm 安装压缩包,移除继承的 Node 解析钩子,并在 `finally` 中删除消费目录;预热 pnpm 存储无法用工作区链接或过期构建输出代替压缩包载荷。[npm 发布决策](2026-08-10-npm-release-sequences.zh.md) 仍负责发布族与发布操作。两个手动发布工作流全部保留在托管运行器上,本改动不增加凭据,也不改变注册表。
+
+## Alternatives considered
+
+始终使用托管演练可以避免持久化主机风险,但会保留全部托管分钟。始终自托管则失去可移植回退。增加调度作业或可复用工作流会多出一个逻辑作业,并隐藏三个简短的设置序列。允许任意引用的手动触发,会让维护者操作获得比明确事件信任规则更广的持久化主机访问权限。
+
+## Consequences
+
+取消变量或将其改为非 `selfhosted` 值,会将后续符合条件的作业路由到托管 Ubuntu。这是运维人员选择的回退,不会自动探测运行器健康,也不会切换已排队的作业。共享虚拟机仍可能与其他受信任作业竞争资源;仓库写权限维护者仍对进入持久化信任域的代码负责。工作流不安装主机系统包,也不修改全局主机配置。
+
+[scripts/tests/ci-release-selfhosted.spec.ts](../../../../scripts/tests/ci-release-selfhosted.spec.ts) 使用受信任事件和 fork、Dependabot、其他仓库、非 master 推送、手动触发、缺失 PR 数据、禁用开关等负向对照求值已提交的选择器。测试固定设置顺序、检出清理、仅托管运行器访问远端缓存、发布隔离和保留命令。真实发布构建与打包安装执行仍由 PR CI 验证;选择器测试不声称重现这些构建。

+ 6 - 0
.agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.md
+2026-09-05-base-default-file-editor.md: 68a86769a99d697f9c7c767e9ba59b0cf61669b1
+2026-09-05-base-default-file-editor.zh.md: e8302b609a8647b1a0b90923361da06e859868ab

+ 31 - 0
.agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.md

@@ -0,0 +1,31 @@
+# Agent Note: Shared base default file editor selection
+
+Status: implemented
+
+English | [中文](2026-09-05-base-default-file-editor.zh.md)
+
+## Problem
+
+The shared base selects both `read`/`write`/`edit` and `str_replace_editor`, which offer overlapping file editing interfaces. [Issue #3599](https://github.com/deepseek-harness/deepseek-harness/issues/3599) requests one default interface for base-backed profiles while preserving the dedicated minimal compositions.
+
+## Decision
+
+The [base patch](../../../../packages/bundle/base/cordis.patch.yml) selects `read`, `write`, and `edit` for file editing. It does not insert `tool-str-replace-editor`; SDK and Web application patches therefore need no disabling override. The editor package remains available to compositions that insert it explicitly.
+
+[Web minimal](../../../../packages/preset/agent-presets/presets/minimal/agent.cordis.yml) inserts its own `str-replace-editor` row in the agent scope. The standalone [sdk-minimal bundle](../../../../packages/bundle/sdk-minimal/cordis.patch.yml) inserts its own row without inheriting base. Both minimal compositions retain their editor.
+
+This refines the shared tool defaults in [one dsh launcher](../architecture/2026-08-22-single-dsh-application-launcher.md). That note remains active for launch ownership, shared services, and patch precedence; no active note is fully superseded.
+
+## Alternatives considered
+
+**Disable the editor separately in each application.** This leaves overlapping defaults in base and requires each consumer to opt out. The base owns the shared choice directly.
+
+**Delete the tool package or remove it from minimal.** The dedicated minimal compositions use this interface for file operations. Keeping the package and their explicit rows preserves that behavior.
+
+## Consequences
+
+Base-backed SDK, headless, ACP, and custom profiles omit the editor schema by default. Web standard also omits it. A profile, home, or invocation patch can add the tool with `insert`; a patch that only sets `disabled: false` requires an existing row and cannot create one. This decision does not make all SDK and Web tools identical.
+
+## Verification
+
+The [SDK process tests](../../../../apps/cli/tests/profiles/sdk/keyless-smoke.e2e.ts) capture actual model requests for default file tools, explicit editor insertion, and the standalone minimal roster. The [headless process test](../../../../apps/cli/tests/profiles/headless/tests/keyless-smoke.e2e.ts) checks the shared default through its application. [Web minimal snapshots](../../../../apps/web/tests/minimal-preset.snapshot.ts) exercise the editor through the minimal preset. The [headless](../../../../snapshots/session/headless.snapshot.ts), [SDK](../../../../snapshots/sdk/sdk.snapshot.ts), and [ACP](../../../../snapshots/acp/acp.snapshot.ts) recorded sessions pin the assembled model-visible outputs, including the SDK fixture that explicitly inserts the editor.

+ 31 - 0
.agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.zh.md

@@ -0,0 +1,31 @@
+# Agent Note: 共享 base 默认文件编辑器选择
+
+Status: implemented
+
+[English](2026-09-05-base-default-file-editor.md) | 中文
+
+## Problem
+
+共享 base 同时选择 `read`/`write`/`edit` 和 `str_replace_editor`,这些工具提供重叠的文件编辑接口。[Issue #3599](https://github.com/deepseek-harness/deepseek-harness/issues/3599) 要求基于 base 的 profile 默认使用一套接口,同时保留专用的极简组合。
+
+## Decision
+
+[base patch](../../../../packages/bundle/base/cordis.patch.yml) 选择 `read`、`write` 和 `edit` 负责文件编辑。它不插入 `tool-str-replace-editor`;因此 SDK 与 Web 应用 patch 无需禁用覆盖。编辑器包仍可供显式插入它的组合使用。
+
+[Web minimal](../../../../packages/preset/agent-presets/presets/minimal/agent.cordis.yml) 在 agent 作用域插入自己的 `str-replace-editor` 配置项。独立的 [sdk-minimal bundle](../../../../packages/bundle/sdk-minimal/cordis.patch.yml) 不继承 base,自行插入配置项。两种极简组合都保留其编辑器。
+
+本决策细化了[统一 dsh 启动器](../architecture/2026-08-22-single-dsh-application-launcher.zh.md)中的共享工具默认值。该文档对启动所有权、共享服务和 patch 优先级仍然有效;没有被完全取代的活跃 Agent Note。
+
+## Alternatives considered
+
+**在每个应用中分别禁用编辑器。** 这会在 base 中保留重叠的默认接口,并要求各消费方主动退出。共享选择由 base 直接负责。
+
+**删除工具包或从 minimal 移除它。** 专用的极简组合通过此接口完成文件操作。保留包及其显式配置项可以保留这一行为。
+
+## Consequences
+
+基于 base 的 SDK、headless、ACP 与自定义 profile 默认不包含该编辑器 schema。Web standard 同样不包含它。profile、home 或逐次调用 patch 可以通过 `insert` 添加工具;只设置 `disabled: false` 的 patch 需要已有配置项,无法创建配置项。本决策不要求 SDK 与 Web 的全部工具一致。
+
+## Verification
+
+[SDK 进程测试](../../../../apps/cli/tests/profiles/sdk/keyless-smoke.e2e.ts) 捕获默认文件工具、显式插入编辑器与独立极简工具清单的实际模型请求。[headless 进程测试](../../../../apps/cli/tests/profiles/headless/tests/keyless-smoke.e2e.ts) 通过所属应用检查共享默认值。[Web minimal 快照](../../../../apps/web/tests/minimal-preset.snapshot.ts) 通过极简 preset 执行编辑器。[headless](../../../../snapshots/session/headless.snapshot.ts)、[SDK](../../../../snapshots/sdk/sdk.snapshot.ts) 与 [ACP](../../../../snapshots/acp/acp.snapshot.ts) 录制会话固定组装后模型可见的输出,包括显式插入编辑器的 SDK fixture。

+ 6 - 0
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
+2026-09-04-session-open-performance-gate.md: 2820c9d7d0e5b7d9382c7f8d6540154440175f26
+2026-09-04-session-open-performance-gate.zh.md: 965b9035074504870bcb2f1ca8166962c264d75a

+ 105 - 0
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md

@@ -0,0 +1,105 @@
+# Agent Note: Required CI performance gate for opening large Sessions
+
+Status: implemented
+
+English | [中文](2026-09-04-session-open-performance-gate.zh.md)
+
+## Problem
+
+The Session format v2 rollout changed two paths whose cost scales with model output: the JSONL backend migrates and publishes a released-v0 log, and the Client folds each settled reply's embedded compact stream. Neither path had an executable performance check, so first open grew from about 35 ms to about 5 s on a 127,400-event synthetic log (and from about 0.3 s to 26 s on a 575,000-chunk real log, with 2.7 GB peak RSS and heap exhaustion under a 512 MB limit), while Client fold grew linearly with streamed deltas instead of compact records; these regressions reached master unnoticed.
+
+Measuring only `SessionPersistence.open()` does not stably describe the result for which a user or Host waits. Work can move among `open()`, `SessionHandle.read()`, Session restoration, and projection, while the first history page and cold Agent resume add separate orchestration above those operations. A single `heapUsed` sample without prior GC also cannot distinguish data still retained by the Session from reclaimable migration temporaries.
+
+## Decision
+
+Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench`. The private `@deepseek-ai/dsh-benchmarks` workspace owns benchmark-only dependencies. The command first builds workspace libraries and dedicated workers under `benchmarks/.dsh-build/`, then invokes `vitest.bench.config.ts`. The [standard hosted runner decision](2026-09-06-standard-hosted-benchmark-runner.md) owns runner selection and the outer job timeout. The job runs the benchmark lane alone; Vitest runs one file at a time and only prepares input, starts measurement children, aggregates results, and enforces budgets. Every timed CPU path executes compiled JavaScript under plain Node with `NODE_OPTIONS` removed and no TypeScript loader; bare workspace imports therefore resolve from `benchmarks/node_modules` through package exports to built `lib/` entries.
+
+Required performance gates live under top-level `benchmarks/`, grouped by measured user path rather than package ownership. Host files use `*.bench.ts`, Client-face files use `*.bench.client.ts`, and scenario-specific workers and fixtures stay beside their benchmark without a benchmark suffix. Package-local `.perf.ts` files remain non-gating diagnostics; `scripts/` owns orchestration rather than benchmark cases.
+
+The Session benchmarks synthesize a released-v0 input from fixed parameters: 200 turns with 500 text deltas and 125 reasoning deltas per turn, for 127,400 logical events. The input uses Zstandard with fixed logical-row grouping and frame partitioning, so every run processes the same events, bytes, and frame distribution. The fixture constructs the immutable released-v0 physical rows directly instead of depending on a current-runtime historical encoder; compression and every measured read or migration entry point still use production code. Setup writes the input into a private temporary directory for each sample before timing starts; benchmarks never use recorded Sessions.
+
+Every Session endpoint runs at two user-lifecycle points. `first-open` starts with only the released V0 generation and therefore includes migration and successor publication. Setup produces `post-upgrade-reopen` once through that same production migration outside measurement, then copies both the unchanged V0 predecessor and published V2 successor into each sample root. Reopen samples use a fresh process, so they measure an upgraded user's later disk open without migration or process-local caches.
+
+Each access-kind and endpoint sample runs in a fresh compiled Node child process. Module imports, Host service initialization, and fixture preparation finish before measurement; the measured process performs no extra parse warm-up. Normal-heap mode runs five independent samples, reports every sample plus minimum, median, and maximum, and enforces access-specific fixed budgets against the median. Another child runs the same path under a fixed 128 MB old-space limit and checks only that it completes; extra GC caused by the constrained heap does not enter the normal timing baseline.
+
+The lane contains three independent Session-opening benchmarks and retains the Client-fold benchmark:
+
+| Benchmark | Measured path | Timing metrics |
+|---|---|---|
+| Phase profile | Executes the real persistence open, handle read, Session restore, and projection for both first open and post-upgrade reopen | `openMs`, `readMs`, `sessionRestoreMs`, and `projectionMs` each have a fixed budget; encoding, writes, verification, and publication awaited by migration all belong to first-open `openMs` |
+| First history | Reads each access kind through the Host Session history controller until it produces the first paginated snapshot | Separate first-open and reopen end-to-end budgets; each includes source stat, reading, restoration, projection, pagination, and snapshot construction, while first open additionally includes migration; both exclude Gateway network transport, Client fold, and browser paint |
+| Agent resume | Calls `ctx.agents.resume()` for each access kind until Agent creation, setup, publication, and loop startup finish | Separate first-open and reopen end-to-end budgets; neither path runs after first-history or reuses that benchmark's cache |
+| Client fold | Folds small and large v2 history windows through the real `ConversationNodeAssembler` and every Chat Definition | The large window's absolute time and scaling relative to the small window each have a fixed budget |
+
+The phase profile invokes each layer's production entry point explicitly and does not copy any decode, migration, restore, or projection algorithm. First-history and Agent-resume each run their real higher-level entry point against fresh first-open and reopen roots, so component measurements do not stand in for end-to-end results and one scenario cannot warm another's process or Session cache. The sum of the four phases is diagnostic only; an outer clock independently measures each end-to-end result.
+
+Normal-heap mode performs a fixed pair of explicit garbage collections after Host initialization and before the cold Session is touched, then records starting memory. It stops operation timing before performing the same garbage-collection sequence while the scenario's intended long-lived objects remain explicitly reachable, then records ending memory. The Agent-resume endpoint retains the Agent, Session, complete events, and normal service caches; its `heapUsed` delta is the primary resident-Session memory budget. Every scenario also reports `external`, `arrayBuffers`, post-GC RSS, and `process.resourceUsage().maxRSS`; the 128 MB mode prevents transient allocation peaks from being hidden by endpoint collection. Explicit garbage-collection time is excluded from operation timing.
+
+The performance gate does not duplicate semantic assertions owned by functional tests; it requires only that the target call completes and reaches its measured endpoint. The Client-fold benchmark continues to use the real `ConversationNodeAssembler` and every Chat Definition, and requires both the large window's absolute time and its scaling relative to the small window to remain below fixed budgets.
+
+Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Except for current-generation `open`, source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. Current-generation `open` uses a directly measured standard-runner expectation of 50 ms with only the 1.25× headroom, rounded up to a 63 ms budget. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
+
+## Calibration evidence
+
+The comparison is orthogonal by user lifecycle, not by artifact representation. Both implementations receive the same fixed V0 bytes for first open. For reopen, each implementation reads the format it considers current in a fresh process: the pre-stack reference remains on V0, while the V2 implementation reads its published V2 successor. This intentionally compares the same user's later-open experience rather than two codecs over one data structure.
+
+Five-sample medians on the same Node 24 reference machine establish the positive and negative controls:
+
+| Access kind | Implementation | Four-phase total | First history | Agent resume | Agent retained heap | 128 MB old space |
+|---|---|---:|---:|---:|---:|---|
+| First open | Pre-stack reference | 249.0 ms | 253.8 ms | 100.7 ms | 26.1 MB | Completes |
+| First open | Repeated-snapshot regression | 4,197.5 ms | 4,284.8 ms | 4,197.9 ms | 4.4 MB | Exhausts heap |
+| Post-upgrade reopen | Pre-stack reference | 251.1 ms | 253.8 ms | 100.7 ms | 26.1 MB | Completes |
+| Post-upgrade reopen | Repeated-snapshot regression | 49.2 ms | 50.4 ms | 43.8 ms | 4.5 MB | Completes |
+
+The pre-stack implementation keeps V0 as its current format, so first open does not change its on-disk representation; its native V0 first-history and Agent-resume measurements therefore apply to both lifecycle rows.
+
+The [standard two-CPU run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384/job/101461539961) at `ca3ffe95dac2c55eefeb16ed9b61067bbd19ee90` uses Node 24.20.0 x64 and Ubuntu image `20260831.293.1`. Its five current-generation `open` samples are 49.2, 47.4, 49.1, 48.6, and 48.1 ms: median 48.6 ms, maximum 49.2 ms. The rounded 50 ms CI expectation gives a 63 ms limit without reapplying the 2× machine scale. The log identifies two available CPUs but not their model; it does not isolate hardware from the Node-version change. This is endpoint-specific runner calibration, not evidence of an application optimization or a new reference-machine measurement. Every other benchmark passes its existing budget. Deterministic controls reject the observed median at the historical 30 ms limit, accept it at 63 ms, reject a synthetic 75 ms reopen median, and reject a synthetic 4,000 ms first-open duration at its unchanged 550 ms limit. These controls verify budget enforcement, not a measured new regression.
+
+The calibrated source budgets are:
+
+| Measurement | Reference expectation | CI budget |
+|---|---:|---:|
+| First-open `open` | 220 ms | 550 ms |
+| Current-generation `open` | 12 ms (historical reference; CI expectation: 50 ms) | 63 ms |
+| Complete read | 8 ms | 20 ms |
+| Session restore | 24 ms | 60 ms |
+| Projection | 14 ms | 35 ms |
+| First-open first history | 220 ms | 550 ms |
+| Current-generation first history | 48 ms | 120 ms |
+| First-open Agent resume | 180 ms | 450 ms |
+| Current-generation Agent resume | 40 ms | 100 ms |
+| Agent retained heap | 26.1 MB | 33 MB |
+| Client-fold absolute time | 16 ms | 40 ms |
+| Client-fold delta scaling | 2.5× | 3.125× |
+| Constrained old space | — | 128 MB |
+
+## Alternatives considered
+
+**Check out the historical commit and compare it on every CI run.** Rejected because a historical checkout requires a separate install, and old and current revisions can assign work to different API phases, adding runtime, dependency, and interface drift. A fixed workload with static budgets calibrated against positive and negative controls is easier to reproduce and review.
+
+**Measure only first open from V0.** Rejected because migration is a one-time upgrade cost and cannot protect later opens of the settled current generation from regressions. The two access kinds need separate measurements and budgets.
+
+**Measure only the four component phases.** Rejected because component measurements locate cost but omit source stat, orchestration, pagination, and snapshot construction, and cannot prove that the complete first-history path remains usable and fast enough.
+
+**Measure only first-history or Agent-resume total time.** Rejected because an end-to-end number protects the result but cannot identify whether storage, reading, Session restoration, or projection regressed; four phase budgets retain actionable attribution.
+
+**Add fine-grained timing instrumentation inside production implementations.** Rejected because those probes would expand production APIs and couple the benchmark to implementation details. Tests use existing service and object boundaries; costs that those boundaries cannot attribute remain part of the end-to-end result.
+
+**Run measured workers from TypeScript source.** Rejected because a source loader changes module resolution and startup behavior, and causes nested workers to select source-only bootstrap paths. Vitest remains an unmeasured orchestrator; every timed worker executes the build output exactly as plain Node consumers do.
+
+**Use only time budgets or only post-GC memory.** Rejected because time does not reveal memory regressions, while endpoint live memory cannot expose transient migration spikes. Normal-heap post-GC deltas and constrained-heap completion cover the two risks separately.
+
+**Benchmark the real recorded corpus.** Rejected because corpus fixtures stay small by policy, recorded material must not become benchmark input, and re-recording would silently move the workload.
+
+**Run benchmarks inside an existing gate aggregate.** Rejected because aggregate gates run concurrently on one runner, so wall-clock measurements inherit neighbouring CPU load.
+
+**Keep each cross-package gate under one participating product package.** Rejected because Session opening spans persistence, migration, projection, Host history, and Agent resume; choosing one participant creates misleading ownership and benchmark-only package dependencies. The repository-level tree owns the integrated user path, while package-local diagnostics remain with their implementation.
+
+**Put benchmark cases under `scripts/`.** Rejected because scripts own commands, generators, and orchestration, while a benchmark case owns typed test files, workers, fixtures, budgets, and lifecycle cleanup. A future reporting or calibration command may consume `benchmarks/` without moving the cases there.
+
+## Consequences
+
+Every pull request pays for one required Linux job; its Session portion runs several short-lived child processes in exchange for cold caches, isolated V8 heaps, explicit GC state, and attributable failures. The repository-level benchmark tree accepts deliberate cross-package test dependencies without changing product package manifests. The fixed Zstandard workload covers both event volume and frame topology; first-open measurements protect the one-time upgrade experience, reopen measurements prevent regressions in later opens, phase budgets locate cost, first-history budgets protect user-visible waiting, Agent-resume budgets and post-GC deltas protect complete cold activation and resident memory, and the 128 MB mode protects the transient allocation ceiling.
+
+The gate does not measure network transfer, browser rendering, or recorded Sessions, and it is not a continuous performance-trend system. A Node or runner change requires resampling the same workload and reviewing the budgets; a business-implementation change must not relax a budget without new positive and negative control data.

+ 105 - 0
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md

@@ -0,0 +1,105 @@
+# Agent Note: 打开大型 Session 的必需 CI 性能 gate
+
+Status: implemented
+
+[English](2026-09-04-session-open-performance-gate.md) | 中文
+
+## 问题
+
+Session format v2 的推出改变了两条成本随模型输出增长的路径:JSONL backend 迁移并发布 released-v0 log,Client fold 每个已结算回复中嵌入的紧凑 stream。两条路径都没有可执行的性能检查,因此首次打开在 127,400 事件的合成 log 上从约 35 ms 增长到约 5 s(在 575,000 chunk 的真实 log 上从约 0.3 s 增长到 26 s,峰值 RSS 2.7 GB,并在 512 MB 堆限制下耗尽堆),Client fold 也随流式 delta 数而不是紧凑记录数线性增长,这些退化未被察觉地进入了 master。
+
+只测 `SessionPersistence.open()` 不能稳定表达用户或 Host 等待的结果。工作可以在 `open()`、`SessionHandle.read()`、Session restore 与 projection 之间移动,而首次历史页和冷 Agent 恢复还包含这些操作之上的独立编排。操作结束时未经 GC 的一次 `heapUsed` 采样也不能区分仍被 Session 持有的数据与可回收的迁移临时对象。
+
+## 决定
+
+Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench`。私有 `@deepseek-ai/dsh-benchmarks` workspace 拥有 benchmark 专属依赖。该命令先构建 workspace library 和 `benchmarks/.dsh-build/` 下的专用 worker,再调用 `vitest.bench.config.ts`。[标准托管运行器决策](2026-09-06-standard-hosted-benchmark-runner.zh.md)拥有运行器选择及外层 job 超时。该 job 单独运行 benchmark lane;Vitest 逐文件运行,只负责准备输入、启动测量子进程、汇总结果和执行预算断言。每条被计时的 CPU 路径都以纯 Node 执行编译后的 JavaScript,并移除 `NODE_OPTIONS` 且不加载 TypeScript runtime;workspace 裸导入因此从 `benchmarks/node_modules` 通过 package exports 解析到构建后的 `lib/` 入口。
+
+必需性能 gate 位于顶层 `benchmarks/`,按被测用户路径而非 package 归属组织。Host 文件使用 `*.bench.ts`,Client 面文件使用 `*.bench.client.ts`,场景专属 worker 与 fixture 留在对应 benchmark 旁且不带 benchmark 后缀。包内 `.perf.ts` 文件仍是非门禁诊断;`scripts/` 负责编排而不承载 benchmark case。
+
+Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮 500 个 text delta 与 125 个 reasoning delta,共 127,400 个逻辑事件。输入使用 Zstandard,并固定 logical rows 的分组与 frame 拆分,使每次运行处理相同的事件、字节与 frame 分布。fixture 直接构造不可变的 released-v0 physical rows,不依赖当前 runtime 的历史 encoder;压缩以及所有被测读取和 migration 入口仍使用生产代码。输入在计时前写入每个样本独占的临时目录;benchmark 不使用录制的 Session。
+
+每个 Session endpoint 都针对用户生命周期中的两个时点运行。`first-open` 最初只有 released V0 generation,因此包含 migration 与后继 generation 发布。测试准备阶段在计时外通过同一套生产 migration 生成一次 `post-upgrade-reopen`,再把未改动的 V0 前代和已发布的 V2 后继一起复制到每个样本目录。Reopen 样本使用全新进程,因此测量用户升级完成后的磁盘再次打开,不包含 migration 或进程内 cache。
+
+每个 access kind 与 endpoint 的样本都在全新、已编译的 Node 子进程中运行。模块加载、Host 服务初始化和 fixture 准备在测量开始前完成;测量进程不执行额外的预热解析。正常堆模式运行五个独立样本,报告全部样本及最小值、中位数和最大值,并以中位数执行各访问状态独立的固定预算。另一个子进程使用固定 128 MB old-space 上限运行同一路径,只判断能否完成;低堆限制引起的额外 GC 不进入正常时间基线。
+
+该 lane 包含三个独立的 Session 打开 benchmark,并保留 Client fold benchmark:
+
+| Benchmark | 被测路径 | 时间指标 |
+|---|---|---|
+| 阶段剖面 | 分别为 first open 与 post-upgrade reopen 执行真实 persistence open、handle read、Session restore 与 projection | `openMs`、`readMs`、`sessionRestoreMs`、`projectionMs` 各自使用固定预算;migration 所等待的编码、写入、verify 与 publish 全部归入 first-open `openMs` |
+| 首屏历史 | 两种 access kind 分别经 Host Session history controller 读取到首个分页 snapshot | First open 与 reopen 各有一个端到端预算;均包含 source stat、读取、Session restore、projection、分页与 snapshot 构造,first open 还包含 migration;两者都不包含 Gateway 网络传输、Client fold 或浏览器 paint |
+| Agent resume | 对两种 access kind 分别调用 `ctx.agents.resume()`,直到 Agent 创建、setup、发布与 loop 启动完成 | First open 与 reopen 各有一个端到端预算;两条路径都不与首屏历史串行,也不依赖它留下的 cache |
+| Client fold | 大小两个 v2 history window 经真实 `ConversationNodeAssembler` 与全部 Chat Definition fold | 大窗口的绝对时间与相对小窗口的缩放比各自使用固定预算 |
+
+阶段剖面显式调用各层正式入口,不复制 decode、migration、restore 或 projection 算法。首屏历史和 Agent resume 分别以新的 first-open 与 reopen 根目录运行真实上层入口,因此组件数据不冒充端到端结果,一个场景也不会给另一个场景预热进程或 Session cache。四阶段之和仅用于解释成本;首屏与 Agent resume 的端到端时间各自由外层时钟直接测量。
+
+正常堆模式在 Host 初始化完成且 Session 尚未访问时执行固定的两轮显式 GC,记录起点内存;操作计时结束后,在该场景要求的长期对象仍明确可达时再次执行同样的 GC,再记录终点内存。Agent resume 场景在终点保留 Agent、Session、完整 events 与正常服务 cache,它的 `heapUsed` 增量是常驻 Session 内存预算的主指标。每个场景同时报告 `external`、`arrayBuffers`、GC 后 RSS 和 `process.resourceUsage().maxRSS`;128 MB 模式继续防止瞬时分配峰值被终点 GC 隐藏。显式 GC 时间不计入操作时间。
+
+性能 gate 不重复功能测试的内容断言,只要求目标调用完成并到达对应的可观察终点。Client fold benchmark 继续使用真实 `ConversationNodeAssembler` 与全部 Chat Definition,要求大窗口的绝对时间和相对小窗口的缩放比均低于固定预算。
+
+预算按各测量终点分别校准。两次 Node 24.19 x64 CI 运行的中位数最大相差 5.2%;其 CPU 密集型壁钟时间是 Node 24.18 arm64 参考运行的 1.95–2.06 倍。除当前 generation `open` 外,源码常量记录参考机器上的预期耗时;`ciTimeBudget()` 将其乘以实测的 2 倍 CI 时间系数和 1.25 倍波动余量。当前 generation `open` 使用标准运行器直接测得的 50 ms 预期值,仅乘以 1.25 倍余量,向上取整得到 63 ms 预算。GC 后增量堆与 Client fold 缩放预算不属于壁钟时间,因此只使用 1.25 倍余量。128 MB 完成性检查仍是独立的瞬时分配限制。由此得到的 first-open 时间上限、受限堆检查与 Client fold 上限都会拒绝已知退化。栈前参考提交固定为 `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5`,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
+
+## 校准证据
+
+比较按用户生命周期正交,而不是按产物表示正交。两种实现的 first open 都接收完全相同的固定 V0 字节。Reopen 时,每种实现都在全新进程中读取自己认定的当前格式:栈前参考版本仍读取 V0,V2 实现则读取它已发布的 V2 后继。这里有意比较同一用户后续打开的体验,而不是让两个 codec 处理同一种数据结构。
+
+同一台 Node 24 参考机器上的五次样本中位数构成正反例:
+
+| Access kind | 实现 | 四阶段总时间 | 首屏历史 | Agent resume | Agent GC 后增量堆 | 128 MB old space |
+|---|---|---:|---:|---:|---:|---|
+| First open | 栈前参考版本 | 249.0 ms | 253.8 ms | 100.7 ms | 26.1 MB | 完成 |
+| First open | 重复 snapshot 退化实现 | 4,197.5 ms | 4,284.8 ms | 4,197.9 ms | 4.4 MB | 堆耗尽 |
+| Post-upgrade reopen | 栈前参考版本 | 251.1 ms | 253.8 ms | 100.7 ms | 26.1 MB | 完成 |
+| Post-upgrade reopen | 重复 snapshot 退化实现 | 49.2 ms | 50.4 ms | 43.8 ms | 完成 |
+
+栈前实现以 V0 作为当前格式,因此 first open 不改变磁盘表示;它的原生 V0 首屏历史与 Agent resume 测量同时适用于两个生命周期行。
+
+`ca3ffe95dac2c55eefeb16ed9b61067bbd19ee90` 上的[标准双 CPU 运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384/job/101461539961)使用 Node 24.20.0 x64 和 Ubuntu 镜像 `20260831.293.1`。当前 generation `open` 的五次样本为 49.2、47.4、49.1、48.6 和 48.1 ms:中位数 48.6 ms,最大值 49.2 ms。取整后的 50 ms CI 预期值给出 63 ms 上限,不重复乘以 2 倍机器系数。日志标明两个可用 CPU,但未记录型号;它无法区分硬件变化与 Node 版本变化的影响。这是端点专属的运行器校准,不是应用优化或参考机器新测量的证据。其他每项 benchmark 均通过既有预算。确定性正反例在历史 30 ms 上限下拒绝实测中位数,在 63 ms 下接受它,拒绝合成的 75 ms reopen 中位数,并以未改变的 550 ms 上限拒绝合成的 4,000 ms 首次打开耗时。这些正反例验证预算执行,不代表测得新的退化。
+
+校准后的源码预算如下:
+
+| 测量项 | 参考机预期 | CI 预算 |
+|---|---:|---:|
+| First-open `open` | 220 ms | 550 ms |
+| 当前 generation `open` | 12 ms(历史参考值;CI 预期值:50 ms) | 63 ms |
+| 完整 read | 8 ms | 20 ms |
+| Session restore | 24 ms | 60 ms |
+| Projection | 14 ms | 35 ms |
+| First-open 首屏历史 | 220 ms | 550 ms |
+| 当前 generation 首屏历史 | 48 ms | 120 ms |
+| First-open Agent resume | 180 ms | 450 ms |
+| 当前 generation Agent resume | 40 ms | 100 ms |
+| Agent GC 后增量堆 | 26.1 MB | 33 MB |
+| Client fold 绝对时间 | 16 ms | 40 ms |
+| Client fold delta 缩放比 | 2.5× | 3.125× |
+| 受限 old space | — | 128 MB |
+
+## 考虑过的替代方案
+
+**每次 CI checkout 历史提交并做相对比较。** 拒绝:历史 checkout 需要独立安装,旧版与当前版还可能把工作放在不同 API 阶段,增加时间、依赖和接口漂移。固定 workload 与经正反例校准的静态预算更容易复现和评审。
+
+**只测从 V0 first open。** 拒绝:migration 是一次性升级成本,不能防止进入稳定当前 generation 后的后续打开发生退化。两种 access kind 需要独立的测量与预算。
+
+**只测四个组件阶段。** 拒绝:组件测量便于定位,但会遗漏 source stat、编排、分页和 snapshot 构造,也不能证明首屏路径整体仍然可用且足够快。
+
+**只测首屏或 Agent resume 总时间。** 拒绝:端到端数字能保护结果,却不能指出退化来自存储、读取、Session restore 还是 projection;四阶段预算保留可操作的归因。
+
+**在生产实现内部添加细粒度计时桩。** 拒绝:这些桩会扩大生产接口并让 benchmark 与实现细节耦合。测试只使用既有服务和对象边界;无法由这些边界解释的成本保留在端到端结果中。
+
+**从 TypeScript 源码运行被测 worker。** 拒绝:源码 loader 会改变模块解析与启动行为,并使嵌套 worker 选择仅适用于源码的启动路径。Vitest 仍可作为不计时的编排层;每个被计时的 worker 都像纯 Node 消费方一样执行构建产物。
+
+**只用时间预算或只看 GC 后内存。** 拒绝:时间无法发现内存退化,终点存活内存也看不到迁移期间的瞬时爆发。正常堆的 GC 后增量与受限堆的完成性分别覆盖两类风险。
+
+**用真实录制语料做 benchmark。** 拒绝:语料 fixture 按策略保持小体量,录制材料不得成为 benchmark 输入,且其重新录制会静默移动 workload。
+
+**把 benchmark 放进现有 gate 聚合中运行。** 拒绝:聚合在一个 runner 上并发运行各 gate,壁钟测量会继承邻居的 CPU 负载。
+
+**把每个跨包 gate 放在一个参与的产品 package 下。** 拒绝:Session 打开跨越 persistence、migration、projection、Host history 与 Agent resume;任选一个参与方都会形成误导性的归属和仅为 benchmark 增加的 package 依赖。仓库级目录拥有集成用户路径,包内诊断仍留在对应实现旁。
+
+**把 benchmark case 放在 `scripts/` 下。** 拒绝:scripts 拥有命令、生成器与编排,而 benchmark case 拥有带类型的测试文件、worker、fixture、预算和生命周期清理。未来的报告或校准命令可以消费 `benchmarks/`,不需要把 case 移入其中。
+
+## 后果
+
+每个 pull request 多付出一个必需 Linux job;该 job 的 Session 部分运行多个短生命周期子进程,以换取冷 cache、独立 V8 heap、明确 GC 状态和可归因的失败。仓库级 benchmark 目录接受有意的跨包测试依赖,而不修改产品 package manifest。固定 Zstandard workload 同时覆盖事件规模与 frame 拓扑;first-open 测量保护一次性升级体验,reopen 测量防止后续打开退化,四阶段预算定位成本归属,首屏预算保护用户可见等待,Agent resume 预算与 GC 后增量保护完整冷恢复及常驻内存,128 MB 模式保护瞬时分配上限。
+
+该 gate 不测量网络传输、浏览器渲染或真实录制 Session,也不是持续性能趋势系统。Node 或 runner 变化需要用同一 workload 重新采样并评审预算;修改业务实现时不得顺带放宽预算而不提供新的正反例数据。

+ 6 - 0
.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.md
+2026-09-06-standard-hosted-benchmark-runner.md: af95af6ee8cef1128a5475863ea7a1aaf66f30b9
+2026-09-06-standard-hosted-benchmark-runner.zh.md: 47095a5b2e91318c2f3ec5ac3cc9f6df3bce04d6

+ 26 - 0
.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.md

@@ -0,0 +1,26 @@
+# Agent Note: Standard hosted runner for required benchmarks
+
+Status: implemented
+
+English | [中文](2026-09-06-standard-hosted-benchmark-runner.zh.md)
+
+## Problem
+
+Wall-clock performance checks need an isolated execution lane and a consistent runner class. Routing them through the enterprise Linux failover switch makes their measurements depend on either larger hosted capacity or a shared self-hosted VM, while also consuming capacity needed by parallel correctness checks.
+
+## Decision
+
+The required benchmark job in [ci.yml](../../../../.github/workflows/ci.yml) uses the standard GitHub-hosted `ubuntu-24.04` runner independently of Linux failover. It always attempts to restore the pnpm store cache and retains a standalone benchmark lane. The complete job has a 15-minute timeout covering setup, installation, builds, and measurements. This bounds infrastructure execution, not an individual performance assertion.
+
+The [Session performance decision](2026-09-04-session-open-performance-gate.md) continues to own workloads, timing and memory budgets, worker isolation, and calibration. Only current-generation `open` uses an endpoint-specific 50 ms standard-runner expectation with the existing 1.25× headroom, giving a 63 ms limit. All other performance budgets and the worker, test, and hook deadlines remain unchanged. Successful raw measurements remain in the Actions log through step-local `DSH_GATE_VERBOSE=1`. The hardware-comparison workflows retain their deliberately different runner sizes.
+
+## Alternatives considered
+
+- Enterprise or shared self-hosted routing retains more build capacity but ties the measurement environment to unrelated failover operations.
+- Increasing performance thresholds without endpoint measurements conflates a bounded CI execution with a regression allowance. Threshold changes require measured calibration and positive and negative controls.
+
+## Consequences
+
+A standard runner trades parallel build capacity for a fixed measurement class without removing the required verdict. Cache misses and runner variation can still affect total duration. Each runner change needs an actual hosted benchmark run before its job timeout is treated as validated; local workflow assertions alone cannot establish execution time.
+
+The owning [workflow tests](../../../../scripts/ci-workflow.spec.ts) pin runner routing, unconditional cache restoration, required status, and the job timeout. Negative controls reject failover routing, a cache condition, and the former 30-minute job bound.

+ 26 - 0
.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.zh.md

@@ -0,0 +1,26 @@
+# Agent Note: 必需 benchmark 使用标准托管运行器
+
+Status: implemented
+
+[English](2026-09-06-standard-hosted-benchmark-runner.md) | 中文
+
+## 问题
+
+壁钟性能检查需要独立执行的 lane 和一致的运行器类别。通过企业 Linux 故障转移开关路由这些检查,会让测量取决于大型托管运行器或共享自托管虚拟机,同时占用并行正确性检查所需的容量。
+
+## 决定
+
+[ci.yml](../../../../.github/workflows/ci.yml) 中的必需 benchmark job 使用标准 GitHub 托管 `ubuntu-24.04` 运行器,不受 Linux 故障转移影响。它始终尝试恢复 pnpm 存储缓存,并保留独立的 benchmark lane。整个 job 的超时为 15 分钟,覆盖准备、安装、构建和测量。这限制的是基础设施执行时间,而非单项性能断言。
+
+[Session 性能决策](2026-09-04-session-open-performance-gate.zh.md) 继续拥有工作负载、时间和内存预算、worker 隔离及校准。仅当前 generation `open` 使用端点专属的 50 ms 标准运行器预期值,乘以既有 1.25 倍余量后得到 63 ms 上限。其他性能预算以及 worker、测试和钩子的截止时间均保持不变。步骤级 `DSH_GATE_VERBOSE=1` 使成功运行的原始测量保留在 Actions 日志中。硬件比较工作流保留有意设置的不同运行器规格。
+
+## 考虑过的替代方案
+
+- 企业或共享自托管路由保留更多构建容量,但使测量环境受无关故障转移操作影响。
+- 没有端点测量就提高性能阈值,会混淆有界 CI 执行与退化容许量。阈值调整需要实测校准及正反例。
+
+## 后果
+
+标准运行器以并行构建容量换取固定测量类别,不移除必需判定。缓存未命中和运行器波动仍会影响总耗时。每次更换运行器都需要实际托管 benchmark 运行,才能认定 job 超时经过验证;本地工作流断言无法单独证明执行耗时。
+
+所属[工作流测试](../../../../scripts/ci-workflow.spec.ts) 固定运行器路由、无条件缓存恢复、必需状态及 job 超时。反例验证拒绝故障转移路由、缓存条件和原来的 30 分钟 job 上限。

+ 6 - 0
.agents/notes/proposed/architecture/2026-09-02-system-prompt-as-surface-node.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/proposed/architecture/2026-09-02-system-prompt-as-surface-node.md
+2026-09-02-system-prompt-as-surface-node.md: 56577a7199235e95f4a7c6500140c8a8841d48dc
+2026-09-02-system-prompt-as-surface-node.zh.md: da864fd300c93cae0210e758562b823c0a61679a

+ 78 - 0
.agents/notes/proposed/architecture/2026-09-02-system-prompt-as-surface-node.md

@@ -0,0 +1,78 @@
+# Agent Note: The system prompt is surface node 0
+
+Status: proposed
+
+English | [中文](2026-09-02-system-prompt-as-surface-node.zh.md)
+
+## Problem
+
+The system prompt has a different durable representation from every other message the model reads. Conversation messages are surface events (`user/message`, `assistant/message`, `tool/result`) folded in seq order by `Session.deriveMessages()`; the system prompt is the `system` field of the log-only `request/header` snapshot, and each DeepSeek serializer prepends it as wire message 0 (`serializeRequest`, `serializeRequestWithImages`). The [reconstructable-requests Agent Note](../../implemented/architecture/2026-07-05-reconstructable-requests.md) made both halves durable, but it left one model-visible fact with two homes: the surface owns the messages, the header owns the message in front of them.
+
+That split forces every reader of "what did the model see" to join two sources. The compaction summarizer (`buildSummarizationInput`) copies `header.system` in front of the region's derived messages; `dsh-token-meter` estimates the system prompt from the header while pricing every other message from the surface; the Web request-prompt card, the trajectory view, and the snapshot normalizer's `{{system}}` placeholder each read the header on their own. The loop's change detection is also split: `headerEquals` compares `system` byte-for-byte beside `config` and `tools`, so a prompt change and a tool change are indistinguishable in the log (`request/header` reason `change`) even though they are different operations on the conversation.
+
+The split also blocks the next step. A model that accepts a mid-conversation `system` message as a prompt replacement needs the harness to append a system-role message to history; with the prompt living in the header there is no surface representation to append, and the header would have to be frozen by special case. The [in-history replacement proposal](../feature/2026-09-02-in-history-system-prompt-replacement.md) depends on this note.
+
+## Proposal
+
+Move the system prompt onto the surface. It becomes an ordinary surface event, `system/message`, and every prompt lifecycle operation is one of the two existing `SurfaceOp` variants applied to that event type. The wire request does not change: the surface fold yields the same message list the serializers already build today, with the system message first.
+
+### The event
+
+`system/message` joins `SurfaceEventType` beside `user/message`, `assistant/message`, and `tool/result`. Its payload mirrors `tool/result`: `{ turn, step, message }`, where `message` is a `Message` with `role: 'system'`, exactly one text block holding the rendered prompt, and source `{ kind: 'plugin', plugin: '@deepseek-ai/dsh-system-prompt' }`. `deriveEventMessage` projects it verbatim, so `deriveMessages()` returns the system message at its surface position and both DeepSeek serializers, which already pass a `role: 'system'` history message through unchanged, emit it as wire message 0. `EpochHeader.system` is removed; the header keeps `config`, `adapterDefaults`, and `tools`.
+
+### The operations
+
+| Situation | Surface operation |
+|---|---|
+| First request of a session with a non-empty rendered prompt | append `system/message` as surface node 0, before the first `user/message` of the step |
+| Rendered prompt differs from the prompt at node 0 | replace node 0: `surfaceOp: { op: 'replace', start: <seq of node 0>, end: <same> }`, `sourceEventSeqs: [<seq of node 0>]` |
+| Rendered prompt is empty on the first request | no system node; a later non-empty prompt appends node 0 when the surface has no system node yet |
+
+Replacing node 0 is today's head rewrite expressed on the surface: the provider prefix changes from the first token, the log records the shadowed node through `sourceEventSeqs`, and `replaceGeneration` advances exactly as it does for a compaction replacement, so the loop's existing `startsSeries` detection (`requestSurfaceGeneration !== surfaceGeneration`) covers the prompt change without a `system` comparison in `headerEquals`. `request/header` keeps reasons `initial`, `resume`, `change`, and `series`; `change` now means config or tools changed.
+
+### Ownership in the loop
+
+`dsh-agent-loop` owns a `SystemPromptProjection` beside `RuntimeContextProjection` in `runtime-context.ts`. It restores the current system node from the log (the latest surviving `system/message` on the surface), follows `session/event` for new system nodes and for replacements whose `sourceEventSeqs` shadow the retained one, and returns the uncommitted append or replace intent when the rendered prompt differs. `turn()` commits that intent immediately before the step's `user/message` events, so the log order is the wire order. `step()` no longer passes `system` to `buildRequest`; the request is `header.config` plus `deriveMessages()` plus `header.tools`. The `dsh-agent-loop/invariant` companion keeps comparing the rebuilt request against the frozen one, now with the system message inside `messages`. `docs/architecture.md` records the new loop step order: claim, assemble, project system prompt, project runtime context, pre-step, commit system node, commit user messages, build request.
+
+### Consumers retargeted
+
+| Consumer | Today | After |
+|---|---|---|
+| DeepSeek serializers (`serializeRequest`, `serializeRequestWithImages`) | prepend `options.system` | serialize `options.messages` only; `GenerateOptions.system` remains for direct one-shot callers such as the summarizer and title providers |
+| `compaction-basic` `buildSummarizationInput` | `header.system` + region messages | node 0's derived message + region messages, still a genuine prefix of the routed request |
+| `compaction-basic` `selectCompactableRange` | head-anchored at `surfaceNodes[0]` | anchored at the first non-system node; node 0 is never inside a compaction range |
+| `dsh-token-meter` system estimate | `header.system` length | the system node is priced like every other surface node; the context breakdown labels it by its source plugin |
+| Web request-prompt card, trajectory request-header node, request inspection | read `header.system` | read the `system/message` node; the card keeps its collapsed inspectable presentation and is never a chat bubble |
+| Snapshot normalizer `{{system}}` placeholder, plan-mode tests asserting `header.system` | header | the system node's text |
+| TypeScript and Python SDK expected outputs | no system event | include the `system/message` event |
+| Human transcript projections (`isAppendSurfaceEvent` readers) | no system events | skip `system/message`; it is model history, not conversation |
+
+`RuntimeContextProjection` and `SystemPromptProjection` are symmetric: both watch owned surface nodes and their shadowing through `sourceEventSeqs`, and both hand the loop an uncommitted message that `turn()` commits. The difference is the role and the operation set — runtime context appends user-role snapshots only, the system prompt appends once and then replaces.
+
+## Alternatives considered
+
+**Keep `header.system` and add `system/message` only for updates.** Two homes for one fact: every consumer above would read the header for message 0 and the surface for later messages, and the loop would need a special case that ignores `system` in `headerEquals` while a surface system node exists. Rejected because the point of the change is one representation.
+
+**A dedicated log-only `system-prompt/change` event that rewrites the header.** Preserves the header as the home of the prompt and records changes as their own event kind, but still cannot express a system message inside history, so the in-history proposal would need a second mechanism anyway. Rejected.
+
+**Synthesize the system message inside the adapter from consecutive headers.** The adapter is stateless per request and never sees the log; a wire history that depends on adapter state is not reconstructable from the surface fold. Rejected.
+
+**Express the prompt as a `user/message` snapshot like runtime context.** Reuses an existing event type but sends the wrong role, so a model that treats a system message as authoritative would not. Rejected.
+
+## Acceptance criteria
+
+- `SurfaceEventType` contains `system/message`; `deriveEventMessage` projects it; `Session.append('system/message', …)` requires a `SurfaceIntent` like the other surface events.
+- `EpochHeader` has no `system` field; `headerEquals` compares `config`, `adapterDefaults`, and `tools` only.
+- A first request with a non-empty rendered prompt appends `system/message` as surface node 0 before the step's first `user/message`; a changed prompt replaces node 0 with `sourceEventSeqs` naming the shadowed node; an unchanged prompt appends nothing.
+- The DeepSeek wire request for every loop step is byte-identical to today's for the same session history: system first, then the folded conversation.
+- Compaction never selects node 0; the summarizer's replayed prefix starts with node 0's derived message.
+- `dsh-token-meter`, the Web request-prompt card, trajectory and inspection views, the snapshot normalizer, plan-mode tests, and both SDK expected outputs read the system node; the `dsh-agent-loop/invariant` companion rebuilds requests with the system message inside `messages`.
+- Keyless recorded snapshots that exercise a mid-session prompt change (plan mode entering and leaving) show a replaced node 0 instead of a `request/header` `change`.
+- `docs/architecture.md`, the `dsh-agent-loop`, `dsh-session`, `dsh-system-prompt`, `dsh-compaction-basic`, and `dsh-token-meter` READMEs, and the reconstructable-requests Agent Note describe the surface node as the home of the system prompt.
+
+## Risks
+
+- Every reader of `header.system` moves in one change; a missed reader fails at compile time because the field is gone, which is the intended failure mode.
+- Compaction region selection gains an invariant (node 0 is never compacted). A compaction provider other than `compaction-basic` that anchors at `surfaceNodes[0]` would shadow the prompt; the `dsh-session` surface manager rejects a replacement whose range covers surface node 0 while node 0 is a `system/message` unless the replacing event is itself a `system/message` covering exactly that node, so the invariant is enforced where the operation happens, not only in the shipped provider. System nodes at later positions carry no such protection: a compaction range may shadow them.
+- Replacing node 0 advances `replaceGeneration`, which today means "compaction happened" to some readers; those readers switch to inspecting the replacement event's type.
+- Recorded snapshot fixtures whose logs contain `header.system` are re-recorded; the fixtures, not the normalizer, change.

+ 78 - 0
.agents/notes/proposed/architecture/2026-09-02-system-prompt-as-surface-node.zh.md

@@ -0,0 +1,78 @@
+# Agent Note: 系统提示词是 surface 的第 0 号节点
+
+Status: proposed
+
+[English](2026-09-02-system-prompt-as-surface-node.md) | 中文
+
+## Problem
+
+系统提示词的持久化表示与模型读到的其他所有消息都不同。对话消息是 surface 事件(`user/message`、`assistant/message`、`tool/result`),由 `Session.deriveMessages()` 按 seq 顺序折叠;系统提示词则是仅记日志的 `request/header` 快照中的 `system` 字段,每个 DeepSeek 序列化器把它前置为协议消息 0(`serializeRequest`、`serializeRequestWithImages`)。[可重建请求 Agent Note](../../implemented/architecture/2026-07-05-reconstructable-requests.zh.md) 让两半都成为持久数据,却让一个模型可见的事实拥有两个归属:surface 拥有消息,header 拥有排在这些消息之前的那条消息。
+
+这种拆分迫使每个想知道「模型看到了什么」的读取方都要合并两个来源。压缩摘要器(`buildSummarizationInput`)把 `header.system` 复制到区域派生消息之前;`dsh-token-meter` 从 header 估算系统提示词,却从 surface 为其他每条消息计价;Web 请求提示词卡片、轨迹视图和快照归一化器的 `{{system}}` 占位符各自单独读取 header。循环的变更检测同样被拆开:`headerEquals` 在 `config` 和 `tools` 旁边逐字节比较 `system`,因此提示词变更与工具变更在日志中无法区分(`request/header` 的 reason 都是 `change`),尽管它们是对对话的两种不同操作。
+
+这种拆分还阻塞了下一步。一个把对话中途的 `system` 消息当作提示词替换来接受的模型,需要 harness 向历史追加一条 system 角色消息;当提示词住在 header 里时,没有可追加的 surface 表示,header 也只能靠特例被冻结。[历史内替换提案](../feature/2026-09-02-in-history-system-prompt-replacement.zh.md) 依赖本 Agent Note。
+
+## Proposal
+
+把系统提示词搬到 surface 上。它成为一个普通的 surface 事件 `system/message`,提示词生命周期中的每个操作都是对该事件类型施加现有两种 `SurfaceOp` 变体之一。协议请求不变:surface 折叠产出的消息列表与序列化器今天构建的完全相同,系统消息在最前面。
+
+### 事件
+
+`system/message` 加入 `SurfaceEventType`,与 `user/message`、`assistant/message`、`tool/result` 并列。它的载荷与 `tool/result` 对称:`{ turn, step, message }`,其中 `message` 是 `role: 'system'` 的 `Message`,恰好一个文本块承载渲染后的提示词,source 为 `{ kind: 'plugin', plugin: '@deepseek-ai/dsh-system-prompt' }`。`deriveEventMessage` 逐字投影它,因此 `deriveMessages()` 在其 surface 位置返回系统消息,而两个 DeepSeek 序列化器本已原样透传 `role: 'system'` 的历史消息,会把它作为协议消息 0 发出。`EpochHeader.system` 被移除;header 保留 `config`、`adapterDefaults` 和 `tools`。
+
+### 操作
+
+| 情形 | surface 操作 |
+|---|---|
+| 会话首个请求且渲染后的提示词非空 | 追加 `system/message` 作为 surface 第 0 号节点,位于该步骤首条 `user/message` 之前 |
+| 渲染后的提示词与第 0 号节点不同 | 替换第 0 号节点:`surfaceOp: { op: 'replace', start: <第 0 号节点的 seq>, end: <同一值> }`,`sourceEventSeqs: [<第 0 号节点的 seq>]` |
+| 首个请求时渲染后的提示词为空 | 没有系统节点;之后出现非空提示词且 surface 尚无系统节点时,追加为第 0 号节点 |
+
+替换第 0 号节点就是今天的头部重写在 surface 上的表达:提供方前缀从第一个 token 起改变,日志通过 `sourceEventSeqs` 记录被遮蔽的节点,`replaceGeneration` 与压缩替换时一样推进,因此循环现有的 `startsSeries` 检测(`requestSurfaceGeneration !== surfaceGeneration`)无需在 `headerEquals` 中比较 `system` 即可覆盖提示词变更。`request/header` 保留 `initial`、`resume`、`change`、`series` 四种 reason;`change` 现在表示 config 或 tools 变更。
+
+### 循环中的归属
+
+`dsh-agent-loop` 在 `runtime-context.ts` 中与 `RuntimeContextProjection` 并列拥有一个 `SystemPromptProjection`。它从日志恢复当前系统节点(surface 上最新存活的 `system/message`),跟随 `session/event` 观察新的系统节点以及 `sourceEventSeqs` 遮蔽了所保留节点的替换,并在渲染后的提示词不同时返回未提交的追加或替换意图。`turn()` 紧接在该步骤的 `user/message` 事件之前提交该意图,因此日志顺序即协议顺序。`step()` 不再向 `buildRequest` 传递 `system`;请求由 `header.config`、`deriveMessages()` 和 `header.tools` 构成。`dsh-agent-loop/invariant` 伴随组件继续把重建的请求与冻结的请求比较,只是系统消息现在位于 `messages` 内。`docs/architecture.md` 记录新的循环步骤顺序:领取、装配、投影系统提示词、投影运行时上下文、pre-step、提交系统节点、提交用户消息、构建请求。
+
+### 消费方迁移
+
+| 消费方 | 现状 | 变更后 |
+|---|---|---|
+| DeepSeek 序列化器(`serializeRequest`、`serializeRequestWithImages`) | 前置 `options.system` | 只序列化 `options.messages`;`GenerateOptions.system` 为摘要器、标题提供方等直接单次调用方保留 |
+| `compaction-basic` 的 `buildSummarizationInput` | `header.system` + 区域消息 | 第 0 号节点的派生消息 + 区域消息,仍是已路由请求的真实前缀 |
+| `compaction-basic` 的 `selectCompactableRange` | 锚定在头部 `surfaceNodes[0]` | 锚定在首个非系统节点;第 0 号节点永不落入压缩范围 |
+| `dsh-token-meter` 的系统提示词估算 | `header.system` 长度 | 系统节点与其他每个 surface 节点一样计价;上下文明细按其 source 插件标注 |
+| Web 请求提示词卡片、轨迹请求 header 节点、请求检视 | 读取 `header.system` | 读取 `system/message` 节点;卡片保持折叠可检视的呈现,永不作为聊天气泡 |
+| 快照归一化器的 `{{system}}` 占位符、断言 `header.system` 的 plan-mode 测试 | header | 系统节点的文本 |
+| TypeScript 与 Python SDK 期望输出 | 没有系统事件 | 包含 `system/message` 事件 |
+| 人类转录投影(`isAppendSurfaceEvent` 的读取方) | 没有系统事件 | 跳过 `system/message`;它是模型历史,不是对话 |
+
+`RuntimeContextProjection` 与 `SystemPromptProjection` 是对称的:两者都通过 `sourceEventSeqs` 观察自己拥有的 surface 节点及其被遮蔽的情况,都把一条未提交的消息交给循环由 `turn()` 提交。区别在于角色与操作集——运行时上下文只追加 user 角色快照,系统提示词追加一次之后只做替换。
+
+## Alternatives considered
+
+**保留 `header.system`,只为更新添加 `system/message`。** 一个事实两个归属:上述每个消费方都要从 header 读消息 0、从 surface 读后续消息,循环还需要一个在 surface 存在系统节点时让 `headerEquals` 忽略 `system` 的特例。被否决,因为本次变更的目的就是单一表示。
+
+**用专门的仅记日志事件 `system-prompt/change` 重写 header。** 保留 header 作为提示词归属,并把变更记录为独立事件种类,但仍无法表达历史内部的系统消息,历史内替换提案还是需要第二套机制。被否决。
+
+**在适配器内根据相邻 header 合成系统消息。** 适配器逐请求无状态且从不接触日志;依赖适配器状态的协议历史无法从 surface 折叠重建。被否决。
+
+**像运行时上下文那样用 `user/message` 快照表达提示词。** 复用了现有事件类型,却发送了错误的角色,因此把系统消息视为权威的模型不会这样对待它。被否决。
+
+## Acceptance criteria
+
+- `SurfaceEventType` 包含 `system/message`;`deriveEventMessage` 投影它;`Session.append('system/message', …)` 与其他 surface 事件一样要求 `SurfaceIntent`。
+- `EpochHeader` 没有 `system` 字段;`headerEquals` 只比较 `config`、`adapterDefaults` 和 `tools`。
+- 渲染后的提示词非空的首个请求在该步骤首条 `user/message` 之前追加 `system/message` 作为 surface 第 0 号节点;提示词变更时以指明被遮蔽节点的 `sourceEventSeqs` 替换第 0 号节点;提示词不变时不追加任何内容。
+- 对同一会话历史,每个循环步骤的 DeepSeek 协议请求与今天逐字节一致:系统消息在先,随后是折叠后的对话。
+- 压缩永不选中第 0 号节点;摘要器回放的前缀以第 0 号节点的派生消息开头。
+- `dsh-token-meter`、Web 请求提示词卡片、轨迹与检视视图、快照归一化器、plan-mode 测试以及两个 SDK 的期望输出都读取系统节点;`dsh-agent-loop/invariant` 伴随组件重建请求时系统消息位于 `messages` 内。
+- 演练会话中途提示词变更(进入与退出 plan 模式)的无密钥录制快照显示被替换的第 0 号节点,而不是 `request/header` 的 `change`。
+- `docs/architecture.md`、`dsh-agent-loop`、`dsh-session`、`dsh-system-prompt`、`dsh-compaction-basic`、`dsh-token-meter` 的 README 以及可重建请求 Agent Note 都把 surface 节点描述为系统提示词的归属。
+
+## Risks
+
+- `header.system` 的每个读取方在一次变更中迁移;遗漏的读取方因字段消失而在编译期失败,这正是预期的失败方式。
+- 压缩范围选择新增一条不变量(第 0 号节点永不被压缩)。除 `compaction-basic` 以外、锚定在 `surfaceNodes[0]` 的压缩提供方会遮蔽提示词;`dsh-session` 的 surface 管理器拒绝在第 0 号节点是 `system/message` 时覆盖第 0 号节点的替换,除非替换事件本身是恰好覆盖该节点的 `system/message`,因此不变量在操作发生处被强制,而不只在随发的提供方中。位于更后位置的系统节点没有此类保护:压缩范围可以遮蔽它们。
+- 替换第 0 号节点会推进 `replaceGeneration`,今天有些读取方把它理解为「发生了压缩」;这些读取方改为检查替换事件的类型。
+- 日志中包含 `header.system` 的录制快照 fixture 需要重新录制;改变的是 fixture,而不是归一化器。

+ 6 - 0
.agents/notes/proposed/feature/2026-09-02-in-history-system-prompt-replacement.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/proposed/feature/2026-09-02-in-history-system-prompt-replacement.md
+2026-09-02-in-history-system-prompt-replacement.md: 229e92500936ec8742f341c4dc18184eb9a2fe0c
+2026-09-02-in-history-system-prompt-replacement.zh.md: f776f25a43f936024564488c1c53e9728fa19827

+ 74 - 0
.agents/notes/proposed/feature/2026-09-02-in-history-system-prompt-replacement.md

@@ -0,0 +1,74 @@
+# Agent Note: In-history system prompt replacement for cache-stable prompt changes
+
+Status: proposed
+
+English | [中文](2026-09-02-in-history-system-prompt-replacement.zh.md)
+
+## Problem
+
+Every system prompt change costs the whole provider prefix cache. The loop renders the prompt on every step; when the bytes differ — a plan-mode section entering or leaving, a skill or tool guidance section registering, an agent-scoped persona shadow, a changed `{{model}}` variable — the request's message 0 changes and the DeepSeek context cache misses from the first token. Long agentic sessions pay this repeatedly, and the [runtime-context snapshot design](../../archived/feature/2026-07-30-current-sandbox-policy-context.md) exists precisely because moving a changing fact out of the prompt was the only way to keep the prefix stable.
+
+A DeepSeek model, provided as an unpublished model fact for this proposal, removes that constraint: it accepts a `system` message at any position of the conversation and treats the latest one as the complete effective system prompt, replacing the leading one. Tool schemas remain part of the cached prefix, so a tool-set change still invalidates the cache. With that model the harness can append the new prompt after the cached history instead of rewriting message 0, and the prefix stays warm.
+
+The harness has the representation for this only after the [system prompt is surface node 0](../architecture/2026-09-02-system-prompt-as-surface-node.md): a prompt change is then an operation on `system/message` surface nodes, and the choice between "replace node 0" and "append a new node" is a per-model decision.
+
+## Proposal
+
+For a model route that declares the capability, the loop appends a new `system/message` surface node instead of replacing node 0 when the rendered prompt changes and the prefix would otherwise survive. Everything else in the [surface-node design](../architecture/2026-09-02-system-prompt-as-surface-node.md) is unchanged: the event type, the projection owner, the serializers, and the presentation.
+
+### Capability
+
+The DeepSeek adapter's catalog model gains a validated optional field, `systemPromptUpdate`, with the single accepted value `'in-history'`; absence means the model needs message 0 rewritten. The adapter surfaces it on `LlmResolvedModelInfo` and `prepareCall()` returns it beside `context.contextWindow`, so the loop reads it from the same registration-bound metadata it already consumes. No default catalog entry declares it until the model is released; a deployment enables it through the `models` list in `cordis.yml`. Models without the field — including every current default entry and every `dsh-llm-pi-ai` route — keep the replace-node-0 behaviour exactly.
+
+### The decision rule
+
+`SystemPromptProjection` tracks the **effective prompt**: the text of the latest surviving `system/message` on the surface (node 0 when no later system node exists). When the rendered prompt differs from the effective prompt:
+
+| Route capability | Prefix state | Operation |
+|---|---|---|
+| none | any | replace node 0 |
+| `in-history` | the current request series continues (no compaction since the last request, no tools or config change) | append a new `system/message` before the step's `user/message` events |
+| `in-history` | a new series starts (compaction replaced the surface, or `request/header` records a `change` for tools or config) and no mid-history system node survives | replace node 0 with the current prompt |
+| `in-history` | a new series starts but a mid-history system node survives | append a new `system/message`; node 0 stays as it is |
+
+The third row exists because a series start already costs the cache; folding the prompt back into node 0 keeps the history short. The fourth row exists because the surface has no delete operation: replacing node 0 while a later system node survives would leave the model reading the later, stale node as authoritative, so the loop appends instead. In-history mode never rewrites node 0 while any later system node survives.
+
+Resume follows the mid-session rule. A new loop instance restores the effective prompt from the log and, when the freshly rendered prompt differs, appends — the provider cache may still be warm across a process boundary, and the `resume` header is not a series start.
+
+### Presentation and accounting
+
+A mid-history `system/message` uses the same collapsed request-prompt inspection card as node 0, labelled as a prompt update at its position in the request; it is never a chat bubble, transcript projections skip it, and SDK projections expose it as a typed event. `dsh-token-meter` prices it like any other surface node, so the per-step context breakdown shows the accumulated cost of retained prompt versions until compaction shadows them. `cacheReadTokens` on the following `assistant/message` usage is the observable effect: for a capable route the value covers the prefix through the last cached message; for a non-capable route it drops to the shared-prefix detection floor.
+
+### Verification plan
+
+- Unit tests in `dsh-agent-loop` for the projection: append on a mid-series change, replace on a series start without surviving mid-history nodes, append on a series start with one, append on resume, no operation when unchanged, and replace-only behaviour for a route without the capability.
+- Unit tests in `dsh-llm-deepseek` for catalog validation (`systemPromptUpdate` accepts `'in-history'` only) and for `prepareCall()` surfacing the field.
+- A keyless recorded snapshot under `snapshots/` whose composition declares the capability on the mock route and toggles plan mode mid-session, pinning the appended `system/message` and the untouched node 0; TypeScript and Python SDK expected outputs include the appended event.
+- A real-API e2e that runs two steps with a prompt change against a capable route and asserts that the second request's `cacheReadTokens` is at least the first request's prompt token count. It resolves its route from the standard credential and base-URL mechanism and self-skips when no capable route is configured.
+
+## Alternatives considered
+
+**Send only the changed sections as a delta.** The model treats the latest system message as the complete prompt, so a delta would silently drop every unchanged section. Rejected on the model contract.
+
+**Enable in-history mode by plugin config instead of a model capability.** A deployment flag could pair a non-capable model with appended system messages, which such a model would read as ordinary history at best. The capability belongs to the route that honours it; the adapter catalog already carries per-model capacities. Rejected.
+
+**Always append, never re-baseline.** One rule, but node 0 would stay stale for the life of the session and every request after compaction would carry the stale head plus the replacement. Re-baselining at a series start costs nothing extra because the cache is already lost there. Rejected.
+
+**Re-baseline on every resume.** Accepts one cache miss per process restart for a simpler resume path. The cache persists across restarts for hours to days, and the log already carries what resume needs. Rejected.
+
+**Place the system message after the step's user messages.** Both positions sit after the cached prefix, but the model then reads the instructions after the input it must apply them to; system-before-user matches the leading position's ordering. Rejected.
+
+## Acceptance criteria
+
+- `DeepSeekCatalogModel.systemPromptUpdate` is validated at load, exposed through `LlmResolvedModelInfo`, and returned by `prepareCall()`; a misspelt value fails at load.
+- On a capable route a mid-series prompt change appends `system/message` before the step's `user/message` events and node 0 is unchanged; on a non-capable route the same change replaces node 0.
+- On a capable route a series start with no surviving mid-history system node replaces node 0; with a surviving one it appends.
+- A resumed loop instance whose rendered prompt differs appends on a capable route.
+- The recorded snapshot and both SDK expected outputs pin the appended event; the e2e asserts the cache-hit inequality when a capable route is configured and skips otherwise.
+- The `dsh-llm-deepseek`, `dsh-agent-loop`, and `dsh-system-prompt` READMEs document the capability, the decision rule, and the KV Cache effect; `docs/config-catalog.md` lists the field.
+
+## Risks
+
+- The model contract is unpublished; the note records it as provided. If the released model narrows it (for example, honouring only the latest system message within a bounded window), the decision rule needs a re-baseline trigger beyond series starts.
+- Retained prompt versions accumulate in history until compaction shadows them. Each version costs its tokens on every request in the series; a deployment whose prompt changes on most steps would be better served by moving that fact into runtime context.
+- A proxy that rewrites or reorders system messages breaks the replacement semantics silently; the e2e's cache-hit assertion is the detector.

+ 74 - 0
.agents/notes/proposed/feature/2026-09-02-in-history-system-prompt-replacement.zh.md

@@ -0,0 +1,74 @@
+# Agent Note: 历史内系统提示词替换,实现缓存稳定的提示词变更
+
+Status: proposed
+
+[English](2026-09-02-in-history-system-prompt-replacement.md) | 中文
+
+## Problem
+
+每一次系统提示词变更都要付出整个提供方前缀缓存的代价。循环在每个步骤渲染提示词;一旦字节不同——plan 模式片段进入或退出、某个 skill 或工具指引片段完成注册、agent 作用域的 persona 遮蔽、`{{model}}` 变量改变——请求的消息 0 随之改变,DeepSeek 上下文缓存从第一个 token 起失效。长时间的 agent 会话反复为此付费,而[运行时上下文快照设计](../../archived/feature/2026-07-30-current-sandbox-policy-context.md)之所以存在,正是因为把会变化的事实移出提示词是保持前缀稳定的唯一办法。
+
+一个 DeepSeek 模型——作为本提案所依据的未公开模型事实——移除了这一限制:它接受对话任意位置的 `system` 消息,并把最新一条视为完整的有效系统提示词,替换最前面那条。工具 schema 仍属于被缓存的前缀,因此工具集变更仍会使缓存失效。有了这样的模型,harness 可以把新提示词追加到已缓存的历史之后而不是重写消息 0,前缀就能保持热态。
+
+只有在[系统提示词成为 surface 第 0 号节点](../architecture/2026-09-02-system-prompt-as-surface-node.zh.md)之后,harness 才拥有实现这一点的表示:提示词变更随之成为对 `system/message` surface 节点的操作,而「替换第 0 号节点」与「追加新节点」之间的选择是逐模型的决定。
+
+## Proposal
+
+对于声明了该能力的模型路由,当渲染后的提示词变化且前缀本可存活时,循环追加一个新的 `system/message` surface 节点而不是替换第 0 号节点。[surface 节点设计](../architecture/2026-09-02-system-prompt-as-surface-node.zh.md)中的其他一切不变:事件类型、投影的拥有者、序列化器和呈现。
+
+### 能力
+
+DeepSeek 适配器的目录模型新增一个经校验的可选字段 `systemPromptUpdate`,唯一接受的值是 `'in-history'`;缺省表示该模型需要重写消息 0。适配器把它暴露在 `LlmResolvedModelInfo` 上,`prepareCall()` 在 `context.contextWindow` 旁边返回它,因此循环从它已经消费的同一份注册绑定元数据中读取。在该模型发布之前,没有默认目录条目声明它;部署方通过 `cordis.yml` 的 `models` 列表启用。没有该字段的模型——包括当前所有默认条目和所有 `dsh-llm-pi-ai` 路由——完全保持替换第 0 号节点的行为。
+
+### 决策规则
+
+`SystemPromptProjection` 跟踪**有效提示词**:surface 上最新存活的 `system/message` 的文本(不存在更后的系统节点时即第 0 号节点)。当渲染后的提示词与有效提示词不同时:
+
+| 路由能力 | 前缀状态 | 操作 |
+|---|---|---|
+| 无 | 任意 | 替换第 0 号节点 |
+| `in-history` | 当前请求序列延续(上次请求以来没有压缩,tools 或 config 没有变更) | 在该步骤的 `user/message` 事件之前追加新的 `system/message` |
+| `in-history` | 新序列开始(压缩替换了 surface,或 `request/header` 记录了 tools 或 config 的 `change`)且没有历史中途的系统节点存活 | 用当前提示词替换第 0 号节点 |
+| `in-history` | 新序列开始但有历史中途的系统节点存活 | 追加新的 `system/message`;第 0 号节点保持原样 |
+
+第三行存在,是因为序列开始已经付出了缓存代价;把提示词折回第 0 号节点能让历史保持简短。第四行存在,是因为 surface 没有删除操作:在更后的系统节点仍存活时替换第 0 号节点,会让模型把更后、已过时的节点当作权威,所以循环改为追加。历史内模式在任何更后的系统节点存活期间永不重写第 0 号节点。
+
+恢复遵循会话中途的规则。新的循环实例从日志恢复有效提示词,当新渲染的提示词不同时执行追加——提供方缓存在进程边界之后可能仍是热的,且 `resume` header 不是序列开始。
+
+### 呈现与记账
+
+历史中途的 `system/message` 使用与第 0 号节点相同的折叠请求提示词检视卡片,在请求中的对应位置标注为提示词更新;它永不作为聊天气泡,转录投影跳过它,SDK 投影把它暴露为带类型的事件。`dsh-token-meter` 像对待其他任何 surface 节点一样为它计价,因此逐步骤的上下文明细会显示被保留的各个提示词版本累计的开销,直到压缩遮蔽它们。随后 `assistant/message` 用量上的 `cacheReadTokens` 是可观察的效果:对具备能力的路由,该值覆盖到最后一条已缓存消息为止的前缀;对不具备能力的路由,它回落到公共前缀检测的下限。
+
+### 验证计划
+
+- `dsh-agent-loop` 中针对投影的单元测试:序列中途变更时追加、没有存活的历史中途节点时在序列开始处替换、有存活节点时在序列开始处追加、恢复时追加、未变更时无操作,以及不具备能力的路由只做替换。
+- `dsh-llm-deepseek` 中针对目录校验(`systemPromptUpdate` 只接受 `'in-history'`)和 `prepareCall()` 暴露该字段的单元测试。
+- `snapshots/` 下的一个无密钥录制快照,其组合在 mock 路由上声明该能力并在会话中途切换 plan 模式,钉住追加的 `system/message` 与未被触及的第 0 号节点;TypeScript 与 Python SDK 的期望输出包含追加的事件。
+- 一个真实 API 的 e2e:针对具备能力的路由运行两个步骤并夹带一次提示词变更,断言第二个请求的 `cacheReadTokens` 不小于第一个请求的提示词 token 数。它通过标准的凭据与 base-URL 机制解析路由,未配置具备能力的路由时自动跳过。
+
+## Alternatives considered
+
+**只发送变化的片段作为增量。** 模型把最新的系统消息当作完整提示词,因此增量会静默丢掉每个未变化的片段。基于模型约定被否决。
+
+**用插件配置而不是模型能力启用历史内模式。** 部署标志可能把不具备能力的模型与追加的系统消息配对,这样的模型最多把它们当作普通历史。该能力属于兑现它的路由;适配器目录已经承载逐模型的容量信息。被否决。
+
+**永远追加,从不重新基线化。** 规则单一,但第 0 号节点会在会话整个生命周期内保持过时,压缩之后的每个请求都要携带过时的头部加替换消息。在序列开始处重新基线化不花额外代价,因为缓存在那里已经丢失。被否决。
+
+**每次恢复都重新基线化。** 为更简单的恢复路径接受每次进程重启一次缓存未命中。缓存跨重启持续数小时到数天,而日志已经承载恢复所需的一切。被否决。
+
+**把系统消息放在该步骤的用户消息之后。** 两个位置都在已缓存前缀之后,但模型会在读到必须应用指令的输入之后才读到指令;system 在 user 之前与最前位置的顺序一致。被否决。
+
+## Acceptance criteria
+
+- `DeepSeekCatalogModel.systemPromptUpdate` 在加载时校验、通过 `LlmResolvedModelInfo` 暴露、由 `prepareCall()` 返回;拼错的值在加载时失败。
+- 在具备能力的路由上,序列中途的提示词变更在该步骤的 `user/message` 事件之前追加 `system/message`,第 0 号节点不变;在不具备能力的路由上,同样的变更替换第 0 号节点。
+- 在具备能力的路由上,没有存活的历史中途系统节点的序列开始替换第 0 号节点;有存活节点时追加。
+- 渲染后的提示词不同的已恢复循环实例在具备能力的路由上追加。
+- 录制快照与两个 SDK 的期望输出钉住追加的事件;配置了具备能力的路由时 e2e 断言缓存命中不等式,否则跳过。
+- `dsh-llm-deepseek`、`dsh-agent-loop`、`dsh-system-prompt` 的 README 记录该能力、决策规则和 KV Cache 效果;`docs/config-catalog.md` 列出该字段。
+
+## Risks
+
+- 模型约定尚未公开;本 Agent Note 按所提供的内容记录。若发布的模型收窄了约定(例如只在有界窗口内兑现最新的系统消息),决策规则需要序列开始之外的重新基线化触发条件。
+- 被保留的提示词版本在历史中累积,直到压缩遮蔽它们。每个版本在该序列的每个请求上都要付出其 token 开销;提示词在多数步骤都变化的部署,更适合把那个事实移入运行时上下文。
+- 重写或重排系统消息的代理会静默破坏替换语义;e2e 的缓存命中断言是探测器。

+ 92 - 0
.agents/skills/dsh-speed-up-perf/SKILL.md

@@ -0,0 +1,92 @@
+---
+name: dsh-speed-up-perf
+description: 'Use when investigating or optimizing DeepSeek Harness performance, designing realistic synthetic benchmarks or CI performance gates, profiling long Sessions or Web responsiveness, or turning performance PR evidence into measured behavior-preserving fixes.'
+---
+
+# Speed Up DeepSeek Harness
+
+Turn a broad “make it faster” request into reproducible user-path measurements and small, evidence-backed fixes. This is guidance, not a quota or a script: survey broadly, follow measured cost, and reject attractive changes that do not improve the workload users actually run.
+
+## Establish scope and current authority
+
+Read [AGENTS.md](../../../AGENTS.md), [architecture](../../../docs/architecture.md), [testing policy](../../../docs/testing.md), [defensive patterns](../../../docs/defensive-patterns.md), and the affected packages’ instructions and Agent Notes. Use [CI test reliability](../dsh-ci-test-reliability/SKILL.md) for processes, clocks, browser tests, and asynchronous cleanup.
+
+Agree on the user-visible endpoint, workload range, resource constraints, acceptable minor behavior differences, and stopping rule. Keep backend and browser end-to-end measurements separate: a fast history iterator or Client fold does not prove fast transport, paint, scrolling, or input response. Exclude model/network latency when measuring local overhead, and state that exclusion rather than calling the result complete product latency.
+
+Inspect the exact current base, not just the running checkout. Study final merged diffs, owning source, tests, and resolved review threads; a PR body can describe an abandoned implementation. Separate merged, closed-unmerged, superseded, estimated, and newly measured evidence. The [performance workflow decision and evidence](../../notes/implemented/process/2026-09-06-evidence-driven-performance-skill.md) supply historical leads, not authority to reintroduce their implementations.
+
+## Survey user paths, then rank candidates
+
+Delegate independent domains when breadth helps; require measurements and production call sites, not guesses. Useful domains include:
+
+- Cold profile startup, first historical read, current-generation reopen, and writable resume.
+- Many-turn and tool-heavy history, large individual messages/results, child Session listing, and repeated navigation among Sessions.
+- Initial history transport and fold, first usable browser paint, older-page loading, scrolling, tool expansion, and inactive-view activation.
+- Live streaming and reconnect, including a long active attempt, interleaved tool work, settlement, cancellation, and teardown.
+
+Vary independent cost drivers: bytes, durable events, compact records, raw deltas, turns, tools, children, and visible DOM nodes are different quantities. Do not call a large count of tiny identical messages “realistic” without checking which user operation it stresses. Include typical and tail workloads, but avoid a combinatorial matrix with no decision value.
+
+Rank candidates by observed user latency, CPU/allocations, retained memory, occurrence, and confidence. For each, name the production consumer, the repeated work, the expected complexity, the smallest falsifiable intervention, and the behavior that must remain stable. A suspicious loop, unused cache, or large file alone is not evidence of a bottleneck.
+
+## Build realistic synthetic benchmarks first
+
+Follow [benchmarks/AGENTS.md](../../../benchmarks/AGENTS.md) and the [performance-gate decision](../../notes/implemented/testing/2026-09-04-session-open-performance-gate.md). Extend the existing required lane rather than creating competing calibration or reporting infrastructure. Package-local diagnostics remain beside their owner; cross-package required cases live under the measured user path in `benchmarks/`.
+
+If the user authorizes local corpus inspection, extract only aggregate workload characteristics. Never copy prompts, outputs, paths, identities, IDs, credentials, recordings, or recognizable snippets into fixtures, logs, screenshots, PRs, or artifacts. Generate fixed inputs from reviewed constants; no benchmark depends on the user’s home, ambient repository, network service, or private data.
+
+Before implementation, record a measurement card:
+
+| Field | Required decision |
+|---|---|
+| User operation | Exact action and externally observable completion condition |
+| Workload | Fixed dimensions, distributions, construction seed/constants, and why they exercise ordinary and tail use |
+| Entry path | Production calls/composition and built artifacts; mocked external boundaries |
+| Clock | Included setup, cold/warm state, timing start/end, and excluded costs |
+| Memory | Reachable endpoint objects, baseline, GC policy, retained versus transient limits |
+| Verdict | Raw samples, chosen aggregate, calibrated absolute/ratio/memory limits, and negative control |
+| Behavior | Owning functional tests/snapshots and permitted minor differences |
+
+Measure built JavaScript under plain Node for CPU workers; source-loader overhead and module resolution are not the shipped path. Browser cases use built product assets and the supported `dsh` profile through the existing test harness. Do not add a production export solely for measurement or copy the algorithm into a “benchmark implementation.”
+
+Use fresh children and private temporary roots for cold/process-memory samples. Warm samples explicitly retain the intended cache; never let fixture setup secretly warm a cold scenario. Keep the same input, validations, completion condition, and reachable output on both sides. A parse-and-discard baseline is not comparable with validated retained history.
+
+Report all samples and the aggregate that decides the result. For the Node lane, use the existing shared time calibration and reviewed variance headroom; do not scale bytes, counts, or dimensionless ratios by CPU speed. Keep manual browser diagnostics threshold-free. A required browser performance case needs an explicit lane decision and repeated measurements on its actual CI browser/runner before adopting timing budgets; the Node machine multiplier alone is not browser calibration. Budgets are source constants, not environment overrides. Serialize measured work against other owned CPU-heavy jobs; measure reference and candidate under comparable conditions. Do not widen a budget or select a lucky run to hide a regression.
+
+Measure end-to-end latency independently from component phases. Track retained memory with intended objects still reachable, and transient pressure separately through constrained-heap completion or an appropriate peak measurement. Faster execution with unbounded retention is not an automatic win.
+
+For browser responsiveness, use real browser input and observe the resulting UI update. Include the final stall in frame/input measurements, distinguish scheduled timers from actual input, and bound synthetic producers so catch-up bursts do not invent a different workload. State whether first paint, scrolling, paging, live updates, and activated-but-hidden views are covered. Node folds, fake DOMs, and custom heartbeat events alone cannot establish browser responsiveness.
+
+## Prove the regression, then remove work
+
+Run the unoptimized workload before changing production code. Save the command, revision, runtime/platform, fixture dimensions, raw measurements, and verdict. Reduce a failing scenario until it still exercises the real bottleneck, then rank falsifiable hypotheses before patching. Use profiles, allocation samples, work counts, or phase timings to distinguish them.
+
+Common patterns worth testing, not automatic prescriptions:
+
+- Keep compact representations compact through downstream readers; avoid per-delta objects when the consumer needs settled content or one aggregate.
+- Remove duplicate parsing, copying, freezing, and validation only after identifying the actual ownership and trust transition. Typed same-process borrowing is not permission to weaken durable or wire parsing.
+- Stream artifact transformations and bound intermediate state rather than retaining every generation. Include publication, verification, and writable-readiness obligations where the user operation requires them.
+- Separate read-only preparation from write/publication work without moving awaited work past a correctness-required endpoint.
+- Defer inactive-view and collapsed-detail work; measure first activation and retained state too. Deferral is not deletion, and viewport highlighting is not full virtualization.
+- Stabilize identities and narrow subscriptions so one changed node does not invalidate an entire history; preserve update ordering and immediate-event behavior.
+- Prefer a suitable data structure to repeated shifting, scanning, or rebuilding. Measure the whole consumer path, not just the isolated container operation.
+- Use revision-keyed reuse or singleflight only with explicit invalidation, bounded retention, independent waiter cancellation, and disposal ownership. Avoid caching expanded representations merely to make repeated benchmarks look fast.
+
+Change one causal factor at a time. Re-run both the focused scenario and its end-to-end parent. Require a negative control: the tightened assertion fails on the original implementation or a controlled reintroduction of the targeted cost. A threshold so generous that the regression passes is not protection; a budget below a verified noise floor is not reliable either.
+
+## Preserve behavior and resource ownership
+
+Performance measurements complement functional evidence; they do not replace it. Run or add the narrow owning tests for output, ordering, paging, stream indexes, errors, cancellation, concurrency, and disposal as applicable. Preserve model-visible/logged equivalence, released-generation immutability, atomic publication, required validation, and writable readiness. Do not silently truncate history, skip tool results, disable invariants, or change lifecycle semantics to reach a number.
+
+State any deliberate minor visible difference and verify it through the owning keyless snapshot. For a product-visible GUI change, include the required browser evidence/GIF. Keep functional expectations independent of benchmark internals; benchmark assertions need enough evidence to reach the real endpoint, not a second semantic test suite.
+
+Reject an optimization when gains disappear end-to-end, a typical workload regresses materially, complexity outweighs a small gain, or cancellation/retention/durability cannot be explained and tested. Record the rejected hypothesis briefly instead of expanding scope to justify it.
+
+## Deliver a bounded, reviewable result
+
+Use [Agent Note rules](../../notes/README.md) for durable rationale, alternatives, calibration, exclusions, and remaining risks. Check relevant notes for supersession without turning performance work into a corpus-wide prose cleanup. Keep the reusable procedure here and scenario-specific truth with its benchmark or package owner.
+
+When the task requests stacked PRs, choose layers before editing and use official GitHub stacks and separate worktrees. Keep each layer mergeable: benchmark infrastructure can protect the measured baseline; the optimization layer carries its fix, functional coverage, and tighter budget. Independent bottlenecks may use separate stacks. Fix a finding in its owning layer before propagating upward.
+
+Apply [pre-push checks](../dsh-pre-push-checks/SKILL.md), report only executed evidence, and inspect CI rather than assuming local timing proves runner stability. After marking ready, evaluate review findings against code and executable evidence; reply with the reason or fix and resolve addressed threads. Do not dismiss a report merely because it came from a bot.
+
+Summarize each result as: workload → before/after absolute values and ratio → endpoint and memory semantics → behavior evidence → negative control → exact checks → exclusions. Separate author-reported historical numbers, fresh local measurements, and CI evidence. Stop at the agreed scenario/fix scope; retain a short ranked follow-up list instead of chasing unrelated opportunities.

+ 16 - 13
.github/ISSUE_TEMPLATE/bug.md

@@ -1,22 +1,25 @@
 ---
 ---
 name: Bug
 name: Bug
 about: 记录现有预期行为的失效
 about: 记录现有预期行为的失效
-title: ''
-labels: ''
-assignees: ''
 type: Bug
 type: Bug
 ---
 ---
 
 
-<!-- 标题写中文行动或结果句;外露正文不超过 50 单位。 -->
-一句话说明错误结果。
+## Summary
 
 
-<details>
-<summary>复现、预期与验收</summary>
+<!-- 简要说明发生了什么错误,以及受影响的用户或场景。 -->
 
 
-- 复现步骤:
-- 实际结果:
-- 预期结果:
-- 环境:
-- 验收条件:
+## Reproduction
 
 
-</details>
+<!-- 列出能稳定触发问题的最小步骤、输入或代码。 -->
+
+## Current behavior
+
+<!-- 说明实际结果,并附上必要的错误信息、日志或截图。 -->
+
+## Expected behavior
+
+<!-- 说明正确结果。 -->
+
+## Environment
+
+<!-- 说明相关版本、平台、配置或运行条件。 -->

+ 0 - 1
.github/ISSUE_TEMPLATE/config.yml

@@ -1,2 +1 @@
 blank_issues_enabled: false
 blank_issues_enabled: false
-contact_links: []

+ 4 - 11
.github/ISSUE_TEMPLATE/feature.md

@@ -1,20 +1,13 @@
 ---
 ---
 name: Feature
 name: Feature
 about: 新增或有意改变可观察行为
 about: 新增或有意改变可观察行为
-title: ''
-labels: ''
-assignees: ''
 type: Feature
 type: Feature
 ---
 ---
 
 
-<!-- 标题写中文行动或结果句;外露正文不超过 50 单位。 -->
-一句话说明预期结果。
+## Motivation
 
 
-<details>
-<summary>验收与细节</summary>
+<!-- 说明当前问题、受影响的用户,以及为什么需要这项变化。 -->
 
 
-- 验收条件:
-- 用户或模型可见变化:
-- 测试证据:
+## Behavior
 
 
-</details>
+<!-- 说明预期的用户、模型或系统可观察行为。 -->

Alguns ficheiros não foram mostrados porque muitos ficheiros mudaram neste diff