Kaynağa Gözat

feat: enforce maintained repository reference policy

Tianyi Cui 5 gün önce
ebeveyn
işleme
6b651380a7
100 değiştirilmiş dosya ile 541 ekleme ve 288 silme
  1. 2 2
      .agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.i18n.yaml
  2. 1 1
      .agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.md
  3. 1 1
      .agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.zh.md
  4. 2 2
      .agents/notes/implemented/architecture/2026-08-27-outbound-proxy-policy.i18n.yaml
  5. 1 1
      .agents/notes/implemented/architecture/2026-08-27-outbound-proxy-policy.md
  6. 1 1
      .agents/notes/implemented/architecture/2026-08-27-outbound-proxy-policy.zh.md
  7. 2 2
      .agents/notes/implemented/architecture/2026-08-30-retain-ignorable-external-session-events.i18n.yaml
  8. 1 1
      .agents/notes/implemented/architecture/2026-08-30-retain-ignorable-external-session-events.md
  9. 1 1
      .agents/notes/implemented/architecture/2026-08-30-retain-ignorable-external-session-events.zh.md
  10. 2 2
      .agents/notes/implemented/bug-fix/2026-08-17-subagent-message-settlement-ordering.i18n.yaml
  11. 1 1
      .agents/notes/implemented/bug-fix/2026-08-17-subagent-message-settlement-ordering.md
  12. 1 1
      .agents/notes/implemented/bug-fix/2026-08-17-subagent-message-settlement-ordering.zh.md
  13. 2 2
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.i18n.yaml
  14. 2 2
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md
  15. 2 2
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.zh.md
  16. 2 2
      .agents/notes/implemented/bug-fix/2026-09-07-typert-package-local-forwarding-imports.i18n.yaml
  17. 1 1
      .agents/notes/implemented/bug-fix/2026-09-07-typert-package-local-forwarding-imports.md
  18. 1 1
      .agents/notes/implemented/bug-fix/2026-09-07-typert-package-local-forwarding-imports.zh.md
  19. 2 2
      .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.i18n.yaml
  20. 12 12
      .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.md
  21. 12 12
      .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.zh.md
  22. 2 2
      .agents/notes/implemented/process/2026-09-06-node-compatibility-selfhosted.i18n.yaml
  23. 1 1
      .agents/notes/implemented/process/2026-09-06-node-compatibility-selfhosted.md
  24. 1 1
      .agents/notes/implemented/process/2026-09-06-node-compatibility-selfhosted.zh.md
  25. 2 2
      .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.i18n.yaml
  26. 2 2
      .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.md
  27. 2 2
      .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.zh.md
  28. 2 2
      .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.i18n.yaml
  29. 1 1
      .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.md
  30. 1 1
      .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.zh.md
  31. 2 2
      .agents/notes/implemented/process/2026-09-09-blacksmith-failover-leg.i18n.yaml
  32. 0 0
      .agents/notes/implemented/process/2026-09-09-blacksmith-failover-leg.md
  33. 0 0
      .agents/notes/implemented/process/2026-09-09-blacksmith-failover-leg.zh.md
  34. 6 0
      .agents/notes/implemented/process/2026-09-12-maintained-repository-references.i18n.yaml
  35. 31 0
      .agents/notes/implemented/process/2026-09-12-maintained-repository-references.md
  36. 31 0
      .agents/notes/implemented/process/2026-09-12-maintained-repository-references.zh.md
  37. 2 2
      .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.i18n.yaml
  38. 1 1
      .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.md
  39. 1 1
      .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.zh.md
  40. 2 2
      .agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-evidence.i18n.yaml
  41. 6 6
      .agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-evidence.md
  42. 6 6
      .agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-evidence.zh.md
  43. 2 2
      .agents/notes/implemented/simplification/2026-09-07-file-content-scan.i18n.yaml
  44. 1 1
      .agents/notes/implemented/simplification/2026-09-07-file-content-scan.md
  45. 1 1
      .agents/notes/implemented/simplification/2026-09-07-file-content-scan.zh.md
  46. 2 2
      .agents/notes/implemented/simplification/2026-09-09-nontransactional-loader.i18n.yaml
  47. 1 1
      .agents/notes/implemented/simplification/2026-09-09-nontransactional-loader.md
  48. 1 1
      .agents/notes/implemented/simplification/2026-09-09-nontransactional-loader.zh.md
  49. 2 2
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml
  50. 4 4
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
  51. 4 4
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md
  52. 2 2
      .agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.i18n.yaml
  53. 5 5
      .agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.md
  54. 5 5
      .agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.zh.md
  55. 2 2
      .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.i18n.yaml
  56. 5 5
      .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md
  57. 5 5
      .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.zh.md
  58. 2 2
      .agents/notes/implemented/testing/2026-09-07-pwsh-ci-observable-completion.i18n.yaml
  59. 2 2
      .agents/notes/implemented/testing/2026-09-07-pwsh-ci-observable-completion.md
  60. 2 2
      .agents/notes/implemented/testing/2026-09-07-pwsh-ci-observable-completion.zh.md
  61. 2 2
      .agents/notes/implemented/testing/2026-09-07-subagent-teardown-test-budgets.i18n.yaml
  62. 1 1
      .agents/notes/implemented/testing/2026-09-07-subagent-teardown-test-budgets.md
  63. 1 1
      .agents/notes/implemented/testing/2026-09-07-subagent-teardown-test-budgets.zh.md
  64. 2 2
      .agents/notes/implemented/testing/2026-09-08-ci-completion-observations.i18n.yaml
  65. 1 1
      .agents/notes/implemented/testing/2026-09-08-ci-completion-observations.md
  66. 1 1
      .agents/notes/implemented/testing/2026-09-08-ci-completion-observations.zh.md
  67. 2 2
      .agents/notes/implemented/testing/2026-09-08-ci-readiness-and-completion.i18n.yaml
  68. 5 5
      .agents/notes/implemented/testing/2026-09-08-ci-readiness-and-completion.md
  69. 5 5
      .agents/notes/implemented/testing/2026-09-08-ci-readiness-and-completion.zh.md
  70. 2 2
      .agents/notes/implemented/testing/2026-09-09-user-patch-hmr-test-delivery.i18n.yaml
  71. 1 1
      .agents/notes/implemented/testing/2026-09-09-user-patch-hmr-test-delivery.md
  72. 1 1
      .agents/notes/implemented/testing/2026-09-09-user-patch-hmr-test-delivery.zh.md
  73. 10 10
      THIRD_PARTY_NOTICES.md
  74. 1 1
      apps/cli/config/examples/github-review/cordis.yml
  75. 1 1
      apps/cli/tests/fixtures/github-webhook/cordis.yml
  76. 2 2
      apps/cli/tests/github-webhook-real.e2e.ts
  77. 4 4
      apps/web/tests/expected/github-ready-review/conversation-expanded.expected.md
  78. 4 4
      apps/web/tests/expected/github-ready-review/conversation.expected.md
  79. 3 3
      apps/web/tests/github-ready-review.e2e.ts
  80. 2 2
      docs/AGENTS.md
  81. 2 2
      docs/session-format-status.i18n.yaml
  82. 3 3
      docs/session-format-status.md
  83. 3 3
      docs/session-format-status.zh.md
  84. 2 2
      docs/user/develop/basic/publish.i18n.yaml
  85. 1 1
      docs/user/develop/basic/publish.md
  86. 1 1
      docs/user/develop/basic/publish.zh.md
  87. 2 0
      native/system/docs/packaging.md
  88. 2 2
      native/system/package.json
  89. 1 1
      native/system/packages/darwin-arm64/package.json
  90. 1 1
      native/system/packages/darwin-x64/package.json
  91. 1 1
      native/system/packages/entry/package.json
  92. 1 1
      native/system/packages/linux-arm64/package.json
  93. 1 1
      native/system/packages/linux-x64/package.json
  94. 72 28
      native/system/scripts/pack-release.mjs
  95. 134 0
      native/system/test/release-packing.test.js
  96. 1 0
      package.json
  97. 3 8
      scripts/check-workspace-constraints.ts
  98. 37 42
      scripts/doc-standard.spec.ts
  99. 6 6
      scripts/gen-third-party-notices.ts
  100. 14 0
      scripts/run-gates.spec.ts

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.md
-2026-08-03-per-session-agent-presets.md: c2c9f670df480662804e086fbf150c773d8f4fc8
-2026-08-03-per-session-agent-presets.zh.md: 9a4760be9e4a9d1d2de0a7f8d2d3bce755c04480
+2026-08-03-per-session-agent-presets.md: 47f2e28c2aac583ce4ca810157e9c7abe58256d5
+2026-08-03-per-session-agent-presets.zh.md: 88b643ea710364f18ff7c73a8ca0e10de0fbb44c

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.md

@@ -27,7 +27,7 @@ The presets the deployment ships are the directories under `packages/preset/agen
 
 Mounting is per-session by default. Measured cost for a twelve-row composition is ~3ms and ~600KB per session, so isolation is the cheaper default than any sharing scheme, and a preset authored by a user or by an agent then has the smallest possible blast radius. A preset that genuinely owns an expensive singleton opts into sharing with Cordis's own `isolate` vocabulary: a named realm label is process-global, so two subtrees naming the same label resolve one instance.
 
-The `agent-presets` user-settings namespace carries `modeSelectionEnabled` and `default`. `modeSelectionEnabled` defaults to `true`: the existing new-session picker remains present and an unnamed session resolves to the saved user `default`, or the composition's deployment `default` when none exists. The Web Settings toggle changes only that policy: disabling selection temporarily uses the deployment default, while re-enabling it restores the saved user `default`. This is a deliberate exception to the ordinary user-over-composition settings precedence established in [#1539](https://github.com/deepseek-harness/deepseek-harness/pull/1539): hiding the chooser disables the user's mode-selection policy without deleting its saved value. The Host policy governs every later session whose caller omits a preset; explicitly named presets and existing sessions remain unchanged. The composition value also keeps the package working with no settings provider, while an enabled user override changes later sessions without editing a deployment-owned `cordis.yml`.
+The `agent-presets` user-settings namespace carries `modeSelectionEnabled` and `default`. `modeSelectionEnabled` defaults to `true`: the existing new-session picker remains present and an unnamed session resolves to the saved user `default`, or the composition's deployment `default` when none exists. The Web Settings toggle changes only that policy: disabling selection temporarily uses the deployment default, while re-enabling it restores the saved user `default`. This is a deliberate exception to the ordinary user-over-composition settings precedence established in #1539: hiding the chooser disables the user's mode-selection policy without deleting its saved value. The Host policy governs every later session whose caller omits a preset; explicitly named presets and existing sessions remain unchanged. The composition value also keeps the package working with no settings provider, while an enabled user override changes later sessions without editing a deployment-owned `cordis.yml`.
 
 ## Consequences
 

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-03-per-session-agent-presets.zh.md

@@ -27,7 +27,7 @@ Status: implemented
 
 挂载默认按会话进行。实测一份十二行组装每会话约 3ms、约 600KB,因此隔离比任何共享方案都更划算;而由用户或 agent 写出的 preset 也因此拥有尽可能小的影响面。确实自带昂贵单例的 preset,可以用 Cordis 自身的 `isolate` 词汇显式选择共享:命名 realm 的 label 是进程级全局的,因此两棵子树只要写同一个 label 就解析到同一个实例。
 
-`agent-presets` 用户设置命名空间同时携带 `modeSelectionEnabled` 与 `default`。`modeSelectionEnabled` 默认为 `true`:既有的新建会话选择器保持显示;未指名会话会解析到已保存的用户 `default`,尚未保存时则使用组装中 `default` 指定的部署默认值。Web 设置开关只改变该策略:关闭选择时临时使用部署默认值,再次开启时恢复已保存的用户 `default`。这是对 [#1539](https://github.com/deepseek-harness/deepseek-harness/pull/1539) 所确立“用户值覆盖组装值”这一普通 settings 优先级的有意例外:隐藏选择器会停用用户的模式选择策略,但不会删除其保存值。该 Host 策略适用于此后所有未显式指定 preset 的会话;显式指定及既有会话不受影响。组装值还使本包在没有 settings 提供方时照常工作;选择器开启后,用户可覆盖默认值来改变后续会话,而无需编辑部署所拥有的 `cordis.yml`。
+`agent-presets` 用户设置命名空间同时携带 `modeSelectionEnabled` 与 `default`。`modeSelectionEnabled` 默认为 `true`:既有的新建会话选择器保持显示;未指名会话会解析到已保存的用户 `default`,尚未保存时则使用组装中 `default` 指定的部署默认值。Web 设置开关只改变该策略:关闭选择时临时使用部署默认值,再次开启时恢复已保存的用户 `default`。这是对 #1539 所确立“用户值覆盖组装值”这一普通 settings 优先级的有意例外:隐藏选择器会停用用户的模式选择策略,但不会删除其保存值。该 Host 策略适用于此后所有未显式指定 preset 的会话;显式指定及既有会话不受影响。组装值还使本包在没有 settings 提供方时照常工作;选择器开启后,用户可覆盖默认值来改变后续会话,而无需编辑部署所拥有的 `cordis.yml`。
 
 ## 后果
 

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-27-outbound-proxy-policy.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-27-outbound-proxy-policy.md
-2026-08-27-outbound-proxy-policy.md: 9db894c45bd1c816e7394ac621e4be18a8a94e74
-2026-08-27-outbound-proxy-policy.zh.md: 30e30e725ade9d6daf94129fa45b81a6aeb97f96
+2026-08-27-outbound-proxy-policy.md: ef6ad24cead3b0c2c2f4ac5fff9ec1dc9b602a5f
+2026-08-27-outbound-proxy-policy.zh.md: d824fc705723b298b1ffc05b6d7996219d01459e

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-27-outbound-proxy-policy.md

@@ -8,7 +8,7 @@ English | [中文](2026-08-27-outbound-proxy-policy.zh.md)
 
 Node's built-in `fetch` ignores `HTTP_PROXY` and `HTTPS_PROXY`. Every other tool a developer runs — curl, git, npm, pip — honours them, so a user behind a proxy exports the variables once and expects everything to follow. The harness did not: `setGlobalDispatcher`, `ProxyAgent`, and `EnvHttpProxyAgent` appeared zero times across `packages/` and `apps/`, so the model request, every web search, `web_fetch`, MCP over HTTP, and the OTLP exporter all connected directly, silently, with no diagnostic anywhere.
 
-The repository had briefly had an answer and lost it without noticing. PR #971 set `NODE_USE_ENV_PROXY=1` in `bin/dsh`; eleven days later `bbb1b1cc38 cleanup: remove managed source installer` deleted that launcher wholesale, taking the flag with it. What survived was one sentence in `apps/cli/reference/README.md` telling the reader to set a variable that nothing consumed any more.
+The repository had briefly had an answer and lost it without noticing. PR #971 set `NODE_USE_ENV_PROXY=1` in `bin/dsh`; eleven days later the “cleanup: remove managed source installer” change deleted that launcher wholesale, taking the flag with it. What survived was one sentence in `apps/cli/reference/README.md` telling the reader to set a variable that nothing consumed any more.
 
 That sentence could not have worked anyway, for three measured reasons. `NODE_USE_ENV_PROXY` samples the environment at process start, while `loadLayeredEnv()` merges the `.env` layers afterwards, so a proxy declared in `$DSH_HOME/.env` is invisible to it. It reaches Node 24.0+ and, on the 22 line, only 22.21+ — while `engines` admits `^22.19.0`, where the variable does not exist and setting it warns about nothing. And it does not reach `web-fetch-http` at all: that provider passes its own `dispatcher` to `fetch`, and an explicit dispatcher overrides the global one whatever the flag says.
 

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-27-outbound-proxy-policy.zh.md

@@ -8,7 +8,7 @@ Status: implemented
 
 Node 内置的 `fetch` 会忽略 `HTTP_PROXY` 与 `HTTPS_PROXY`。开发者运行的其他工具——curl、git、npm、pip——都遵循它们,所以代理后面的用户导出一次变量就期待一切随之生效。Harness 并没有:`setGlobalDispatcher`、`ProxyAgent` 与 `EnvHttpProxyAgent` 在 `packages/` 与 `apps/` 中出现次数为零,因此模型请求、每次 web 搜索、`web_fetch`、走 HTTP 的 MCP 与 OTLP 导出器全部直连,且是静默的,任何地方都没有诊断。
 
-仓库曾短暂拥有过答案,又在无人察觉时弄丢了。PR #971 在 `bin/dsh` 里设置了 `NODE_USE_ENV_PROXY=1`;十一天后 `bbb1b1cc38 cleanup: remove managed source installer` 整体删除了那个启动器,把该标志一并带走。留下的只有 `apps/cli/reference/README.md` 里的一句话,让读者去设置一个已经无人消费的变量。
+仓库曾短暂拥有过答案,又在无人察觉时弄丢了。PR #971 在 `bin/dsh` 里设置了 `NODE_USE_ENV_PROXY=1`;十一天后的“cleanup: remove managed source installer”改动整体删除了那个启动器,把该标志一并带走。留下的只有 `apps/cli/reference/README.md` 里的一句话,让读者去设置一个已经无人消费的变量。
 
 即便照做,那句话也不可能生效,原因有三条且都经过实测。`NODE_USE_ENV_PROXY` 在进程启动时对环境取快照,而 `loadLayeredEnv()` 是在之后才合并 `.env` 层,因此写在 `$DSH_HOME/.env` 中的代理对它不可见。它只覆盖 Node 24.0+,在 22 线上只覆盖 22.21+——而 `engines` 允许 `^22.19.0`,那里根本没有这个变量,设置了也不会有任何警告。它也完全触及不到 `web-fetch-http`:该提供方向 `fetch` 传入自己的 `dispatcher`,而显式 dispatcher 无论标志如何都会覆盖全局的那个。
 

+ 2 - 2
.agents/notes/implemented/architecture/2026-08-30-retain-ignorable-external-session-events.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/architecture/2026-08-30-retain-ignorable-external-session-events.md
-2026-08-30-retain-ignorable-external-session-events.md: 1fe3a6d99a16daa6ad88f6717baa18f18e8c7355
-2026-08-30-retain-ignorable-external-session-events.zh.md: 632b7b418252c2b299162f3e00a4d41169169509
+2026-08-30-retain-ignorable-external-session-events.md: e085e429a6ad49cf742694cec7dda05770a5b8a1
+2026-08-30-retain-ignorable-external-session-events.zh.md: 1fe4f81441503287961b993eae444fa4385f4803

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-30-retain-ignorable-external-session-events.md

@@ -6,7 +6,7 @@ English | [中文](2026-08-30-retain-ignorable-external-session-events.zh.md)
 
 ## Problem
 
-The session event envelope carries `ignorable?: true` so a reader can accept an unrecognized informational event without treating every vocabulary addition as a new session format. [PR #3087](https://github.com/deepseek-harness/deepseek-harness/pull/3087) removed the field after finding no first-party producer and made every unknown event required-on-read.
+The session event envelope carries `ignorable?: true` so a reader can accept an unrecognized informational event without treating every vocabulary addition as a new session format. PR #3087 removed the field after finding no first-party producer and made every unknown event required-on-read.
 
 That producer inventory did not cover a third-party plugin that currently depends on the field. Without `ignorable`, a first-party reader rejects a stored session containing the plugin's informational event because the event is outside the repository-generated `KNOWN_SESSION_EVENT_TYPES`. The plugin has no replacement registration or versioning mechanism, so deleting the field before a replacement exists breaks a current external consumer.
 

+ 1 - 1
.agents/notes/implemented/architecture/2026-08-30-retain-ignorable-external-session-events.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-会话事件信封包含 `ignorable?: true`,读取器因此可以接受不认识的信息性事件,而不必把每次词汇增加都视为新的会话格式。[PR #3087](https://github.com/deepseek-harness/deepseek-harness/pull/3087) 在没有发现第一方生产方后删除了该字段,并把每个未知事件都改为读取必需项。
+会话事件信封包含 `ignorable?: true`,读取器因此可以接受不认识的信息性事件,而不必把每次词汇增加都视为新的会话格式。PR #3087 在没有发现第一方生产方后删除了该字段,并把每个未知事件都改为读取必需项。
 
 该生产方清单没有覆盖当前依赖此字段的一个第三方插件。没有 `ignorable` 时,第一方读取器会拒绝包含该插件信息性事件的已存会话,因为该事件不在仓库生成的 `KNOWN_SESSION_EVENT_TYPES` 中。插件没有可替代的注册或版本机制,因此在替代机制存在前删除该字段会破坏当前外部消费方。
 

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-08-17-subagent-message-settlement-ordering.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-08-17-subagent-message-settlement-ordering.md
-2026-08-17-subagent-message-settlement-ordering.md: 7462a670766664745e46204dcb01e578b4f86219
-2026-08-17-subagent-message-settlement-ordering.zh.md: bba8fe0e7c543982176def4dbcc6f94dda204339
+2026-08-17-subagent-message-settlement-ordering.md: b383dc290949a3a22ed48e1e639b145e39dbb211
+2026-08-17-subagent-message-settlement-ordering.zh.md: 330c19b20b966d9f719fb61ee1bf5272e0cc8287

+ 1 - 1
.agents/notes/implemented/bug-fix/2026-08-17-subagent-message-settlement-ordering.md

@@ -6,7 +6,7 @@ English | [中文](2026-08-17-subagent-message-settlement-ordering.zh.md)
 
 ## Problem
 
-A continuable child can send selected content and later produce an unconditional manager-authored settlement notice. If those two messages enter queues with different claim priority, the later settlement notice can reach the parent model before the earlier child message. The first step of a turn claims the complete `next-step` batch before one `next-turn` message, so mixing a FIFO later-turn send with a next-step settlement reverses causal order. [Issue #2600](https://github.com/deepseek-harness/deepseek-harness/issues/2600) records the defect.
+A continuable child can send selected content and later produce an unconditional manager-authored settlement notice. If those two messages enter queues with different claim priority, the later settlement notice can reach the parent model before the earlier child message. The first step of a turn claims the complete `next-step` batch before one `next-turn` message, so mixing a FIFO later-turn send with a next-step settlement reverses causal order. Issue #2600 records the defect.
 
 The child instruction says to send a finding whenever it changes what the parent should do next. Deferring that message to a later turn contradicts its scheduling meaning and separates causally ordered messages across queues with different claim priority.
 

+ 1 - 1
.agents/notes/implemented/bug-fix/2026-08-17-subagent-message-settlement-ordering.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-可继续 child 可以发送选中内容,之后还会产生一条由管理器编写且无条件投递的结算通知。如果这两条消息进入领取优先级不同的队列,较晚的结算通知可能先于较早的 child 消息到达 parent 模型。一个轮次的第一个 step 会先领取完整 `next-step` 批次,再领取一条 `next-turn` 消息,因此混用 FIFO 后续轮次发送与 next-step 结算会颠倒因果顺序。[Issue #2600](https://github.com/deepseek-harness/deepseek-harness/issues/2600)记录了该缺陷。
+可继续 child 可以发送选中内容,之后还会产生一条由管理器编写且无条件投递的结算通知。如果这两条消息进入领取优先级不同的队列,较晚的结算通知可能先于较早的 child 消息到达 parent 模型。一个轮次的第一个 step 会先领取完整 `next-step` 批次,再领取一条 `next-turn` 消息,因此混用 FIFO 后续轮次发送与 next-step 结算会颠倒因果顺序。Issue #2600记录了该缺陷。
 
 child 指令要求在发现会改变 parent 下一步动作时发送该发现。把这条消息推迟到后续轮次既违背其调度含义,也会把具有因果顺序的消息拆到领取优先级不同的队列。
 

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md
-2026-09-06-windows-python-console-spawn-wait.md: 92443bcf8a6e4e5609dc469efa4ebd1d82ab127f
-2026-09-06-windows-python-console-spawn-wait.zh.md: dba2f324b29955580fc11e7cea7a0525a8bc8c86
+2026-09-06-windows-python-console-spawn-wait.md: 036062bdf459e4cbb69a6c31d3160c78ec1fad4c
+2026-09-06-windows-python-console-spawn-wait.zh.md: d0870be84066b885e1622438a9fbc277f56a6bcf

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-06-windows-python-console-spawn-wait.zh.md)
 
 ## Problem
 
-The installed Python `dsh.exe` console command intermittently exits with Windows access violation `0xc0000005` before initializing a profile. Its smoke assertion omitted the process status and reported only empty streams. A [native faulthandler probe](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34030851888) captures the fault in Python 3.10 `os._execvpe`, called by the runtime console entry, rather than in the bundled Node executable. Direct executable controls pass.
+The installed Python `dsh.exe` console command intermittently exits with Windows access violation `0xc0000005` before initializing a profile. Its smoke assertion omitted the process status and reported only empty streams. A native faulthandler probe (run 34030851888) captures the fault in Python 3.10 `os._execvpe`, called by the runtime console entry, rather than in the bundled Node executable. Direct executable controls pass.
 
 ## Decision
 
@@ -24,4 +24,4 @@ The [installed-wheel smoke](../../../../scripts/smoke-python-runtime.py) reports
 
 Windows keeps a Python parent until the runtime exits; it no longer depends on CRT overlay behavior. The standard synchronous subprocess implementation owns waiting and interruption cleanup. No custom process-tree manager or global host setting is added.
 
-[Runtime-resolution tests](../../../../python/sdk/tests/test_runtime_resolution.py) retain POSIX forwarding and cover Windows argument/environment forwarding, statuses 0/37/513, real child completion, Unicode streams and arguments with spaces. Native Windows owns the wide exit-status case because POSIX truncates process statuses to eight bits. The [native fixed-count comparison](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34031142773) passes all four patched launches with compile caching enabled; all four unpatched controls also pass in that batch, so it is not a same-batch reproduction. Full installed-wheel CI must validate the final artifact separately from local branch-level tests.
+[Runtime-resolution tests](../../../../python/sdk/tests/test_runtime_resolution.py) retain POSIX forwarding and cover Windows argument/environment forwarding, statuses 0/37/513, real child completion, Unicode streams and arguments with spaces. Native Windows owns the wide exit-status case because POSIX truncates process statuses to eight bits. The native fixed-count comparison (run 34031142773) passes all four patched launches with compile caching enabled; all four unpatched controls also pass in that batch, so it is not a same-batch reproduction. Full installed-wheel CI must validate the final artifact separately from local branch-level tests.

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-Python 安装的 `dsh.exe` 控制台命令会在初始化 profile 前间歇性地以 Windows 访问冲突 `0xc0000005` 退出。其冒烟断言遗漏进程状态,只报告空标准流。[原生 faulthandler 探测](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34030851888) 将故障定位在运行时控制台入口调用的 Python 3.10 `os._execvpe`,而非打包的 Node 可执行文件。直接启动可执行文件的对照通过。
+Python 安装的 `dsh.exe` 控制台命令会在初始化 profile 前间歇性地以 Windows 访问冲突 `0xc0000005` 退出。其冒烟断言遗漏进程状态,只报告空标准流。原生 faulthandler 探测 (run 34030851888) 将故障定位在运行时控制台入口调用的 Python 3.10 `os._execvpe`,而非打包的 Node 可执行文件。直接启动可执行文件的对照通过。
 
 ## 决策
 
@@ -24,4 +24,4 @@ Python 安装的 `dsh.exe` 控制台命令会在初始化 profile 前间歇性
 
 Windows 保留 Python 父进程直到运行时退出,不再依赖 CRT overlay 行为。标准同步子进程实现负责等待和中断清理。不添加自定义进程树管理器或全局主机设置。
 
-[运行时解析测试](../../../../python/sdk/tests/test_runtime_resolution.py) 保留 POSIX 转发验证,并覆盖 Windows 参数/环境转发、状态 0/37/513、真实子进程完成、Unicode 标准流和带空格的参数。宽退出状态由原生 Windows 验证,因为 POSIX 会将进程状态截断为八位。[原生固定次数对照](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34031142773) 中,启用编译缓存的四次修复后启动全部通过;该批次四次未修复对照也全部通过,因此它不是同批次复现。完整安装后 wheel CI 必须独立于本地分支级测试,验证最终产物。
+[运行时解析测试](../../../../python/sdk/tests/test_runtime_resolution.py) 保留 POSIX 转发验证,并覆盖 Windows 参数/环境转发、状态 0/37/513、真实子进程完成、Unicode 标准流和带空格的参数。宽退出状态由原生 Windows 验证,因为 POSIX 会将进程状态截断为八位。原生固定次数对照 (run 34031142773) 中,启用编译缓存的四次修复后启动全部通过;该批次四次未修复对照也全部通过,因此它不是同批次复现。完整安装后 wheel CI 必须独立于本地分支级测试,验证最终产物。

+ 2 - 2
.agents/notes/implemented/bug-fix/2026-09-07-typert-package-local-forwarding-imports.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-07-typert-package-local-forwarding-imports.md
-2026-09-07-typert-package-local-forwarding-imports.md: f50dc7bfc8c9d83c2b6f2b584e1d1119b8df817b
-2026-09-07-typert-package-local-forwarding-imports.zh.md: 7e012d220df3e7356b35a105784f89e4df148802
+2026-09-07-typert-package-local-forwarding-imports.md: 33e4334edef732b9c498ad2397151bd9c2c16f1f
+2026-09-07-typert-package-local-forwarding-imports.zh.md: ea4babdb4b58981fb16f499c7986279fa18dbc58

+ 1 - 1
.agents/notes/implemented/bug-fix/2026-09-07-typert-package-local-forwarding-imports.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-07-typert-package-local-forwarding-imports.zh.md)
 
 ## Problem
 
-`WorkspaceAnalyzer` resolves every type reference to its original declaration before classifying it, then reads only the referencing file's own `import` statement to decide whether the reference crossed a package through a public export. A package that re-exports another package's type from one of its own modules, and imports that module by relative path elsewhere, therefore fails with `crosses a package without an explicit package import` although the package import exists one hop away. The failure is deterministic for every batch size and package order; it surfaces in whichever analysis selects the referencing package as a root, which is why [issue 3525](https://github.com/deepseek-harness/deepseek-harness/issues/3525) observed it as batch-dependent.
+`WorkspaceAnalyzer` resolves every type reference to its original declaration before classifying it, then reads only the referencing file's own `import` statement to decide whether the reference crossed a package through a public export. A package that re-exports another package's type from one of its own modules, and imports that module by relative path elsewhere, therefore fails with `crosses a package without an explicit package import` although the package import exists one hop away. The failure is deterministic for every batch size and package order; it surfaces in whichever analysis selects the referencing package as a root, which is why issue 3525 observed it as batch-dependent.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/bug-fix/2026-09-07-typert-package-local-forwarding-imports.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## Problem
 
-`WorkspaceAnalyzer` 先把每个类型引用解析到原始声明再分类,然后只读引用所在文件自己的 `import` 语句来判断该引用是否经由公开导出跨包。一个包若在自己的某个模块里重新导出另一个包的类型,并在别处用相对路径导入该模块,就会报 `crosses a package without an explicit package import`,尽管包导入只隔一跳。这个失败在任何批次大小和包顺序下都会稳定出现;它出现在哪次分析里,取决于哪次分析把引用方的包选为根,因此 [issue 3525](https://github.com/deepseek-harness/deepseek-harness/issues/3525) 观察到的现象像是与批次相关。
+`WorkspaceAnalyzer` 先把每个类型引用解析到原始声明再分类,然后只读引用所在文件自己的 `import` 语句来判断该引用是否经由公开导出跨包。一个包若在自己的某个模块里重新导出另一个包的类型,并在别处用相对路径导入该模块,就会报 `crosses a package without an explicit package import`,尽管包导入只隔一跳。这个失败在任何批次大小和包顺序下都会稳定出现;它出现在哪次分析里,取决于哪次分析把引用方的包选为根,因此 issue 3525 观察到的现象像是与批次相关。
 
 ## Decision
 

+ 2 - 2
.agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.md
-2026-09-06-evidence-driven-performance-skill.md: 5b15cce1adbd7ff47e5668f7332cba8d1b59e5fe
-2026-09-06-evidence-driven-performance-skill.zh.md: c1fbbd76740badd87ed0a95218e2c4082f2c0e8b
+2026-09-06-evidence-driven-performance-skill.md: 221287ef03386f53d61458db7a280142dfde690a
+2026-09-06-evidence-driven-performance-skill.zh.md: 90e2c1aa5778d23492bec4f2049ce158433e2363

+ 12 - 12
.agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.md

@@ -22,18 +22,18 @@ These are author-reported historical measurements, not benchmarks rerun for this
 
 | Evidence | Measured path and result | Reusable lesson |
 |---|---|---|
-| [#3535](https://github.com/deepseek-harness/deepseek-harness/pull/3535), merged | The [final benchmark design](https://github.com/deepseek-harness/deepseek-harness/pull/3535#issuecomment-5552779119) reports a 4,394 ms first-open negative control against 550 ms, first-history 4,452 against 550, resume 4,333 against 450, and 128 MB heap failures. Client fold: 123.9 ms / 10.84× against 40 ms / 3.125×. | Built-JS user-path gates and positive/negative controls matter more than an earlier PR-body design. |
-| [#3536](https://github.com/deepseek-harness/deepseek-harness/pull/3536), closed unmerged | Repeated snapshot/freeze work occupied about 70% of profiled CPU; synthetic open improved from 4,734–4,921 to 707–823 ms. | Streaming migration superseded this identity-registry proposal. Do not revive it without current ownership evidence. |
-| [#3585](https://github.com/deepseek-harness/deepseek-harness/pull/3585), merged | Historical physical decode: 7.527 s / 7,219 MB peak RSS to 1.467 s / 908 MB; streaming migration with serial publication: 6.241 s, 2.107 GB peak, 477 MB retained. Settled 500,000-delta Client fold: 3.2 ms. | Keep representations compact across consumers; bound intermediate state. Attribution estimates overlap and cannot be added. |
-| [#3586](https://github.com/deepseek-harness/deepseek-harness/pull/3586), merged | Current-v2 opening snapshot: 2,011.4→1,027.9 ms; restore: 598.5→16 ms; retained heap: 1,025.3→478.7 MB. | Separate read-only preparation from awaited write publication; share immutable ownership with revision-keyed preparation and caller-local cancellation. |
-| [#3537](https://github.com/deepseek-harness/deepseek-harness/pull/3537), merged | Synthetic 200-turn projection: 28→5.4 ms; total: 76.9→50 ms; peak RSS: 137.2→94.9 MB. | Read stats, usage, text and image references per compact record. Expanded-stream caching retains unnecessary representation cost. Chat/Trajectory belong to the preceding migration change. |
-| [#2587](https://github.com/deepseek-harness/deepseek-harness/pull/2587), merged | Historical 416,756 events represented by 696 records: client history 4,682→276 ms; sampled additional V8 peak 612.5→199.4 MB. | Preserve compactness through validation and folding; [baseline review](https://github.com/deepseek-harness/deepseek-harness/pull/2587#discussion_r3803082730) requires equal validation and retained output, not parse-and-discard. |
-| [#3331](https://github.com/deepseek-harness/deepseek-harness/pull/3331), merged | 10,000 collapsed tool rows: 22.5→7.5 ms, retained 12.2→1.6 MiB; inactive Trajectory flushes: 4,082→15.5 ms. | Defer unused parsing and materialization; first activation and retained Context still cost work. |
-| [#3391](https://github.com/deepseek-harness/deepseek-harness/pull/3391) and [#3383](https://github.com/deepseek-harness/deepseek-harness/pull/3383), merged | Narrow subscriptions, stable identities, batched publication, and viewport-triggered highlighting. The 10,000-node timing table is estimated, not browser measurement. | Deferral is not virtualization: visited token DOM remains retained. |
-| [#3292](https://github.com/deepseek-harness/deepseek-harness/pull/3292), merged | Two-million-item FIFO drain: 9.656 ms median, excluding enqueue. | A deque removes shift copying, not queue admission or backpressure obligations. |
-| [#1161](https://github.com/deepseek-harness/deepseek-harness/pull/1161), merged | Keyless 100,000-chunk browser stress at 128 chunks per 16 ms. | [Producer catch-up](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970161) and [final heartbeat stalls](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970162) can distort measurements; scheduled events are not trusted keyboard/pointer input. |
-
-The [cancellation review](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940578092), [source-revision review](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940569241), and [typed-reader review](https://github.com/deepseek-harness/deepseek-harness/pull/3537#discussion_r3942974015) illustrate why removing repeated work does not authorize deleting validation or publication obligations. A [standby-runner review](https://github.com/deepseek-harness/deepseek-harness/pull/3535#discussion_r3927945561) distinguishes a dedicated job from an isolated physical host.
+| #3535, merged | The final benchmark design (PR #3535, comment 5552779119) reports a 4,394 ms first-open negative control against 550 ms, first-history 4,452 against 550, resume 4,333 against 450, and 128 MB heap failures. Client fold: 123.9 ms / 10.84× against 40 ms / 3.125×. | Built-JS user-path gates and positive/negative controls matter more than an earlier PR-body design. |
+| #3536, closed unmerged | Repeated snapshot/freeze work occupied about 70% of profiled CPU; synthetic open improved from 4,734–4,921 to 707–823 ms. | Streaming migration superseded this identity-registry proposal. Do not revive it without current ownership evidence. |
+| #3585, merged | Historical physical decode: 7.527 s / 7,219 MB peak RSS to 1.467 s / 908 MB; streaming migration with serial publication: 6.241 s, 2.107 GB peak, 477 MB retained. Settled 500,000-delta Client fold: 3.2 ms. | Keep representations compact across consumers; bound intermediate state. Attribution estimates overlap and cannot be added. |
+| #3586, merged | Current-v2 opening snapshot: 2,011.4→1,027.9 ms; restore: 598.5→16 ms; retained heap: 1,025.3→478.7 MB. | Separate read-only preparation from awaited write publication; share immutable ownership with revision-keyed preparation and caller-local cancellation. |
+| #3537, merged | Synthetic 200-turn projection: 28→5.4 ms; total: 76.9→50 ms; peak RSS: 137.2→94.9 MB. | Read stats, usage, text and image references per compact record. Expanded-stream caching retains unnecessary representation cost. Chat/Trajectory belong to the preceding migration change. |
+| #2587, merged | Historical 416,756 events represented by 696 records: client history 4,682→276 ms; sampled additional V8 peak 612.5→199.4 MB. | Preserve compactness through validation and folding; baseline review (PR #2587, comment 3803082730) requires equal validation and retained output, not parse-and-discard. |
+| #3331, merged | 10,000 collapsed tool rows: 22.5→7.5 ms, retained 12.2→1.6 MiB; inactive Trajectory flushes: 4,082→15.5 ms. | Defer unused parsing and materialization; first activation and retained Context still cost work. |
+| #3391 and #3383, merged | Narrow subscriptions, stable identities, batched publication, and viewport-triggered highlighting. The 10,000-node timing table is estimated, not browser measurement. | Deferral is not virtualization: visited token DOM remains retained. |
+| #3292, merged | Two-million-item FIFO drain: 9.656 ms median, excluding enqueue. | A deque removes shift copying, not queue admission or backpressure obligations. |
+| #1161, merged | Keyless 100,000-chunk browser stress at 128 chunks per 16 ms. | Producer catch-up (PR #1161, comment 3699970161) and final heartbeat stalls (PR #1161, comment 3699970162) can distort measurements; scheduled events are not trusted keyboard/pointer input. |
+
+The cancellation review (PR #3586, comment 3940578092), source-revision review (PR #3586, comment 3940569241), and typed-reader review (PR #3537, comment 3942974015) illustrate why removing repeated work does not authorize deleting validation or publication obligations. A standby-runner review (PR #3535, comment 3927945561) distinguishes a dedicated job from an isolated physical host.
 
 ## Alternatives considered
 

+ 12 - 12
.agents/notes/implemented/process/2026-09-06-evidence-driven-performance-skill.zh.md

@@ -22,18 +22,18 @@ Status: implemented
 
 | 证据 | 测量路径与结果 | 可复用经验 |
 |---|---|---|
-| [#3535](https://github.com/deepseek-harness/deepseek-harness/pull/3535),已合并 | [最终基准设计](https://github.com/deepseek-harness/deepseek-harness/pull/3535#issuecomment-5552779119)报告首次打开负向对照 4,394 ms,预算 550 ms;首屏历史 4,452,预算 550;恢复 4,333,预算 450;128 MB 堆检查失败。Client fold:123.9 ms / 10.84×,预算 40 ms / 3.125×。 | built-JS 用户路径门禁与正/负向对照比早期 PR 正文设计更重要。 |
-| [#3536](https://github.com/deepseek-harness/deepseek-harness/pull/3536),关闭未合并 | 重复 snapshot/freeze 工作占采样 CPU 的约 70%;合成打开从 4,734–4,921 改善为 707–823 ms。 | 流式迁移替代了该身份注册表提案。没有当前所有权证据时,不恢复它。 |
-| [#3585](https://github.com/deepseek-harness/deepseek-harness/pull/3585),已合并 | 历史物理解码:7.527 s / 7,219 MB 峰值 RSS 降至 1.467 s / 908 MB;流式迁移加串行发布:6.241 s,2.107 GB 峰值,477 MB 保留。已结算的 500,000-delta Client fold:3.2 ms。 | 跨消费者保持紧凑表示;限制中间状态。归因估计重叠,不能相加。 |
-| [#3586](https://github.com/deepseek-harness/deepseek-harness/pull/3586),已合并 | 当前 v2 打开快照:2,011.4→1,027.9 ms;恢复:598.5→16 ms;保留堆:1,025.3→478.7 MB。 | 分离只读准备与必须等待的写发布;通过按修订号共享准备和调用方局部取消共享不可变所有权。 |
-| [#3537](https://github.com/deepseek-harness/deepseek-harness/pull/3537),已合并 | 合成 200 轮投影:28→5.4 ms;总计:76.9→50 ms;峰值 RSS:137.2→94.9 MB。 | 按紧凑记录读取统计、usage、文本和图像引用。展开流缓存保留不必要的表示成本。Chat/Trajectory 属于前置迁移改动。 |
-| [#2587](https://github.com/deepseek-harness/deepseek-harness/pull/2587),已合并 | 历史 416,756 事件由 696 记录表示:Client 历史 4,682→276 ms;采样额外 V8 峰值 612.5→199.4 MB。 | 验证和折叠过程保持紧凑;[基线审查](https://github.com/deepseek-harness/deepseek-harness/pull/2587#discussion_r3803082730)要求相同验证与保留输出,而不是解析后丢弃。 |
-| [#3331](https://github.com/deepseek-harness/deepseek-harness/pull/3331),已合并 | 10,000 个折叠工具行:22.5→7.5 ms,保留 12.2→1.6 MiB;非活动 Trajectory 刷新:4,082→15.5 ms。 | 延迟未使用的解析和实体化;首次激活与保留 Context 仍有成本。 |
-| [#3391](https://github.com/deepseek-harness/deepseek-harness/pull/3391)[#3383](https://github.com/deepseek-harness/deepseek-harness/pull/3383),已合并 | 缩小订阅范围、稳定身份、批量发布和视口触发高亮。10,000 节点计时表是估计,不是浏览器测量。 | 延迟不等于虚拟化:访问过的 token DOM 仍被保留。 |
-| [#3292](https://github.com/deepseek-harness/deepseek-harness/pull/3292),已合并 | 两百万条 FIFO 排空:中位数 9.656 ms,不含入队。 | deque 删除 shift 复制,不删除队列准入或背压义务。 |
-| [#1161](https://github.com/deepseek-harness/deepseek-harness/pull/1161),已合并 | 无密钥的 100,000-chunk 浏览器压力测试,每 16 ms 推送 128 个 chunk。 | [生产者追赶](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970161)和[最后一次心跳停顿](https://github.com/deepseek-harness/deepseek-harness/pull/1161#discussion_r3699970162)可能扭曲测量;定时派发事件不是真实键盘/指针输入。 |
-
-[取消审查](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940578092)、[源修订审查](https://github.com/deepseek-harness/deepseek-harness/pull/3586#discussion_r3940569241)和[类型化读取器审查](https://github.com/deepseek-harness/deepseek-harness/pull/3537#discussion_r3942974015)说明删除重复工作不等于允许删除验证或发布义务。[备用 runner 审查](https://github.com/deepseek-harness/deepseek-harness/pull/3535#discussion_r3927945561)区分独立 job 与隔离的物理主机。
+| #3535,已合并 | 最终基准设计 (PR #3535, comment 5552779119)报告首次打开负向对照 4,394 ms,预算 550 ms;首屏历史 4,452,预算 550;恢复 4,333,预算 450;128 MB 堆检查失败。Client fold:123.9 ms / 10.84×,预算 40 ms / 3.125×。 | built-JS 用户路径门禁与正/负向对照比早期 PR 正文设计更重要。 |
+| #3536,关闭未合并 | 重复 snapshot/freeze 工作占采样 CPU 的约 70%;合成打开从 4,734–4,921 改善为 707–823 ms。 | 流式迁移替代了该身份注册表提案。没有当前所有权证据时,不恢复它。 |
+| #3585,已合并 | 历史物理解码:7.527 s / 7,219 MB 峰值 RSS 降至 1.467 s / 908 MB;流式迁移加串行发布:6.241 s,2.107 GB 峰值,477 MB 保留。已结算的 500,000-delta Client fold:3.2 ms。 | 跨消费者保持紧凑表示;限制中间状态。归因估计重叠,不能相加。 |
+| #3586,已合并 | 当前 v2 打开快照:2,011.4→1,027.9 ms;恢复:598.5→16 ms;保留堆:1,025.3→478.7 MB。 | 分离只读准备与必须等待的写发布;通过按修订号共享准备和调用方局部取消共享不可变所有权。 |
+| #3537,已合并 | 合成 200 轮投影:28→5.4 ms;总计:76.9→50 ms;峰值 RSS:137.2→94.9 MB。 | 按紧凑记录读取统计、usage、文本和图像引用。展开流缓存保留不必要的表示成本。Chat/Trajectory 属于前置迁移改动。 |
+| #2587,已合并 | 历史 416,756 事件由 696 记录表示:Client 历史 4,682→276 ms;采样额外 V8 峰值 612.5→199.4 MB。 | 验证和折叠过程保持紧凑;基线审查 (PR #2587, comment 3803082730)要求相同验证与保留输出,而不是解析后丢弃。 |
+| #3331,已合并 | 10,000 个折叠工具行:22.5→7.5 ms,保留 12.2→1.6 MiB;非活动 Trajectory 刷新:4,082→15.5 ms。 | 延迟未使用的解析和实体化;首次激活与保留 Context 仍有成本。 |
+| #3391 和 #3383,已合并 | 缩小订阅范围、稳定身份、批量发布和视口触发高亮。10,000 节点计时表是估计,不是浏览器测量。 | 延迟不等于虚拟化:访问过的 token DOM 仍被保留。 |
+| #3292,已合并 | 两百万条 FIFO 排空:中位数 9.656 ms,不含入队。 | deque 删除 shift 复制,不删除队列准入或背压义务。 |
+| #1161,已合并 | 无密钥的 100,000-chunk 浏览器压力测试,每 16 ms 推送 128 个 chunk。 | 生产者追赶 (PR #1161, comment 3699970161)和最后一次心跳停顿 (PR #1161, comment 3699970162)可能扭曲测量;定时派发事件不是真实键盘/指针输入。 |
+
+取消审查 (PR #3586, comment 3940578092)、源修订审查 (PR #3586, comment 3940569241)和类型化读取器审查 (PR #3537, comment 3942974015)说明删除重复工作不等于允许删除验证或发布义务。备用 runner 审查 (PR #3535, comment 3927945561)区分独立 job 与隔离的物理主机。
 
 ## 考虑过的替代方案
 

+ 2 - 2
.agents/notes/implemented/process/2026-09-06-node-compatibility-selfhosted.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-node-compatibility-selfhosted.md
-2026-09-06-node-compatibility-selfhosted.md: c361d21d3e1093dd5c87bf2ba085bdd1acacb5d8
-2026-09-06-node-compatibility-selfhosted.zh.md: 80a9d8519d6084f5e01944b277882a51b2291e94
+2026-09-06-node-compatibility-selfhosted.md: 046b9aab7a35b2398e7ed99262bb2208d0857a5f
+2026-09-06-node-compatibility-selfhosted.zh.md: 1a37f3a14306de2551dbd484772b71146f02624d

+ 1 - 1
.agents/notes/implemented/process/2026-09-06-node-compatibility-selfhosted.md

@@ -32,4 +32,4 @@ The pool receives three additional jobs per trusted PR; each retains gate concur
 
 The focused [workflow regression](../../../../scripts/ci-compatible-selfhosted.spec.ts) executes the actual routing expressions and environment setup. A negative control removing the fork condition fails the hosted-fallback assertion. It checks Dependabot reruns by a maintainer, repository mismatch, fork flags, disabled variables, and runner-scoped cache paths.
 
-[Successful standby run 33984559660](https://github.com/deepseek-harness/deepseek-harness/actions/runs/33984559660) at the implementation base supplies Linux Node 24.19.0 and Windows Node 24.20.0 baseline evidence. Linux job 101359402557 uses runner-specific temporary and tool directories on the data volume. [Read-only capability probe 34012679056](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34012679056/job/101431064925) reports Linux x64, 192 online logical CPUs, GCC/G++ 13.3, Make 4.3, and Python 3.12.3. Python 3.10 is absent, reinforcing the separate SDK provisioning requirement. [PR run 34013779750](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34013779750) at `282519d2` verifies Node 22.19.0, 24.9.0, and 26.8.1 on self-hosted Linux, including setup, executable-path checks, compatibility tests, and post actions. The executables reside under each runner’s `_temp/node-compat-toolcache/node/<version>/x64/bin`; the completed jobs take 228s, 94s, and 101s respectively. These observations establish version and path compatibility, not an exclusive-host capacity guarantee.
+Successful standby run 33984559660 at the implementation base supplies Linux Node 24.19.0 and Windows Node 24.20.0 baseline evidence. Linux job 101359402557 uses runner-specific temporary and tool directories on the data volume. Read-only capability probe 34012679056 (job 101431064925) reports Linux x64, 192 online logical CPUs, GCC/G++ 13.3, Make 4.3, and Python 3.12.3. Python 3.10 is absent, reinforcing the separate SDK provisioning requirement. PR run 34013779750 at `282519d2` verifies Node 22.19.0, 24.9.0, and 26.8.1 on self-hosted Linux, including setup, executable-path checks, compatibility tests, and post actions. The executables reside under each runner’s `_temp/node-compat-toolcache/node/<version>/x64/bin`; the completed jobs take 228s, 94s, and 101s respectively. These observations establish version and path compatibility, not an exclusive-host capacity guarantee.

+ 1 - 1
.agents/notes/implemented/process/2026-09-06-node-compatibility-selfhosted.zh.md

@@ -32,4 +32,4 @@ Status: implemented
 
 聚焦的[工作流回归测试](../../../../scripts/ci-compatible-selfhosted.spec.ts) 执行真实的路由表达式和环境设置。移除 fork 条件的负对照使托管回退断言失败。它检查维护者重跑 Dependabot PR、仓库不匹配、fork 标志、禁用变量以及运行器范围内的缓存路径。
 
-实施基线上的[成功热备运行 33984559660](https://github.com/deepseek-harness/deepseek-harness/actions/runs/33984559660) 提供 Linux Node 24.19.0 和 Windows Node 24.20.0 基线证据。Linux 作业 101359402557 使用数据卷上运行器专属的临时目录和工具目录。[只读能力探测 34012679056](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34012679056/job/101431064925) 报告 Linux x64、192 个在线逻辑 CPU、GCC/G++ 13.3、Make 4.3 和 Python 3.12.3。Python 3.10 缺失,进一步说明 SDK 需要单独配置。`282519d2` 上的 [PR 运行 34013779750](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34013779750) 验证了自托管 Linux 上的 Node 22.19.0、24.9.0 和 26.8.1,包括设置、可执行文件路径检查、兼容性测试和 post actions。可执行文件位于各运行器的 `_temp/node-compat-toolcache/node/<version>/x64/bin` 下;完成的作业分别耗时 228s、94s 和 101s。这些观测证明版本与路径兼容性,而非独占主机的容量保证。
+实施基线上的成功热备运行 33984559660 提供 Linux Node 24.19.0 和 Windows Node 24.20.0 基线证据。Linux 作业 101359402557 使用数据卷上运行器专属的临时目录和工具目录。只读能力探测 34012679056 (job 101431064925) 报告 Linux x64、192 个在线逻辑 CPU、GCC/G++ 13.3、Make 4.3 和 Python 3.12.3。Python 3.10 缺失,进一步说明 SDK 需要单独配置。`282519d2` 上的 PR 运行 34013779750 验证了自托管 Linux 上的 Node 22.19.0、24.9.0 和 26.8.1,包括设置、可执行文件路径检查、兼容性测试和 post actions。可执行文件位于各运行器的 `_temp/node-compat-toolcache/node/<version>/x64/bin` 下;完成的作业分别耗时 228s、94s 和 101s。这些观测证明版本与路径兼容性,而非独占主机的容量保证。

+ 2 - 2
.agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.md
-2026-09-06-preview-hosted-runner-sizing.md: 87298e94f11aa7e483afde31e0963a56f523febc
-2026-09-06-preview-hosted-runner-sizing.zh.md: 285b21d60d755e76db582e4e55c5cde913f7fc2e
+2026-09-06-preview-hosted-runner-sizing.md: 414a17bf8a47f901e61ed271e9b2cf3dee3415b8
+2026-09-06-preview-hosted-runner-sizing.zh.md: 91a67b90b6617fea4dfa77672cb7e78412f462ee

+ 2 - 2
.agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.md

@@ -14,7 +14,7 @@ The [preview workflow](../../../../.github/workflows/build-preview-cloudflare.ym
 
 ### Measurements
 
-[Experiment 34012729982](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34012729982) succeeds for all eight size/cache combinations plus one cache seed. Every measured job checks out SHA `9149d7e7ef945b5601711badd3cf63d58ab384f5`, uses Node 24.19.0 and pnpm 11.7.0, and executes immutable install, full workspace build, preview/VFS packing, and local upload shaping with gzip integrity verification. Warm jobs restore one exact run-private pnpm cache; cold jobs skip restoration but contain pnpm bootstrap files. No compiled outputs are restored.
+Experiment 34012729982 succeeds for all eight size/cache combinations plus one cache seed. Every measured job checks out the same experiment revision, uses Node 24.19.0 and pnpm 11.7.0, and executes immutable install, full workspace build, preview/VFS packing, and local upload shaping with gzip integrity verification. Warm jobs restore one exact run-private pnpm cache; cold jobs skip restoration but contain pnpm bootstrap files. No compiled outputs are restored.
 
 | Runner | Cold / warm job seconds | Rounded minutes each | USD each | Workspace seconds cold / warm | Preview seconds cold / warm |
 |---|---:|---:|---:|---:|---:|
@@ -29,7 +29,7 @@ Standard jobs expose two vCPUs and 7.75 GiB RAM. Workspace maximum process RSS i
 
 The comparison fixes source, lockfile, commands, and runtime versions, not physical CPUs or image release: standard and 4-core use image 20260831.293.1; 8-core and 16-core use 20260823.283.1. CPUs vary among AMD EPYC 9V74/7763 and Intel Xeon 8370C/8573C. One sample per cache state measures the offered labels, not isolated CPU scaling or statistical repeatability.
 
-The experiment does not deploy or access Cloudflare credentials. Measurement upload takes zero to one second; warm-cache restore takes six to ten seconds. For context, [production job 101428009994](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34011495156/job/101428009994) spends 14 seconds uploading, one second verifying, and two seconds commenting on a different SHA. Adding that overhead to this experiment is a projection, not a measured standard-runner publication result. The actual PR preview workflow owns deployment confirmation.
+The experiment does not deploy or access Cloudflare credentials. Measurement upload takes zero to one second; warm-cache restore takes six to ten seconds. For context, production job 101428009994 (run 34011495156) spends 14 seconds uploading, one second verifying, and two seconds commenting on a different SHA. Adding that overhead to this experiment is a projection, not a measured standard-runner publication result. The actual PR preview workflow owns deployment confirmation.
 
 ## Alternatives considered
 

+ 2 - 2
.agents/notes/implemented/process/2026-09-06-preview-hosted-runner-sizing.zh.md

@@ -14,7 +14,7 @@ PR(Pull Request)预览构建完整工作区及浏览器 worker VFS 镜像。
 
 ### 测量
 
-[实验 34012729982](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34012729982) 的八种规格/缓存组合及一个缓存预热作业均成功。每个测量作业检出 SHA `9149d7e7ef945b5601711badd3cf63d58ab384f5`,使用 Node 24.19.0 与 pnpm 11.7.0,并执行不可变安装、完整工作区构建、预览/VFS 打包,以及含 gzip 完整性验证的本地上传内容整理。热作业恢复同一个运行私有精确 pnpm 缓存;冷作业跳过恢复,但包含 pnpm 引导安装文件。不恢复编译产物。
+实验 34012729982 的八种规格/缓存组合及一个缓存预热作业均成功。每个测量作业检出同一个实验修订,使用 Node 24.19.0 与 pnpm 11.7.0,并执行不可变安装、完整工作区构建、预览/VFS 打包,以及含 gzip 完整性验证的本地上传内容整理。热作业恢复同一个运行私有精确 pnpm 缓存;冷作业跳过恢复,但包含 pnpm 引导安装文件。不恢复编译产物。
 
 | 运行器 | 冷 / 热作业秒数 | 各自取整分钟数 | 各自美元费用 | 冷 / 热工作区秒数 | 冷 / 热预览秒数 |
 |---|---:|---:|---:|---:|---:|
@@ -29,7 +29,7 @@ PR(Pull Request)预览构建完整工作区及浏览器 worker VFS 镜像。
 
 比较固定源代码、锁文件、命令和运行时版本,但不固定物理 CPU 或镜像版本:标准与 4 核使用镜像 20260831.293.1;8 核与 16 核使用 20260823.283.1。CPU 包括 AMD EPYC 9V74/7763 与 Intel Xeon 8370C/8573C。每种缓存状态的单个样本测量所提供的标签,而非独立 CPU 扩展性或统计可重复性。
 
-实验不部署,也不访问 Cloudflare 凭据。测量上传耗时零至一秒;热缓存恢复耗时六至十秒。作为背景,[生产作业 101428009994](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34011495156/job/101428009994) 在不同 SHA 上上传耗时 14 秒、验证一秒、评论两秒。将该开销加至本实验属于推算,而非已测量的标准运行器发布结果。实际 PR 预览工作流负责部署确认。
+实验不部署,也不访问 Cloudflare 凭据。测量上传耗时零至一秒;热缓存恢复耗时六至十秒。作为背景,生产作业 101428009994 (run 34011495156) 在不同 SHA 上上传耗时 14 秒、验证一秒、评论两秒。将该开销加至本实验属于推算,而非已测量的标准运行器发布结果。实际 PR 预览工作流负责部署确认。
 
 ## 考虑过的替代方案
 

+ 2 - 2
.agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.md
-2026-09-06-python-runtime-windows-hosted.md: 8ae69d9836a07d9760c856d12bea29b4e09d1461
-2026-09-06-python-runtime-windows-hosted.zh.md: ea0b07c8131402efb60e226a6583b1c0e8faf4f0
+2026-09-06-python-runtime-windows-hosted.md: 95b762f2247c791536ae0902287ba40e20066258
+2026-09-06-python-runtime-windows-hosted.zh.md: e3f27788b0b620b4a7215b98bdc42b7d6957d5d7

+ 1 - 1
.agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-06-python-runtime-windows-hosted.zh.md)
 
 ## Problem
 
-The Windows x64 target in [build-exe-for-python-sdk.yml](../../../../.github/workflows/build-exe-for-python-sdk.yml) started resolving through `DSH_CI_FAILOVER_WINDOWS=selfhosted` for trusted pull-request CI when #3629 added the failover selector and the job-private Windows toolchain. The shared `dsh-win-ci` pool did not make the lane more reliable. On 2026-09-06 (all times UTC; every run executed the #3629 migration workflow's selector, which was live from the 07:56 merge) the installed-wheel smoke passed at 09:12 on `dsh-win-ci-16` for [commit `ca3ffe95` of PR #3640 (`ci/benchmark-standard-runner`)](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384), then failed at 10:06 on `dsh-win-ci-21` for [PR #3337 (`feat/visualizer-host-plugin`)](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34026500701) and at 10:46 on `dsh-win-ci-04` for [PR #3640 at its final head `c5ba873f`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34028339888/job/101473395734), where `smoke_sdk_profile_plugin`'s packaged `dsh plugin add` child exited without output while the Linux and macOS cells of that run passed; a job rerun at 11:29 repeated the same silent death. The migration proposal ([#3629](https://github.com/deepseek-harness/deepseek-harness/pull/3629)) remained `proposed` because its throughput and shared-load acceptance criteria were never measured.
+The Windows x64 target in [build-exe-for-python-sdk.yml](../../../../.github/workflows/build-exe-for-python-sdk.yml) started resolving through `DSH_CI_FAILOVER_WINDOWS=selfhosted` for trusted pull-request CI when #3629 added the failover selector and the job-private Windows toolchain. The shared `dsh-win-ci` pool did not make the lane more reliable. On 2026-09-06 (all times UTC; every run executed the #3629 migration workflow's selector, which was live from the 07:56 merge) the installed-wheel smoke passed at 09:12 on `dsh-win-ci-16` for the initial tested revision of PR #3640 (`ci/benchmark-standard-runner`) (run 34023970384), then failed at 10:06 on `dsh-win-ci-21` for PR #3337 (`feat/visualizer-host-plugin`) (run 34026500701) and at 10:46 on `dsh-win-ci-04` for PR #3640 at its final tested revision (run 34028339888, job 101473395734), where `smoke_sdk_profile_plugin`'s packaged `dsh plugin add` child exited without output while the Linux and macOS cells of that run passed; a job rerun at 11:29 repeated the same silent death. The migration proposal (#3629) remained `proposed` because its throughput and shared-load acceptance criteria were never measured.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/process/2026-09-06-python-runtime-windows-hosted.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-当 #3629 加入故障切换选择器与作业私有的 Windows 工具链后,[build-exe-for-python-sdk.yml](../../../../.github/workflows/build-exe-for-python-sdk.yml) 中的 Windows x64 目标开始对受信任的 PR CI 通过 `DSH_CI_FAILOVER_WINDOWS=selfhosted` 解析运行器。共享的 `dsh-win-ci` 池并未让该通道更可靠。2026-09-06(所有时间均为 UTC;每次运行都执行 #3629 迁移工作流的选择器,该选择器自 07:56 合并起生效):安装后 wheel 冒烟测试在 09:12 于 `dsh-win-ci-16` 上为[PR #3640(`ci/benchmark-standard-runner`)的提交 `ca3ffe95`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384)通过,随后 10:06 在 `dsh-win-ci-21` 上为[PR #3337(`feat/visualizer-host-plugin`)](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34026500701)失败,10:46 在 `dsh-win-ci-04` 上为[PR #3640 的最终 head `c5ba873f`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34028339888/job/101473395734)失败——`smoke_sdk_profile_plugin` 打包的 `dsh plugin add` 子进程无输出即退出,而该次运行的 Linux 与 macOS 单元均通过;11:29 的作业重试再次出现相同的无声死亡。迁移提案([#3629](https://github.com/deepseek-harness/deepseek-harness/pull/3629))保持 `proposed`,因为其吞吐量与共享负载验收标准从未实测。
+当 #3629 加入故障切换选择器与作业私有的 Windows 工具链后,[build-exe-for-python-sdk.yml](../../../../.github/workflows/build-exe-for-python-sdk.yml) 中的 Windows x64 目标开始对受信任的 PR CI 通过 `DSH_CI_FAILOVER_WINDOWS=selfhosted` 解析运行器。共享的 `dsh-win-ci` 池并未让该通道更可靠。2026-09-06(所有时间均为 UTC;每次运行都执行 #3629 迁移工作流的选择器,该选择器自 07:56 合并起生效):安装后 wheel 冒烟测试在 09:12 于 `dsh-win-ci-16` 上为PR #3640(`ci/benchmark-standard-runner`)的最初实测修订 (run 34023970384)通过,随后 10:06 在 `dsh-win-ci-21` 上为PR #3337(`feat/visualizer-host-plugin`) (run 34026500701)失败,10:46 在 `dsh-win-ci-04` 上为PR #3640 的最终实测修订 (run 34028339888, job 101473395734)失败——`smoke_sdk_profile_plugin` 打包的 `dsh plugin add` 子进程无输出即退出,而该次运行的 Linux 与 macOS 单元均通过;11:29 的作业重试再次出现相同的无声死亡。迁移提案(#3629)保持 `proposed`,因为其吞吐量与共享负载验收标准从未实测。
 
 ## 决策
 

+ 2 - 2
.agents/notes/implemented/process/2026-09-09-blacksmith-failover-leg.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-09-blacksmith-failover-leg.md
-2026-09-09-blacksmith-failover-leg.md: eaaa6f4d5c8326ae7686dc9f485143781d065313
-2026-09-09-blacksmith-failover-leg.zh.md: 562aed96327a86c24e38b502f2ef4756a11b91e1
+2026-09-09-blacksmith-failover-leg.md: d4be1ad25c6c5b66e295439d4609726c8918e137
+2026-09-09-blacksmith-failover-leg.zh.md: cd6e74aefba2cba9481fa1ae9510bb1ce020325d

Dosya farkı çok büyük olduğundan ihmal edildi
+ 0 - 0
.agents/notes/implemented/process/2026-09-09-blacksmith-failover-leg.md


Dosya farkı çok büyük olduğundan ihmal edildi
+ 0 - 0
.agents/notes/implemented/process/2026-09-09-blacksmith-failover-leg.zh.md


+ 6 - 0
.agents/notes/implemented/process/2026-09-12-maintained-repository-references.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-09-12-maintained-repository-references.md
+2026-09-12-maintained-repository-references.md: c65f8b984c147bd323dd6fbd865402c07c0f4ad8
+2026-09-12-maintained-repository-references.zh.md: e613f332d4d3338e44acd8a1f55271ca044f06eb

+ 31 - 0
.agents/notes/implemented/process/2026-09-12-maintained-repository-references.md

@@ -0,0 +1,31 @@
+# Agent Note: Maintained repository references
+
+Status: implemented
+
+English | [中文](2026-09-12-maintained-repository-references.zh.md)
+
+## Problem
+
+Historical evidence needs recognizable release, PR, and measured-run identities. Maintained files also serve the public source home, while native publishing uses the repository identity supplied by its workflow. Concrete commit references and deployment-specific organization URLs do not express that distinction.
+
+## Decision
+
+Use release tags and PR, run, or job identifiers for historical evidence, and relative links for current repository files. The [reference gate](../../../../scripts/verify-repository-references.ts) scans tracked files and nonignored new files. Vendor sources and frozen Agent Notes retain their existing exclusions; active notes and historical release records remain checked.
+
+The gate resolves hexadecimal candidates against the local Git object database and rejects only unambiguous commit identities. Pairing hashes that identify blobs, schema digests, unrelated hexadecimal values, and hexadecimal branch names that resolve to different object identities remain valid. Organization URL detection shares decoding and normalization with the existing repository-link policy.
+
+Native source manifests identify the public source home. During workflow packing, the [native packer](../../../../native/system/scripts/pack-release.mjs) projects the workflow repository into disposable inputs so npm trusted publishing can verify the artifact identity. It preserves source manifests and publishes the resulting tarballs unchanged. Local packing without workflow context retains the public source metadata.
+
+## Alternatives considered
+
+**Reject every hexadecimal string.** Persistence schemas and bilingual pairing legitimately use hashes. Git object-type verification separates commit references from these values.
+
+**Check only new diff lines.** The policy applies to every maintained file, including existing references. A complete scan also catches staged and untracked additions before publication.
+
+**Delete required publishing metadata.** Native publishing still needs the workflow repository identity. Projecting that identity during packing removes the source literal without changing the identity npm verifies.
+
+## Consequences
+
+The gate performs no network fetch. Local shallow checkouts can identify only objects they contain; CI static checks use full history. Reference validation establishes the maintained-file policy, not remote tag immutability or historical compatibility.
+
+Vendored notices link checked-in source locations while retaining upstream names, licenses, and the unchanged vendor manifest.

+ 31 - 0
.agents/notes/implemented/process/2026-09-12-maintained-repository-references.zh.md

@@ -0,0 +1,31 @@
+# Agent Note: 维护中的仓库引用
+
+Status: implemented
+
+[English](2026-09-12-maintained-repository-references.md) | 中文
+
+## 问题
+
+历史证据需要可辨认的发行版本、PR 和测量运行标识。维护中的文件也面向公开源码主页,而原生发布使用工作流提供的仓库身份。具体提交引用和与部署绑定的组织 URL 无法表达这一区分。
+
+## 决策
+
+历史证据使用发行 tag 以及 PR、运行或 job 标识,当前仓库文件使用相对链接。[引用检查](../../../../scripts/verify-repository-references.ts)扫描已跟踪文件和未被忽略的新文件。Vendor 源码与冻结的 Agent Note 保留既有排除规则;活跃笔记和历史发行记录仍参与检查。
+
+检查器针对本地 Git 对象库解析十六进制候选值,仅拒绝能明确标识提交的值。标识 blob 的配对哈希、schema 摘要、无关十六进制值,以及解析到不同对象标识的十六进制分支名仍然有效。组织 URL 检测与既有仓库链接策略共享解码和规范化逻辑。
+
+原生源码清单标识公开源码主页。工作流打包期间,[原生打包器](../../../../native/system/scripts/pack-release.mjs)将工作流仓库投影到临时输入,让 npm 可信发布可以核验产物身份。它保留源码清单,生成的 tarball 也不在发布时改写。没有工作流上下文的本地打包保留公开源码元数据。
+
+## 考虑过的替代方案
+
+**拒绝每个十六进制字符串。** 持久化 schema 和双语配对合理使用哈希。Git 对象类型校验将提交引用与这些值区分开。
+
+**仅检查新 diff 行。** 策略适用于每个维护中的文件,包括既有引用。完整扫描也会在发布前发现暂存或尚未跟踪的新增文件。
+
+**删除发布所需的元数据。** 原生发布仍需要工作流仓库身份。在打包阶段投影该身份,可以移除源码中的字面量,同时保持 npm 核验的身份不变。
+
+## 影响
+
+检查器不进行网络获取。本地浅检出只能识别其包含的对象;CI 静态检查使用完整历史。引用校验建立维护中文件的策略,不证明远端 tag 不可变或历史兼容性。
+
+Vendored 声明链接仓库内的源码位置,同时保留上游名称、许可证和未变的 vendor 清单。

+ 2 - 2
.agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.md
-2026-09-05-base-default-file-editor.md: b87d7986e7beeab30ff5414908239ba95022d231
-2026-09-05-base-default-file-editor.zh.md: db3c70d0a28527391073ddc34a72e9faabe1c3be
+2026-09-05-base-default-file-editor.md: e1e15718962322ba8c79107fda31446448b87056
+2026-09-05-base-default-file-editor.zh.md: d43df3313ecaeff8b2123208620779df2048b818

+ 1 - 1
.agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-05-base-default-file-editor.zh.md)
 
 ## Problem
 
-The shared base selects both `read`/`write`/`edit` and `str_replace_editor`, which offer overlapping file editing interfaces. [Issue #3599](https://github.com/deepseek-harness/deepseek-harness/issues/3599) requests one default interface for base-backed profiles while preserving the dedicated minimal compositions.
+The shared base selects both `read`/`write`/`edit` and `str_replace_editor`, which offer overlapping file editing interfaces. Issue #3599 requests one default interface for base-backed profiles while preserving the dedicated minimal compositions.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/simplification/2026-09-05-base-default-file-editor.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## Problem
 
-共享 base 同时选择 `read`/`write`/`edit` 和 `str_replace_editor`,这些工具提供重叠的文件编辑接口。[Issue #3599](https://github.com/deepseek-harness/deepseek-harness/issues/3599) 要求基于 base 的 profile 默认使用一套接口,同时保留专用的极简组合。
+共享 base 同时选择 `read`/`write`/`edit` 和 `str_replace_editor`,这些工具提供重叠的文件编辑接口。Issue #3599 要求基于 base 的 profile 默认使用一套接口,同时保留专用的极简组合。
 
 ## Decision
 

+ 2 - 2
.agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-evidence.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-evidence.md
-2026-09-06-agent-request-freeze-evidence.md: e597ba64ca31da5e6a9be0c32b922737c9a1c074
-2026-09-06-agent-request-freeze-evidence.zh.md: b56b14ca0d69261eddd46507deffba9fa945004c
+2026-09-06-agent-request-freeze-evidence.md: e370b2556f212455239bceca70ec2661ce7787e0
+2026-09-06-agent-request-freeze-evidence.zh.md: 970e5f84f2593be738388aad828a2a83146a3a34

+ 6 - 6
.agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-evidence.md

@@ -36,9 +36,9 @@ An earlier original-code run at 06:58:28 UTC overlaps a sibling build because of
 
 ### Standard hosted CI calibration
 
-The standard two-CPU `ubuntu-24.04` lane runs Node 24.20.0. [Run 34033336380, job 101487280801](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336380/job/101487280801) measures the optimized request path at merge commit `8fba64d9ae06d1a9a778a95487bb915d24cb0644` in Azure eastus: 183.355397, 184.468253, 185.042397, 182.160790, 182.924728 ms; median 183.355397 ms. Every sample completes the same 40 requests and 13,923 events. All five exceed the historical 175 ms budget without changing the WeakSet implementation or workload.
+The standard two-CPU `ubuntu-24.04` lane runs Node 24.20.0. Run 34033336380, job 101487280801 measures the optimized request path at merge commit `8fba64d9ae06d1a9a778a95487bb915d24cb0644` in Azure eastus: 183.355397, 184.468253, 185.042397, 182.160790, 182.924728 ms; median 183.355397 ms. Every sample completes the same 40 requests and 13,923 events. All five exceed the historical 175 ms budget without changing the WeakSet implementation or workload.
 
-A second hosted run of the same request implementation, [run 34033336246, job 101487216170](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336246/job/101487216170), records 145.644577, 144.204300, 143.072572, 145.985903, 146.834474 ms; median 145.644577 ms. It uses the same Ubuntu image and Node version but a different worker in Azure westus3 at merge commit `c366e49`. This faster run does not replace the eastus evidence or establish why the workers differ. The older self-hosted `VM-7-113-ubuntu-ci-9` run with Node 24.18.1 ([run 34021903421, job 101456015028](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34021903421/job/101456015028)) records 110.025154, 119.958978, 108.266860, 107.557950, 108.538902 ms; median 108.538902 ms. Its runner and Node version do not calibrate the standard hosted lane.
+A second hosted run of the same request implementation, run 34033336246, job 101487216170, records 145.644577, 144.204300, 143.072572, 145.985903, 146.834474 ms; median 145.644577 ms. It uses the same Ubuntu image and Node version but a different worker in Azure westus3 at merge commit `c366e49`. This faster run does not replace the eastus evidence or establish why the workers differ. The older self-hosted `VM-7-113-ubuntu-ci-9` run with Node 24.18.1 (run 34021903421, job 101456015028) records 110.025154, 119.958978, 108.266860, 107.557950, 108.538902 ms; median 108.538902 ms. Its runner and Node version do not calibrate the standard hosted lane.
 
 The current request-history median limit is 297 ms. It is the largest integer within a 25% increase from the initial 238 ms limit: `floor(238 × 1.25) = 297`, an increase of 24.79%. This allowance belongs only to `agent-continuation/request-history`; the shared time scale, variance headroom, other time limits, memory limits, sample count, and workload remain unchanged.
 
@@ -46,13 +46,13 @@ The standard GitHub Actions `ubuntu-24.04` runner group reports the same image `
 
 | Hosted measurement | Request-history raw totals (ms) | Median (ms) |
 |---|---|---:|
-| [Release `f778396b2e`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34233932940/job/102086694864) | 152.649609, 154.616595, 144.588531, 144.261013, 154.017377 | 152.649609 |
-| [Release `a0a61a8237`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34235890227/job/102095345914) | 246.876615, 246.881047, 272.370218, 265.796833, 272.507507 | 265.796833 |
-| [Reference `35fcb95275`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34232504298/job/102084171717) | 279.689489, 297.792849, 263.178391, 252.660292, 251.267361 | 263.178391 |
+| PR #3797 merge (run 34233932940, job 102086694864) | 152.649609, 154.616595, 144.588531, 144.261013, 154.017377 | 152.649609 |
+| PR #3799 merge (run 34235890227, job 102095345914) | 246.876615, 246.881047, 272.370218, 265.796833, 272.507507 | 265.796833 |
+| Integration reference (run 34232504298, job 102084171717) | 279.689489, 297.792849, 263.178391, 252.660292, 251.267361 | 263.178391 |
 
 The first two jobs also differ across tool continuation (483.877/698.657 ms), catalog (612.127/1009.367 ms), and profile continuation (2438.362/3854.950 ms). These observations establish broad hosted execution-time variation; they do not identify a hardware fault or a runtime regression. Among these four continuation scenarios, only request history crosses its limit in the slower release run.
 
-A bounded profile of `a0a61a8237` on Apple M4 Pro / Node 24.19.0 retains five fresh-process totals: 70.198916, 67.432250, 65.151208, 66.473292, 71.049667 ms; median 67.432250 ms. Every sample completes the same 40 requests and 13,925 events. Sampling attributes 38.082 ms inclusive time to adapter dispatch, including 10.878 ms of required file-content traversal; system-node scanning takes 4.127 ms, while the one-time restored-event reversal takes 0.291 ms outside the timed turns. Removing the latter cannot explain the observed turn cost. Caching projected content or system nodes adds immutability or invalidation obligations beyond this bounded allowance. Runtime code is unchanged.
+A bounded profile of the PR #3799 merge on Apple M4 Pro / Node 24.19.0 retains five fresh-process totals: 70.198916, 67.432250, 65.151208, 66.473292, 71.049667 ms; median 67.432250 ms. Every sample completes the same 40 requests and 13,925 events. Sampling attributes 38.082 ms inclusive time to adapter dispatch, including 10.878 ms of required file-content traversal; system-node scanning takes 4.127 ms, while the one-time restored-event reversal takes 0.291 ms outside the timed turns. Removing the latter cannot explain the observed turn cost. Caching projected content or system nodes adds immutability or invalidation obligations beyond this bounded allowance. Runtime code is unchanged.
 
 Deterministic controls call the timed case's `assertRequestHistoryBudget`. They accept the recorded 185.042397 ms maximum and the two slower hosted medians, while rejecting a synthetic 310 ms median from 308, 310, 312, 311, 309 ms inputs. The slower-host acceptance control reproduces `265.796833 > 238` before the allowance; the complete owner file passes 11 tests at 297 ms. Replaying recorded values validates the assertion, not a new hosted run. The historical 250 ms synthetic case and the 246.130875 ms original M4 measurement fit this allowance and are no longer rejection controls; original/optimized M4 measurements remain evidence of the freeze implementation's gain.
 

+ 6 - 6
.agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-evidence.zh.md

@@ -36,9 +36,9 @@ Apple M4 Pro、macOS arm64、Node 24.19.0;worktree 使用独立依赖和构建
 
 ### 标准托管 CI 校准
 
-标准双 CPU `ubuntu-24.04` 测试通道运行 Node 24.20.0。[运行 34033336380、任务 101487280801](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336380/job/101487280801)在 Azure eastus 上测量合并提交 `8fba64d9ae06d1a9a778a95487bb915d24cb0644` 的优化请求路径:183.355397, 184.468253, 185.042397, 182.160790, 182.924728 ms;中位数 183.355397 ms。每个样本都完成相同的 40 个请求和 13,923 个事件。在 WeakSet 实现与工作负载未变的情况下,全部五个样本均超过历史 175 ms 预算。
+标准双 CPU `ubuntu-24.04` 测试通道运行 Node 24.20.0。运行 34033336380、任务 101487280801在 Azure eastus 上测量合并提交 `8fba64d9ae06d1a9a778a95487bb915d24cb0644` 的优化请求路径:183.355397, 184.468253, 185.042397, 182.160790, 182.924728 ms;中位数 183.355397 ms。每个样本都完成相同的 40 个请求和 13,923 个事件。在 WeakSet 实现与工作负载未变的情况下,全部五个样本均超过历史 175 ms 预算。
 
-相同请求实现的另一次托管运行,[运行 34033336246、任务 101487216170](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336246/job/101487216170),记录了 145.644577, 144.204300, 143.072572, 145.985903, 146.834474 ms;中位数 145.644577 ms。它在合并提交 `c366e49` 上使用相同的 Ubuntu 镜像和 Node 版本,但运行于 Azure westus3 的另一台工作机。较快的运行不能替代 eastus 证据,也不能证明工作机差异的原因。较早的自托管 `VM-7-113-ubuntu-ci-9` 运行使用 Node 24.18.1([运行 34021903421、任务 101456015028](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34021903421/job/101456015028)),记录了 110.025154, 119.958978, 108.266860, 107.557950, 108.538902 ms;中位数 108.538902 ms。其运行器和 Node 版本不能校准标准托管通道。
+相同请求实现的另一次托管运行,运行 34033336246、任务 101487216170,记录了 145.644577, 144.204300, 143.072572, 145.985903, 146.834474 ms;中位数 145.644577 ms。它在合并提交 `c366e49` 上使用相同的 Ubuntu 镜像和 Node 版本,但运行于 Azure westus3 的另一台工作机。较快的运行不能替代 eastus 证据,也不能证明工作机差异的原因。较早的自托管 `VM-7-113-ubuntu-ci-9` 运行使用 Node 24.18.1(运行 34021903421、任务 101456015028),记录了 110.025154, 119.958978, 108.266860, 107.557950, 108.538902 ms;中位数 108.538902 ms。其运行器和 Node 版本不能校准标准托管通道。
 
 当前请求历史中位数上限为 297 ms。这是在最初 238 ms 上限基础上增加不超过 25% 的最大整数:`floor(238 × 1.25) = 297`,增加 24.79%。该余量仅属于 `agent-continuation/request-history`;共享时间系数、波动余量、其他时间上限、内存上限、采样次数与工作负载均保持不变。
 
@@ -46,13 +46,13 @@ Apple M4 Pro、macOS arm64、Node 24.19.0;worktree 使用独立依赖和构建
 
 | 托管测量 | 请求历史原始总耗时(ms) | 中位数(ms) |
 |---|---|---:|
-| [Release `f778396b2e`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34233932940/job/102086694864) | 152.649609, 154.616595, 144.588531, 144.261013, 154.017377 | 152.649609 |
-| [Release `a0a61a8237`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34235890227/job/102095345914) | 246.876615, 246.881047, 272.370218, 265.796833, 272.507507 | 265.796833 |
-| [参考 `35fcb95275`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34232504298/job/102084171717) | 279.689489, 297.792849, 263.178391, 252.660292, 251.267361 | 263.178391 |
+| PR #3797 merge (run 34233932940, job 102086694864) | 152.649609, 154.616595, 144.588531, 144.261013, 154.017377 | 152.649609 |
+| PR #3799 merge (run 34235890227, job 102095345914) | 246.876615, 246.881047, 272.370218, 265.796833, 272.507507 | 265.796833 |
+| 集成参考 (run 34232504298, job 102084171717) | 279.689489, 297.792849, 263.178391, 252.660292, 251.267361 | 263.178391 |
 
 前两次任务的工具续跑(483.877/698.657 ms)、catalog(612.127/1009.367 ms)和 profile 续跑(2438.362/3854.950 ms)也存在差异。这些观测证明托管执行时间存在广泛波动,不能据此断言硬件故障或运行时回归。在上述四个续跑场景中,较慢的 release 运行仅请求历史超过其上限。
 
-对 `a0a61a8237` 在 Apple M4 Pro / Node 24.19.0 上做的有界性能分析保留五次新进程总耗时:70.198916, 67.432250, 65.151208, 66.473292, 71.049667 ms;中位数 67.432250 ms。每个样本均完成相同的 40 个请求和 13,925 个事件。采样将适配器分派的包含后代耗时记为 38.082 ms,其中必需的文件内容遍历占 10.878 ms;system 节点扫描占 4.127 ms,一次性的恢复事件倒序则在计时轮次之外占 0.291 ms。删除后者不能解释观测到的轮次耗时。缓存投影内容或 system 节点会引入超出本次有界余量调整的不可变性或失效管理义务。运行时代码保持不变。
+对 PR #3799 合并结果 在 Apple M4 Pro / Node 24.19.0 上做的有界性能分析保留五次新进程总耗时:70.198916, 67.432250, 65.151208, 66.473292, 71.049667 ms;中位数 67.432250 ms。每个样本均完成相同的 40 个请求和 13,925 个事件。采样将适配器分派的包含后代耗时记为 38.082 ms,其中必需的文件内容遍历占 10.878 ms;system 节点扫描占 4.127 ms,一次性的恢复事件倒序则在计时轮次之外占 0.291 ms。删除后者不能解释观测到的轮次耗时。缓存投影内容或 system 节点会引入超出本次有界余量调整的不可变性或失效管理义务。运行时代码保持不变。
 
 确定性对照调用计时场景使用的 `assertRequestHistoryBudget`。它们接受已记录的 185.042397 ms 最大值和两次较慢托管中位数,同时拒绝由 308, 310, 312, 311, 309 ms 输入得到的合成 310 ms 中位数。较慢运行器接受对照在增加余量前复现 `265.796833 > 238`;完整所属文件在 297 ms 下通过 11 个测试。回放已记录数值验证的是断言,而非新的托管运行。历史合成 250 ms 场景与原版 M4 的 246.130875 ms 测量符合该余量,不再作为拒绝对照;原版/优化版 M4 测量仍保留为冻结实现收益的证据。
 

+ 2 - 2
.agents/notes/implemented/simplification/2026-09-07-file-content-scan.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/simplification/2026-09-07-file-content-scan.md
-2026-09-07-file-content-scan.md: 5a539f0b9c4e00545d9894647599aed713285014
-2026-09-07-file-content-scan.zh.md: 9a8fe390b8b430baf7975bfc1391ccd09ae1db6e
+2026-09-07-file-content-scan.md: bf4f03ee5a7134b11c7cbec322b0a37a3f2a6304
+2026-09-07-file-content-scan.zh.md: 5efb1589b036f80b46752a80103c080952c82396

+ 1 - 1
.agents/notes/implemented/simplification/2026-09-07-file-content-scan.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-07-file-content-scan.zh.md)
 
 ## Problem
 
-Every model dispatch checks complete message content for files, including nested tool results. A request-history CPU profile attributes 23.540 ms of self time to `contentHasFile` and 5.584 ms to its callback. This traversal remains necessary even after [loop-owned freeze evidence](2026-09-06-agent-request-freeze-evidence.md) removes repeated request freezing. The hot LLM source is identical at master `bd5917`, master `112a5`, and the measured `f834b002826453e7918eeb558d052b2c24c56a76`; these observations do not establish PR causality.
+Every model dispatch checks complete message content for files, including nested tool results. A request-history CPU profile attributes 23.540 ms of self time to `contentHasFile` and 5.584 ms to its callback. This traversal remains necessary even after [loop-owned freeze evidence](2026-09-06-agent-request-freeze-evidence.md) removes repeated request freezing. The hot LLM source is identical at both sampled master revisions and the measured V3 integration revision; these observations do not establish PR causality.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/simplification/2026-09-07-file-content-scan.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## Problem
 
-每次模型分发都检查完整消息内容中的文件,包括嵌套工具结果。请求历史 CPU profile 将 23.540 ms 自身时间归于 `contentHasFile`,将 5.584 ms 归于其回调。即使[循环自有冻结证据](2026-09-06-agent-request-freeze-evidence.zh.md)消除了重复请求冻结,这次遍历仍然必需。master `bd5917`、master `112a5` 与实测的 `f834b002826453e7918eeb558d052b2c24c56a76` 的 LLM 热点源码完全相同;这些观察不能证明 PR 因果关系。
+每次模型分发都检查完整消息内容中的文件,包括嵌套工具结果。请求历史 CPU profile 将 23.540 ms 自身时间归于 `contentHasFile`,将 5.584 ms 归于其回调。即使[循环自有冻结证据](2026-09-06-agent-request-freeze-evidence.zh.md)消除了重复请求冻结,这次遍历仍然必需。两次抽样的 master 修订与实测的 V3 集成修订 的 LLM 热点源码完全相同;这些观察不能证明 PR 因果关系。
 
 ## Decision
 

+ 2 - 2
.agents/notes/implemented/simplification/2026-09-09-nontransactional-loader.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/simplification/2026-09-09-nontransactional-loader.md
-2026-09-09-nontransactional-loader.md: 7ea4db0abf0fb65e443126dbca9fd822e47ce2d0
-2026-09-09-nontransactional-loader.zh.md: 5293efb58998f68a25e1143a2d043da32ef06bfb
+2026-09-09-nontransactional-loader.md: 6a3cb1476e1f296199067a8d048c6d7c8bba4069
+2026-09-09-nontransactional-loader.zh.md: 4e91a31ed86be4e45313e1322db8af486477aea7

+ 1 - 1
.agents/notes/implemented/simplification/2026-09-09-nontransactional-loader.md

@@ -10,7 +10,7 @@ Transactional config reload preserves an old plugin generation after a failed ed
 
 ## Decision
 
-Revert the five commits in [#932](https://github.com/deepseek-harness/deepseek-harness/pull/932), resolving package moves and retaining independent later behavior. The reported merge commit belongs to the larger #936 dependency chain; reverting its first-parent diff would remove unrelated repository-plugin support. The [vendor ledger](../../../../vendor/README.md#local-modifications) records every retained source change against the unchanged pins.
+Revert the five commits in #932, resolving package moves and retaining independent later behavior. The reported merge commit belongs to the larger #936 dependency chain; reverting its first-parent diff would remove unrelated repository-plugin support. The [vendor ledger](../../../../vendor/README.md#local-modifications) records every retained source change against the unchanged pins.
 
 Loader changes entry options eagerly. EntryGroup starts siblings concurrently and logs application failures; EntryTree waits for outstanding work without rejecting failed fibers. Neither restores a previous plugin or configuration. Include retains parse validation and patch reapplication, but plugin failures can leave a partially applied tree.
 

+ 1 - 1
.agents/notes/implemented/simplification/2026-09-09-nontransactional-loader.zh.md

@@ -10,7 +10,7 @@ Status: implemented
 
 ## 决策
 
-撤销 [#932](https://github.com/deepseek-harness/deepseek-harness/pull/932) 中的五个提交,解决包移动冲突并保留后续独立行为。记录的合并提交属于更大的 #936 依赖链;撤销其第一父提交差异还会删除无关的仓库插件支持。[Vendor 修改记录](../../../../vendor/README.md#local-modifications) 按不变的固定来源记录每项保留的源码更改。
+撤销 #932 中的五个提交,解决包移动冲突并保留后续独立行为。记录的合并提交属于更大的 #936 依赖链;撤销其第一父提交差异还会删除无关的仓库插件支持。[Vendor 修改记录](../../../../vendor/README.md#local-modifications) 按不变的固定来源记录每项保留的源码更改。
 
 Loader 立即更改条目选项。EntryGroup 并发启动同级条目并记录应用失败;EntryTree 等待未完成的工作,但不因失败的 fiber 而拒绝。两者均不恢复旧插件或配置。Include 保留解析校验和 patch 重应用,但插件失败可能留下部分应用的配置树。
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
-2026-09-04-session-open-performance-gate.md: c0c337d1adbfda18d4d91631720b37caa651ae53
-2026-09-04-session-open-performance-gate.zh.md: d9989a6050dbc97d5ecb4ad28d7e1c0a0052c121
+2026-09-04-session-open-performance-gate.md: eaf7c6045bca12329d81cf9c968fb191229b0add
+2026-09-04-session-open-performance-gate.zh.md: e9bf13460aafbd962a1d15b7a7d240f0b4be6b4e

+ 4 - 4
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md

@@ -37,7 +37,7 @@ Normal-heap mode performs a fixed pair of explicit garbage collections after Hos
 
 A small untimed fixture prerequisite verifies current migration, message preservation, immutable V0 bytes, and successor reopen. Worker failures retain the first and last ten stderr lines, or fatal heap diagnostics, so setup rejection remains distinguishable from a budget breach. The timed performance cases do not duplicate semantic assertions owned by functional tests; it requires only that the target call completes and reaches its measured endpoint. The Client-fold benchmark continues to use the real `ConversationNodeAssembler` and every Chat Definition, and requires both the large window's absolute time and its scaling relative to the small window to remain below fixed budgets.
 
-Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Except for current-generation `open` and first-open Agent resume, source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. Current-generation `open` uses a directly measured standard-runner expectation of 50 ms with only the 1.25× headroom, rounded up to a 63 ms budget. First-open Agent resume uses a reviewed 562 ms hosted limit. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
+Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Except for current-generation `open` and first-open Agent resume, source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. Current-generation `open` uses a directly measured standard-runner expectation of 50 ms with only the 1.25× headroom, rounded up to a 63 ms budget. First-open Agent resume uses a reviewed 562 ms hosted limit. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. The PR #3533 merge is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
 
 ## Calibration evidence
 
@@ -54,11 +54,11 @@ Five-sample medians on the same Node 24 reference machine establish the positive
 
 The pre-stack implementation keeps V0 as its current format, so first open does not change its on-disk representation; its native V0 first-history and Agent-resume measurements therefore apply to both lifecycle rows.
 
-The [standard two-CPU run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384/job/101461539961) at `ca3ffe95dac2c55eefeb16ed9b61067bbd19ee90` uses Node 24.20.0 x64 and Ubuntu image `20260831.293.1`. Its five current-generation `open` samples are 49.2, 47.4, 49.1, 48.6, and 48.1 ms: median 48.6 ms, maximum 49.2 ms. The rounded 50 ms CI expectation gives a 63 ms limit without reapplying the 2× machine scale. The log identifies two available CPUs but not their model; it does not isolate hardware from the Node-version change. This is endpoint-specific runner calibration, not evidence of an application optimization or a new reference-machine measurement. Every other benchmark passes its existing budget. Deterministic controls reject the observed median at the historical 30 ms limit, accept it at 63 ms, reject a synthetic 75 ms reopen median, and reject a synthetic 4,000 ms first-open duration at its unchanged 550 ms limit. These controls verify budget enforcement, not a measured new regression.
+The standard two-CPU run (run 34023970384, job 101461539961) for PR #3640 uses Node 24.20.0 x64 and Ubuntu image `20260831.293.1`. Its five current-generation `open` samples are 49.2, 47.4, 49.1, 48.6, and 48.1 ms: median 48.6 ms, maximum 49.2 ms. The rounded 50 ms CI expectation gives a 63 ms limit without reapplying the 2× machine scale. The log identifies two available CPUs but not their model; it does not isolate hardware from the Node-version change. This is endpoint-specific runner calibration, not evidence of an application optimization or a new reference-machine measurement. Every other benchmark passes its existing budget. Deterministic controls reject the observed median at the historical 30 ms limit, accept it at 63 ms, reject a synthetic 75 ms reopen median, and reject a synthetic 4,000 ms first-open duration at its unchanged 550 ms limit. These controls verify budget enforcement, not a measured new regression.
 
-A cold-verifier packaging change removes runtime workspace-module loading without changing these budgets or the measured endpoint. On macOS arm64, Node 24.18.0, the same 127,400-event fixture at `ac48359b195558806ee5a2286697074fd1a52815` takes 164.2, 162.4, 159.7, 149.3, and 167.7 ms for first writable resume (median 162.4 ms). Bundling the verifier through the workspace build gives 119.8, 120.9, 121.9, 122.3, and 121.8 ms (median 121.8 ms, 25% lower). Retained heap stays at 5.4 MB; median peak RSS changes from 144.9 to 143.7 MB. Reopen medians are 27.5 and 27.1 ms, and all 16 Session cases, including the 128 MB completion checks, pass. A CPU profile attributes part of the old verifier cost to module resolution and compilation. The isolated-package built-worker test fails on the original worker because its workspace imports cannot resolve, and passes with the bundled worker, including rejection of an incorrect event count. These local results do not establish Linux runner timing; those cases use the 450 ms CI limit.
+A cold-verifier packaging change removes runtime workspace-module loading without changing these budgets or the measured endpoint. On macOS arm64, Node 24.18.0, the same 127,400-event fixture before verifier bundling takes 164.2, 162.4, 159.7, 149.3, and 167.7 ms for first writable resume (median 162.4 ms). Bundling the verifier through the workspace build gives 119.8, 120.9, 121.9, 122.3, and 121.8 ms (median 121.8 ms, 25% lower). Retained heap stays at 5.4 MB; median peak RSS changes from 144.9 to 143.7 MB. Reopen medians are 27.5 and 27.1 ms, and all 16 Session cases, including the 128 MB completion checks, pass. A CPU profile attributes part of the old verifier cost to module resolution and compilation. The isolated-package built-worker test fails on the original worker because its workspace imports cannot resolve, and passes with the bundled worker, including rejection of an incorrect event count. These local results do not establish Linux runner timing; those cases use the 450 ms CI limit.
 
-The [hosted run at `a7884138be`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34265057987/job/102192211510) includes the bundled verifier and reports first-open Agent-resume samples of 454.2, 454.8, 455.4, 457.8, and 459.8 ms: median 455.4 ms against 450 ms. The code-equivalent [preceding run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34263062688/job/102185561214) reports a 436.2 ms median; only the bilingual request-history README and its pairing record differ between those heads. The reviewed ceiling is 562 ms, `floor(450 × 1.25)`, a 24.89% increase that leaves 23.4% above the observed 455.4 ms median. Five fresh M4 Pro / Node 24.19 samples span 150.07–158.22 ms with a 152.57 ms median. A bounded main-thread profile identifies no obvious small optimization; it excludes verifier-thread CPU and does not establish the cause of hosted variation. Controls reject the recorded median at 450 ms, accept it at 562 ms, and reject 600 ms. The bundled-verifier improvement, workload, other time budgets, shared scaling, and memory limits remain intact.
+The hosted run (run 34265057987, job 102192211510) includes the bundled verifier and reports first-open Agent-resume samples of 454.2, 454.8, 455.4, 457.8, and 459.8 ms: median 455.4 ms against 450 ms. The code-equivalent preceding run (run 34263062688, job 102185561214) reports a 436.2 ms median; only the bilingual request-history README and its pairing record differ between those heads. The reviewed ceiling is 562 ms, `floor(450 × 1.25)`, a 24.89% increase that leaves 23.4% above the observed 455.4 ms median. Five fresh M4 Pro / Node 24.19 samples span 150.07–158.22 ms with a 152.57 ms median. A bounded main-thread profile identifies no obvious small optimization; it excludes verifier-thread CPU and does not establish the cause of hosted variation. Controls reject the recorded median at 450 ms, accept it at 562 ms, and reject 600 ms. The bundled-verifier improvement, workload, other time budgets, shared scaling, and memory limits remain intact.
 
 The calibrated source budgets are:
 

+ 4 - 4
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md

@@ -37,7 +37,7 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 一个不计时的小型 fixture 前置用例验证当前 migration、消息保留、V0 字节不变及后继再次打开。Worker 失败时保留 stderr 首尾各十行或致命堆错误,使准备阶段拒绝与预算超限可区分。计时性能用例不重复功能测试的内容断言,只要求目标调用完成并到达对应的可观察终点。Client fold benchmark 继续使用真实 `ConversationNodeAssembler` 与全部 Chat Definition,要求大窗口的绝对时间和相对小窗口的缩放比均低于固定预算。
 
-预算按各测量终点分别校准。两次 Node 24.19 x64 CI 运行的中位数最大相差 5.2%;其 CPU 密集型壁钟时间是 Node 24.18 arm64 参考运行的 1.95–2.06 倍。除当前 generation `open` 和 first-open Agent resume 外,源码常量记录参考机器上的预期耗时;`ciTimeBudget()` 将其乘以实测的 2 倍 CI 时间系数和 1.25 倍波动余量。当前 generation `open` 使用标准运行器直接测得的 50 ms 预期值,仅乘以 1.25 倍余量,向上取整得到 63 ms 预算。First-open Agent resume 使用经审查的 562 ms 托管上限。GC 后增量堆与 Client fold 缩放预算不属于壁钟时间,因此只使用 1.25 倍余量。128 MB 完成性检查仍是独立的瞬时分配限制。由此得到的 first-open 时间上限、受限堆检查与 Client fold 上限都会拒绝已知退化。栈前参考提交固定为 `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5`,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
+预算按各测量终点分别校准。两次 Node 24.19 x64 CI 运行的中位数最大相差 5.2%;其 CPU 密集型壁钟时间是 Node 24.18 arm64 参考运行的 1.95–2.06 倍。除当前 generation `open` 和 first-open Agent resume 外,源码常量记录参考机器上的预期耗时;`ciTimeBudget()` 将其乘以实测的 2 倍 CI 时间系数和 1.25 倍波动余量。当前 generation `open` 使用标准运行器直接测得的 50 ms 预期值,仅乘以 1.25 倍余量,向上取整得到 63 ms 预算。First-open Agent resume 使用经审查的 562 ms 托管上限。GC 后增量堆与 Client fold 缩放预算不属于壁钟时间,因此只使用 1.25 倍余量。128 MB 完成性检查仍是独立的瞬时分配限制。由此得到的 first-open 时间上限、受限堆检查与 Client fold 上限都会拒绝已知退化。参考固定为 PR #3533 的合并结果,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
 
 ## 校准证据
 
@@ -54,11 +54,11 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 栈前实现以 V0 作为当前格式,因此 first open 不改变磁盘表示;它的原生 V0 首屏历史与 Agent resume 测量同时适用于两个生命周期行。
 
-`ca3ffe95dac2c55eefeb16ed9b61067bbd19ee90` 上的[标准双 CPU 运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384/job/101461539961)使用 Node 24.20.0 x64 和 Ubuntu 镜像 `20260831.293.1`。当前 generation `open` 的五次样本为 49.2、47.4、49.1、48.6 和 48.1 ms:中位数 48.6 ms,最大值 49.2 ms。取整后的 50 ms CI 预期值给出 63 ms 上限,不重复乘以 2 倍机器系数。日志标明两个可用 CPU,但未记录型号;它无法区分硬件变化与 Node 版本变化的影响。这是端点专属的运行器校准,不是应用优化或参考机器新测量的证据。其他每项 benchmark 均通过既有预算。确定性正反例在历史 30 ms 上限下拒绝实测中位数,在 63 ms 下接受它,拒绝合成的 75 ms reopen 中位数,并以未改变的 550 ms 上限拒绝合成的 4,000 ms 首次打开耗时。这些正反例验证预算执行,不代表测得新的退化。
+PR #3640 的标准双 CPU 运行 (run 34023970384, job 101461539961)使用 Node 24.20.0 x64 和 Ubuntu 镜像 `20260831.293.1`。当前 generation `open` 的五次样本为 49.2、47.4、49.1、48.6 和 48.1 ms:中位数 48.6 ms,最大值 49.2 ms。取整后的 50 ms CI 预期值给出 63 ms 上限,不重复乘以 2 倍机器系数。日志标明两个可用 CPU,但未记录型号;它无法区分硬件变化与 Node 版本变化的影响。这是端点专属的运行器校准,不是应用优化或参考机器新测量的证据。其他每项 benchmark 均通过既有预算。确定性正反例在历史 30 ms 上限下拒绝实测中位数,在 63 ms 下接受它,拒绝合成的 75 ms reopen 中位数,并以未改变的 550 ms 上限拒绝合成的 4,000 ms 首次打开耗时。这些正反例验证预算执行,不代表测得新的退化。
 
-一次冷 verifier 打包调整移除了运行时 workspace 模块加载,未改变这些预算或测量终点。在 macOS arm64、Node 24.18.0 上,`ac48359b195558806ee5a2286697074fd1a52815` 对同一份 127,400-event fixture 的首次 writable resume 耗时为 164.2、162.4、159.7、149.3、167.7 ms(中位数 162.4 ms)。通过 workspace build 打包 verifier 后为 119.8、120.9、121.9、122.3、121.8 ms(中位数 121.8 ms,降低 25%)。Retained heap 保持 5.4 MB;peak RSS 中位数从 144.9 变为 143.7 MB。Reopen 中位数为 27.5 和 27.1 ms,包含 128 MB completion check 的全部 16 项 Session 用例通过。CPU profile 将旧 verifier 的部分成本归因于模块解析和编译。隔离 package 的 built-worker 测试在旧 worker 上因无法解析 workspace import 而失败,在打包后的 worker 上通过,同时验证错误的 event count 会被拒绝。这些本地结果不能证明 Linux runner 耗时;这些用例使用 450 ms CI 上限。
+一次冷 verifier 打包调整移除了运行时 workspace 模块加载,未改变这些预算或测量终点。在 macOS arm64、Node 24.18.0 上,verifier 打包前对同一份 127,400-event fixture 的首次 writable resume 耗时为 164.2、162.4、159.7、149.3、167.7 ms(中位数 162.4 ms)。通过 workspace build 打包 verifier 后为 119.8、120.9、121.9、122.3、121.8 ms(中位数 121.8 ms,降低 25%)。Retained heap 保持 5.4 MB;peak RSS 中位数从 144.9 变为 143.7 MB。Reopen 中位数为 27.5 和 27.1 ms,包含 128 MB completion check 的全部 16 项 Session 用例通过。CPU profile 将旧 verifier 的部分成本归因于模块解析和编译。隔离 package 的 built-worker 测试在旧 worker 上因无法解析 workspace import 而失败,在打包后的 worker 上通过,同时验证错误的 event count 会被拒绝。这些本地结果不能证明 Linux runner 耗时;这些用例使用 450 ms CI 上限。
 
-[`a7884138be` 的托管运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34265057987/job/102192211510)包含已打包的 verifier,first-open Agent-resume 样本为 454.2、454.8、455.4、457.8 和 459.8 ms:中位数 455.4 ms,超过 450 ms。[代码等价的前一次运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34263062688/job/102185561214)报告 436.2 ms 中位数;两个 head 之间只有请求历史双语 README 及其配对记录不同。经审查的上限为 562 ms,即 `floor(450 × 1.25)`,增加 24.89%,比观测到的 455.4 ms 中位数高 23.4%。五次新进程 M4 Pro / Node 24.19 样本范围为 150.07–158.22 ms,中位数为 152.57 ms。一次有界的主线程 profile 未发现明显的小型优化;它不包含 verifier 线程 CPU,也不能证明托管耗时变化的原因。对照在 450 ms 下拒绝已记录中位数,在 562 ms 下接受该值,并拒绝 600 ms。Verifier 打包优化、工作负载、其他时间预算、共享缩放和内存限制均保持不变。
+托管运行 (run 34265057987, job 102192211510)包含已打包的 verifier,first-open Agent-resume 样本为 454.2、454.8、455.4、457.8 和 459.8 ms:中位数 455.4 ms,超过 450 ms。代码等价的前一次运行 (run 34263062688, job 102185561214)报告 436.2 ms 中位数;两个 head 之间只有请求历史双语 README 及其配对记录不同。经审查的上限为 562 ms,即 `floor(450 × 1.25)`,增加 24.89%,比观测到的 455.4 ms 中位数高 23.4%。五次新进程 M4 Pro / Node 24.19 样本范围为 150.07–158.22 ms,中位数为 152.57 ms。一次有界的主线程 profile 未发现明显的小型优化;它不包含 verifier 线程 CPU,也不能证明托管耗时变化的原因。对照在 450 ms 下拒绝已记录中位数,在 562 ms 下接受该值,并拒绝 600 ms。Verifier 打包优化、工作负载、其他时间预算、共享缩放和内存限制均保持不变。
 
 校准后的源码预算如下:
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.md
-2026-09-06-backend-continuation-performance.md: d72e3145d9bfe0ffabab2822c8a3313785fdd44f
-2026-09-06-backend-continuation-performance.zh.md: 66085e1310fbfb92cb384875015d3baa648650c5
+2026-09-06-backend-continuation-performance.md: a78dcb8bcc51b106123071951363090c20b98820
+2026-09-06-backend-continuation-performance.zh.md: 66f10d0b811cc354c18f0f59c6407b4d56de0723

+ 5 - 5
.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.md

@@ -31,7 +31,7 @@ The parent bounds every child to 60 seconds, checks timeout, signal, exit, and r
 
 ## Calibration evidence
 
-The implementation reference is `925e012340f033f0521e802ba8569ce6dd7ef1ac` on Apple M4 Pro, macOS arm64, Node 24.19.0. Two exclusive five-sample runs use the same seed and no product optimization. Durations below are milliseconds; source expectations round above the observed run medians rather than imposing an unimplemented optimization target.
+The implementation reference is the PR #3537 merge on Apple M4 Pro, macOS arm64, Node 24.19.0. Two exclusive five-sample runs use the same seed and no product optimization. Durations below are milliseconds; source expectations round above the observed run medians rather than imposing an unimplemented optimization target.
 
 | Case | Run 1 raw totals | Run 2 raw totals | Medians | Historical M4 expectation | Historical scaled budget |
 |---|---|---|---|---:|---:|
@@ -45,13 +45,13 @@ A separate plain-Node request-history CPU profile attributes 132.876 ms of sampl
 
 The shipped SDK variant completes 100 turns, 200 requests, and 800 real file reads. Its five-sample smoke totals are 1,521.773, 1,463.465, 1,689.701, 1,365.485, and 1,417.106 ms (median 1,463.465 ms); a full-suite repeat reports 1,596.183, 1,784.536, 2,120.082, 1,405.365, and 1,355.894 ms (median 1,596.183 ms). Its 1,700 ms reference expectation yields a 4,250 ms CI budget. The repeat also slows the unchanged service cases, so it is validation under variable host load rather than evidence to relax their exclusive calibration. The SDK process receives an allowlisted environment and private home/workspace. A 40-second deadline starts SDK shutdown; every path awaits the same memoized close promise before the outer worker’s 60-second deadline. Profile timing includes boot, all turns, and shutdown, reported separately; no parent-process CPU or heap metric is presented as server memory. The adapter does not serialize requests for an external model provider.
 
-The first Linux x64 CI measurement at commit `1dc3296eba631d51fbb3bb50e249bf3cc0fce9f6` ran on `VM-7-113-ubuntu-ci-10` with Node 24.18.1 ([run 34017868081, attempt 1, job 101444810498](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34017868081/job/101444810498)). The SDK median was 2,753.441 ms against its 4,250 ms budget, and tool-continuation retained-heap median was 22.274 MiB against 28.75 MiB. Request-history and tool-continuation time budgets failed: 785.498 ms against 550 ms and 1,077.285 ms against 850 ms, respectively. The unchanged Session-reopen open phase also failed at 31.6 ms against 30 ms. [Attempt 2, job 101447076381](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34017868081/job/101447076381) passed every benchmark on the same commit and unchanged budgets, but used `VM-7-113-ubuntu-ci-29` with Node 24.19.0. The gate runner suppressed successful child output, so that attempt supplies a passing verdict rather than raw medians. The changed runner and Node version prevent attributing the difference solely to contention or claiming stable repeated CI calibration; neither the budgets nor the shared scale are changed on this evidence.
+The first Linux x64 CI measurement at commit `1dc3296eba631d51fbb3bb50e249bf3cc0fce9f6` ran on `VM-7-113-ubuntu-ci-10` with Node 24.18.1 (run 34017868081, attempt 1, job 101444810498). The SDK median was 2,753.441 ms against its 4,250 ms budget, and tool-continuation retained-heap median was 22.274 MiB against 28.75 MiB. Request-history and tool-continuation time budgets failed: 785.498 ms against 550 ms and 1,077.285 ms against 850 ms, respectively. The unchanged Session-reopen open phase also failed at 31.6 ms against 30 ms. Attempt 2, job 101447076381 (run 34017868081) passed every benchmark on the same commit and unchanged budgets, but used `VM-7-113-ubuntu-ci-29` with Node 24.19.0. The gate runner suppressed successful child output, so that attempt supplies a passing verdict rather than raw medians. The changed runner and Node version prevent attributing the difference solely to contention or claiming stable repeated CI calibration; neither the budgets nor the shared scale are changed on this evidence.
 
-Catalog uses an explicit 900 ms expected CI duration and only the existing 1.25× headroom, yielding 1,125 ms without applying the reference-machine scale again. The standard two-CPU hosted `ubuntu-24.04` [run 34033336380, job 101487280801](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336380/job/101487280801) reports five unchanged-catalog totals of 797.374, 883.157, 858.364, 790.569, and 904.579 ms: median 858.364 ms exceeds the historical 800 ms budget. The 320 ms M4 expectation above remains historical evidence, not a CI measurement. This follows the explicit-CI calibration used by Session reopening (50 ms expected CI); shared factors, workloads, timing endpoints, and product implementations remain unchanged. Deterministic controls use the same assertion as the measured verdict: the unrounded recorded median passes 1,125 ms and fails 800 ms, while a synthetic 1,400 ms median fails 1,125 ms. A passing run on a faster host does not calibrate the standard hosted runner.
+Catalog uses an explicit 900 ms expected CI duration and only the existing 1.25× headroom, yielding 1,125 ms without applying the reference-machine scale again. The standard two-CPU hosted `ubuntu-24.04` run 34033336380, job 101487280801 reports five unchanged-catalog totals of 797.374, 883.157, 858.364, 790.569, and 904.579 ms: median 858.364 ms exceeds the historical 800 ms budget. The 320 ms M4 expectation above remains historical evidence, not a CI measurement. This follows the explicit-CI calibration used by Session reopening (50 ms expected CI); shared factors, workloads, timing endpoints, and product implementations remain unchanged. Deterministic controls use the same assertion as the measured verdict: the unrounded recorded median passes 1,125 ms and fails 800 ms, while a synthetic 1,400 ms median fails 1,125 ms. A passing run on a faster host does not calibrate the standard hosted runner.
 
-Tool continuation also uses a 900 ms expected CI duration with 1.25× headroom (1,125 ms). At unchanged implementation `79c052ab29`, standard two-CPU hosted [run 34034524265, job 101490056074](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34034524265/job/101490056074) reports totals of 917.007, 892.091, 887.839, 905.659, and 898.252 ms: median 898.252 ms exceeds the historical 850 ms budget. The 340 ms M4 expectation remains historical evidence. The same measured-verdict assertion accepts the recorded unrounded median under 1,125 ms, rejects it under 850 ms, and rejects a synthetic 1,400 ms regression. Workload, timing, product code, and the 28.75 MiB retained-heap budget remain unchanged.
+Tool continuation also uses a 900 ms expected CI duration with 1.25× headroom (1,125 ms). At unchanged implementation `79c052ab29`, standard two-CPU hosted run 34034524265, job 101490056074 reports totals of 917.007, 892.091, 887.839, 905.659, and 898.252 ms: median 898.252 ms exceeds the historical 850 ms budget. The 340 ms M4 expectation remains historical evidence. The same measured-verdict assertion accepts the recorded unrounded median under 1,125 ms, rejects it under 850 ms, and rejects a synthetic 1,400 ms regression. Workload, timing, product code, and the 28.75 MiB retained-heap budget remain unchanged.
 
-Baseline request history uses a 600 ms expected CI duration with 1.25× headroom (750 ms). At unchanged implementation `54d1190a75`, standard two-CPU hosted [run 34035306987, job 101492163630](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34035306987/job/101492163630) reports totals of 618.598, 618.606, 582.035, 582.304, and 581.832 ms: median 582.304 ms exceeds the historical 550 ms budget. The 220 ms M4 expectation remains historical evidence. The same measured-verdict assertion accepts the recorded unrounded median under 750 ms, rejects it under 550 ms, and rejects a synthetic 900 ms regression. This calibrates the unoptimized baseline only; workload, timing, product code, and memory budgets remain unchanged.
+Baseline request history uses a 600 ms expected CI duration with 1.25× headroom (750 ms). At unchanged implementation `54d1190a75`, standard two-CPU hosted run 34035306987, job 101492163630 reports totals of 618.598, 618.606, 582.035, 582.304, and 581.832 ms: median 582.304 ms exceeds the historical 550 ms budget. The 220 ms M4 expectation remains historical evidence. The same measured-verdict assertion accepts the recorded unrounded median under 750 ms, rejects it under 550 ms, and rejects a synthetic 900 ms regression. This calibrates the unoptimized baseline only; workload, timing, product code, and memory budgets remain unchanged.
 
 ## Alternatives considered
 

+ 5 - 5
.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.zh.md

@@ -31,7 +31,7 @@ SDK fixture 通过 profile patch 显式插入 `fs-local` 和 `str_replace_editor
 
 ## 校准证据
 
-实现参考为 Apple M4 Pro、macOS arm64、Node 24.19.0 上的 `925e012340f033f0521e802ba8569ce6dd7ef1ac`。两轮独占的五样本运行使用相同播种数据,没有产品优化。下表时间单位为毫秒;源码期望值向上取整至实测各轮中位数以上,而不是施加尚未实现的优化目标。
+实现参考为 Apple M4 Pro、macOS arm64、Node 24.19.0 上的 PR #3537 合并结果。两轮独占的五样本运行使用相同播种数据,没有产品优化。下表时间单位为毫秒;源码期望值向上取整至实测各轮中位数以上,而不是施加尚未实现的优化目标。
 
 | 用例 | 第一轮原始总时间 | 第二轮原始总时间 | 中位数 | 历史 M4 期望 | 历史缩放预算 |
 |---|---|---|---|---:|---:|
@@ -45,13 +45,13 @@ SDK fixture 通过 profile patch 显式插入 `fs-local` 和 `str_replace_editor
 
 已发布 SDK 变体完成 100 个轮次、200 次请求和 800 次真实文件读取。五样本 smoke 总时间为 1,521.773、1,463.465、1,689.701、1,365.485 和 1,417.106 ms(中位数 1,463.465 ms);完整套件重复运行报告 1,596.183、1,784.536、2,120.082、1,405.365 和 1,355.894 ms(中位数 1,596.183 ms)。1,700 ms 参考期望对应 4,250 ms CI 预算。重复运行中未改变的服务用例也变慢,因此这是可变主机负载下的验证,不是放宽其独占校准预算的依据。SDK 进程使用白名单环境和私有主目录/工作区。40 秒截止时间启动 SDK 关闭;所有路径等待同一个记忆化 close Promise,并早于外层 worker 的 60 秒截止时间。Profile 时间包含启动、全部轮次和关闭,分别报告;不把父进程 CPU 或堆指标当作服务端内存。适配器不为外部模型服务商序列化请求。
 
-提交 `1dc3296eba631d51fbb3bb50e249bf3cc0fce9f6` 的首次 Linux x64 CI 测量使用 `VM-7-113-ubuntu-ci-10` 和 Node 24.18.1([run 34017868081,attempt 1,job 101444810498](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34017868081/job/101444810498))。SDK 中位数为 2,753.441 ms,预算为 4,250 ms;工具续聊保留堆中位数为 22.274 MiB,预算为 28.75 MiB。请求历史与工具续聊时间预算失败:分别为 785.498 ms 对 550 ms、1,077.285 ms 对 850 ms。未修改的 Session 重开 open 阶段也以 31.6 ms 对 30 ms 失败。[Attempt 2,job 101447076381](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34017868081/job/101447076381) 在同一提交和未修改预算下通过全部基准,但使用 `VM-7-113-ubuntu-ci-29` 和 Node 24.19.0。门禁运行器隐藏成功子进程的输出,因此该次运行只提供通过结论,不提供原始中位数。Runner 与 Node 版本同时变化,不能把差异仅归因于资源争用,也不能宣称已获得稳定的重复 CI 校准;这些证据不改变预算或共享比例。
+提交 `1dc3296eba631d51fbb3bb50e249bf3cc0fce9f6` 的首次 Linux x64 CI 测量使用 `VM-7-113-ubuntu-ci-10` 和 Node 24.18.1(run 34017868081,attempt 1,job 101444810498)。SDK 中位数为 2,753.441 ms,预算为 4,250 ms;工具续聊保留堆中位数为 22.274 MiB,预算为 28.75 MiB。请求历史与工具续聊时间预算失败:分别为 785.498 ms 对 550 ms、1,077.285 ms 对 850 ms。未修改的 Session 重开 open 阶段也以 31.6 ms 对 30 ms 失败。Attempt 2,job 101447076381 (run 34017868081) 在同一提交和未修改预算下通过全部基准,但使用 `VM-7-113-ubuntu-ci-29` 和 Node 24.19.0。门禁运行器隐藏成功子进程的输出,因此该次运行只提供通过结论,不提供原始中位数。Runner 与 Node 版本同时变化,不能把差异仅归因于资源争用,也不能宣称已获得稳定的重复 CI 校准;这些证据不改变预算或共享比例。
 
-目录用例使用显式的 900 ms CI 期望时间,仅乘现有 1.25× 余量,得到 1,125 ms,不再应用参考机器比例。标准双 CPU 托管 `ubuntu-24.04` 的 [run 34033336380,job 101487280801](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336380/job/101487280801) 报告未修改目录实现的五个总时间为 797.374、883.157、858.364、790.569 和 904.579 ms:中位数 858.364 ms 超出历史 800 ms 预算。上表 320 ms M4 期望保留为历史证据,不是 CI 测量。此方法与 Session 重开使用的显式 CI 校准一致(CI 期望为 50 ms);共享系数、负载、计时终点和产品实现均不改变。确定性对照与实测判定使用同一断言:未经舍入的录制中位数通过 1,125 ms 并被 800 ms 拒绝,合成的 1,400 ms 中位数则被 1,125 ms 拒绝。更快主机上的通过结果不能校准标准托管 runner。
+目录用例使用显式的 900 ms CI 期望时间,仅乘现有 1.25× 余量,得到 1,125 ms,不再应用参考机器比例。标准双 CPU 托管 `ubuntu-24.04` 的 run 34033336380,job 101487280801 报告未修改目录实现的五个总时间为 797.374、883.157、858.364、790.569 和 904.579 ms:中位数 858.364 ms 超出历史 800 ms 预算。上表 320 ms M4 期望保留为历史证据,不是 CI 测量。此方法与 Session 重开使用的显式 CI 校准一致(CI 期望为 50 ms);共享系数、负载、计时终点和产品实现均不改变。确定性对照与实测判定使用同一断言:未经舍入的录制中位数通过 1,125 ms 并被 800 ms 拒绝,合成的 1,400 ms 中位数则被 1,125 ms 拒绝。更快主机上的通过结果不能校准标准托管 runner。
 
-工具续聊同样使用 900 ms CI 期望时间与 1.25× 余量(1,125 ms)。未修改实现的 `79c052ab29` 在标准双 CPU 托管 [run 34034524265,job 101490056074](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34034524265/job/101490056074) 中报告总时间为 917.007、892.091、887.839、905.659 和 898.252 ms:中位数 898.252 ms 超出历史 850 ms 预算。340 ms M4 期望保留为历史证据。与实测判定相同的断言在 1,125 ms 下接受未经舍入的录制中位数,在 850 ms 下拒绝它,并拒绝合成的 1,400 ms 回退。负载、计时、产品代码和 28.75 MiB 保留堆预算均不改变。
+工具续聊同样使用 900 ms CI 期望时间与 1.25× 余量(1,125 ms)。未修改实现的 `79c052ab29` 在标准双 CPU 托管 run 34034524265,job 101490056074 中报告总时间为 917.007、892.091、887.839、905.659 和 898.252 ms:中位数 898.252 ms 超出历史 850 ms 预算。340 ms M4 期望保留为历史证据。与实测判定相同的断言在 1,125 ms 下接受未经舍入的录制中位数,在 850 ms 下拒绝它,并拒绝合成的 1,400 ms 回退。负载、计时、产品代码和 28.75 MiB 保留堆预算均不改变。
 
-基线请求历史使用 600 ms CI 期望时间与 1.25× 余量(750 ms)。未修改实现的 `54d1190a75` 在标准双 CPU 托管 [run 34035306987,job 101492163630](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34035306987/job/101492163630) 中报告总时间为 618.598、618.606、582.035、582.304 和 581.832 ms:中位数 582.304 ms 超出历史 550 ms 预算。220 ms M4 期望保留为历史证据。与实测判定相同的断言在 750 ms 下接受未经舍入的录制中位数,在 550 ms 下拒绝它,并拒绝合成的 900 ms 回退。此校准仅针对未优化基线;负载、计时、产品代码和内存预算均不改变。
+基线请求历史使用 600 ms CI 期望时间与 1.25× 余量(750 ms)。未修改实现的 `54d1190a75` 在标准双 CPU 托管 run 34035306987,job 101492163630 中报告总时间为 618.598、618.606、582.035、582.304 和 581.832 ms:中位数 582.304 ms 超出历史 550 ms 预算。220 ms M4 期望保留为历史证据。与实测判定相同的断言在 750 ms 下接受未经舍入的录制中位数,在 550 ms 下拒绝它,并拒绝合成的 900 ms 回退。此校准仅针对未优化基线;负载、计时、产品代码和内存预算均不改变。
 
 ## 考虑过的替代方案
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md
-2026-09-06-frontend-performance-budgets.md: 4d69dda8a04d4e9807c18c8f1b978ca7fc87a883
-2026-09-06-frontend-performance-budgets.zh.md: c9a174df55a2c175232eb2f4d8468b054e532b9c
+2026-09-06-frontend-performance-budgets.md: 349de386521f9ef5642cc808f7d2a9f849299ba4
+2026-09-06-frontend-performance-budgets.zh.md: 6231827472ece3992d994f9f021c0a34c0d398ec

+ 5 - 5
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md

@@ -22,7 +22,7 @@ Reconnect uses three fresh compiled plain-Node children. Each creates a 100,000-
 
 ## Calibration
 
-Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at product revision `925e012340`, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. The following historical reference table uses 8 ms replay pacing and includes a two-frame wait in first-reply timing. Standard-hosted open, paging, Trajectory, and reconnect expectations are recorded separately below; other source reference constants retain these allowances. The bounded-observer 261.60 ms paging median exceeds its 260 ms reference allowance but remains below its 650 ms CI limit; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The shared scale originates in Node CI calibration. Both actual x64 browser runs below pass the fixed budgets on unchanged benchmark code; this supplies repeated-run evidence for these runners, not a universal browser speed ratio.
+Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at the PR #3537 merge, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. The following historical reference table uses 8 ms replay pacing and includes a two-frame wait in first-reply timing. Standard-hosted open, paging, Trajectory, and reconnect expectations are recorded separately below; other source reference constants retain these allowances. The bounded-observer 261.60 ms paging median exceeds its 260 ms reference allowance but remains below its 650 ms CI limit; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The shared scale originates in Node CI calibration. Both actual x64 browser runs below pass the fixed budgets on unchanged benchmark code; this supplies repeated-run evidence for these runners, not a universal browser speed ratio.
 
 | Endpoint | Measured median | Reference allowance | Historical CI limit |
 |---|---:|---:|---:|
@@ -40,7 +40,7 @@ Draft typing spans 124.97–504.96 ms across the three isolated samples; the ref
 
 ### Actual CI runs
 
-[Run 34020120425, benchmark job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) passes the complete benchmark inventory at `6d1ba089e5052680961825c08aa4de19b4fe137a`. The runner is `VM-7-113-ubuntu-ci-19` in `dsh-selfhosted-ci`, using x64 Node 24.19.0 and Chromium 149.0.7827.55. [Attempt 2, benchmark job 101453296071](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101453296071) also passes the complete inventory at the same commit, on `VM-7-113-ubuntu-ci-25` with the same Node and Chromium versions. The following medians use three fresh samples per scenario in each run and leave the local reference table and source budgets unchanged.
+Run 34020120425, benchmark job 101451135853 passes the complete benchmark inventory at `6d1ba089e5052680961825c08aa4de19b4fe137a`. The runner is `VM-7-113-ubuntu-ci-19` in `dsh-selfhosted-ci`, using x64 Node 24.19.0 and Chromium 149.0.7827.55. Attempt 2, benchmark job 101453296071 (run 34020120425) also passes the complete inventory at the same commit, on `VM-7-113-ubuntu-ci-25` with the same Node and Chromium versions. The following medians use three fresh samples per scenario in each run and leave the local reference table and source budgets unchanged.
 
 | Endpoint | First CI median | Second CI median |
 |---|---:|---:|
@@ -58,13 +58,13 @@ All six browser samples report `inputOverlapped: true` and finish after the 241s
 
 ### Standard hosted expectations and input scheduling
 
-[Run 34033336246, job 101487216170](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336246/job/101487216170) on standard hosted Ubuntu with two CPUs records reconnect replacements of 46.574411, 46.067910, and 44.193704 ms, with 23.028 MiB retained heap. The endpoint-specific expectation is 50 ms; the existing 1.25× headroom gives a 63 ms integer ceiling. The 30 MiB memory budget and shared machine factor remain unchanged. Browser open records 681.276514 and 541.051233 ms before the third sample fails input overlap; both exceed the historical 500 ms limit. Repeated hosted open measurements below set its expectation and ceiling. Deterministic controls pass these recorded values and reject values above the new ceilings through the same assertions as the measured verdicts. Complete repeated hosted verdicts remain required; the two open values are not a three-sample median.
+Run 34033336246, job 101487216170 on standard hosted Ubuntu with two CPUs records reconnect replacements of 46.574411, 46.067910, and 44.193704 ms, with 23.028 MiB retained heap. The endpoint-specific expectation is 50 ms; the existing 1.25× headroom gives a 63 ms integer ceiling. The 30 MiB memory budget and shared machine factor remain unchanged. Browser open records 681.276514 and 541.051233 ms before the third sample fails input overlap; both exceed the historical 500 ms limit. Repeated hosted open measurements below set its expectation and ceiling. Deterministic controls pass these recorded values and reject values above the new ceilings through the same assertions as the measured verdicts. Complete repeated hosted verdicts remain required; the two open values are not a three-sample median.
 
 A local diagnostic with temporary 3× Chromium CPU throttling reproduces the overlap failure: the first marker becomes visible at 1321 ms, two animation frames finish at 1370 ms, and the composer click finishes at 1660 ms; the actual input is trusted but already sees DONE. Removing the frame wait and installing the witness before Send still leaves a run with first visibility at 1415 ms and click completion at 1726 ms, after the original 992 ms scripted stream. The fixed 16 ms cadence keeps the same 120 deltas and payload, providing 1984 ms of scripted pacing for this workload. Only that pacing term changes in the complete-wall allowance (4484 ms); input, first-reply, and main-thread overhead allowances remain unchanged. With the same diagnostic slowdown, three 16 ms samples reach first visibility at 1307/1479/1599 ms and accept trusted input before DONE; their post-DONE controls reject it. The diagnostic is not a CPU-ratio calibration. Each measured sample still requires trusted input while FIRST is present and DONE absent; a post-measurement trusted key after DONE must fail that same assertion. Host settlement and the 241st rendered turn-tail remain completion witnesses.
 
-[Run 34034524861, job 101490135303](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34034524861/job/101490135303) records three complete browser samples with trusted input overlap and passing post-DONE rejection controls. Slowest-page samples are 843.941625/672.834329/684.461818 ms (median 684.461818); first-Trajectory samples are 605.788061/367.754027/485.931656 ms (median 485.931656). Their endpoint-specific hosted expectations are 700 and 500 ms, with the same 1.25× headroom producing 875 and 625 ms limits. Recorded-median controls reject the historical 650/400 ms limits, accept these hosted limits, and reject one millisecond above each limit through the measured verdict's assertion. Open, first reply, main-thread task, input, and complete-wall medians are 713.910/1486.206/2806.415/947.398/2986.983 ms; only the open limit is recalibrated by the repeated measurements below. This run supplies calibration data, not a passing benchmark verdict; a complete hosted repeat remains required.
+Run 34034524861, job 101490135303 records three complete browser samples with trusted input overlap and passing post-DONE rejection controls. Slowest-page samples are 843.941625/672.834329/684.461818 ms (median 684.461818); first-Trajectory samples are 605.788061/367.754027/485.931656 ms (median 485.931656). Their endpoint-specific hosted expectations are 700 and 500 ms, with the same 1.25× headroom producing 875 and 625 ms limits. Recorded-median controls reject the historical 650/400 ms limits, accept these hosted limits, and reject one millisecond above each limit through the measured verdict's assertion. Open, first reply, main-thread task, input, and complete-wall medians are 713.910/1486.206/2806.415/947.398/2986.983 ms; only the open limit is recalibrated by the repeated measurements below. This run supplies calibration data, not a passing benchmark verdict; a complete hosted repeat remains required.
 
-[Run 34036109842, job 101494445658](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34036109842/job/101494445658) records open samples of 875.306861/1083.683529/814.700998 ms, with a median of 875.306861 ms versus 713.909727 ms in the preceding hosted run. The endpoint-specific expectation is 900 ms, rounding up the larger repeated median rather than adding an epsilon to the 875 ms limit; unchanged 1.25× headroom gives 1125 ms. The same enforced assertion accepts the recorded median, rejects it at both historical 500 and 875 ms limits, and rejects a synthetic 1126 ms value at the current limit. All three samples retain trusted input overlap and post-DONE rejection; every other frontend median remains within its unchanged limit. This calibration does not claim a green CI run.
+Run 34036109842, job 101494445658 records open samples of 875.306861/1083.683529/814.700998 ms, with a median of 875.306861 ms versus 713.909727 ms in the preceding hosted run. The endpoint-specific expectation is 900 ms, rounding up the larger repeated median rather than adding an epsilon to the 875 ms limit; unchanged 1.25× headroom gives 1125 ms. The same enforced assertion accepts the recorded median, rejects it at both historical 500 and 875 ms limits, and rejects a synthetic 1126 ms value at the current limit. All three samples retain trusted input overlap and post-DONE rejection; every other frontend median remains within its unchanged limit. This calibration does not claim a green CI run.
 
 A controlled mouse-refocus delay waits for the real DONE marker without pausing replay: the mouse path rejects a trusted input after DONE, while Enter submission and keyboard-only draft input pass all three samples under the same control. The delay is diagnostic-only. A clean three-sample run on arm64 Node 24.19.0 / Chromium 149.0.7827.55 reports first-reply/input/complete-wall medians of 288.823/418.868/2567.328 ms, with actual overlap and post-DONE rejection in every sample. This proves removal of the mouse-action scheduling dependency, not the cause of a particular hosted stall; all workload constants and budgets remain fixed.
 

+ 5 - 5
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.zh.md

@@ -22,7 +22,7 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 ## 校准
 
-在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和产品版本 `925e012340` 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。下列历史参考表使用 8 ms 重放节奏,首段回复计时包含两帧等待。标准托管打开、分页、Trajectory 和重连预期在下文单独记录;其他源码参考常量保留这些额度。受限观察器的分页中位数 261.60 ms 超过 260 ms 参考额度,但仍低于 650 ms CI 限制;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。共享倍率源自 Node CI 校准。下述两次实际 x64 浏览器运行在基准代码不变的情况下均通过固定预算;这提供这些 runner 的重复运行证据,而非普遍适用的浏览器速度比。
+在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和PR #3537 合并结果 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。下列历史参考表使用 8 ms 重放节奏,首段回复计时包含两帧等待。标准托管打开、分页、Trajectory 和重连预期在下文单独记录;其他源码参考常量保留这些额度。受限观察器的分页中位数 261.60 ms 超过 260 ms 参考额度,但仍低于 650 ms CI 限制;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。共享倍率源自 Node CI 校准。下述两次实际 x64 浏览器运行在基准代码不变的情况下均通过固定预算;这提供这些 runner 的重复运行证据,而非普遍适用的浏览器速度比。
 
 | 终点 | 实测中位数 | 参考额度 | 历史 CI 限制 |
 |---|---:|---:|---:|
@@ -40,7 +40,7 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 ### 实际 CI 运行
 
-[运行 34020120425,基准 job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) 在 `6d1ba089e5052680961825c08aa4de19b4fe137a` 上通过完整基准清单。runner 为 `dsh-selfhosted-ci` 中的 `VM-7-113-ubuntu-ci-19`,使用 x64 Node 24.19.0 和 Chromium 149.0.7827.55。[第 2 次执行,基准 job 101453296071](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101453296071) 在相同 commit 上也通过完整清单,runner 为 `VM-7-113-ubuntu-ci-25`,Node 和 Chromium 版本相同。下列中位数来自每次运行中每个场景的三个全新样本,本地参考表和源码预算保持不变。
+运行 34020120425,基准 job 101451135853 在 `6d1ba089e5052680961825c08aa4de19b4fe137a` 上通过完整基准清单。runner 为 `dsh-selfhosted-ci` 中的 `VM-7-113-ubuntu-ci-19`,使用 x64 Node 24.19.0 和 Chromium 149.0.7827.55。第 2 次执行,基准 job 101453296071 (run 34020120425) 在相同 commit 上也通过完整清单,runner 为 `VM-7-113-ubuntu-ci-25`,Node 和 Chromium 版本相同。下列中位数来自每次运行中每个场景的三个全新样本,本地参考表和源码预算保持不变。
 
 | 终点 | 首次 CI 中位数 | 第二次 CI 中位数 |
 |---|---:|---:|
@@ -58,13 +58,13 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 ### 标准托管预期与输入调度
 
-[运行 34033336246,job 101487216170](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336246/job/101487216170) 在双 CPU 标准托管 Ubuntu 上记录重连替换时间 46.574411、46.067910 和 44.193704 ms,保留 heap 为 23.028 MiB。该终点的预期为 50 ms;现有 1.25× 余量产生向上取整后的 63 ms 上限。30 MiB 内存预算及共享机器倍率不变。浏览器打开记录 681.276514 和 541.051233 ms,第三个样本因输入重叠失败而中止;两个值均超过历史 500 ms 上限。下文的托管打开重复测量决定其预期与上限。确定性对照通过这些记录值,并使用与测量判定相同的断言拒绝超过新上限的值。仍需完整的托管重复运行判定;这两个打开值不是三样本中位数。
+运行 34033336246,job 101487216170 在双 CPU 标准托管 Ubuntu 上记录重连替换时间 46.574411、46.067910 和 44.193704 ms,保留 heap 为 23.028 MiB。该终点的预期为 50 ms;现有 1.25× 余量产生向上取整后的 63 ms 上限。30 MiB 内存预算及共享机器倍率不变。浏览器打开记录 681.276514 和 541.051233 ms,第三个样本因输入重叠失败而中止;两个值均超过历史 500 ms 上限。下文的托管打开重复测量决定其预期与上限。确定性对照通过这些记录值,并使用与测量判定相同的断言拒绝超过新上限的值。仍需完整的托管重复运行判定;这两个打开值不是三样本中位数。
 
 临时使用 3× Chromium CPU 降速的本地诊断复现重叠失败:首个标记在 1321 ms 可见,两次动画帧在 1370 ms 结束,输入框点击在 1660 ms 完成;实际输入是真实事件,但已看到 DONE。移除帧等待并在发送前安装观察器后,一次运行仍在 1415 ms 才看到首个标记,点击在 1726 ms 完成,晚于原先 992 ms 的脚本流。固定 16 ms 节奏保留相同的 120 个 delta 和负载,为该工作负载提供 1984 ms 脚本节奏。完整壁钟额度仅改变该节奏项(4484 ms);输入、首段回复及主线程额外开销额度不变。在相同诊断降速下,三个 16 ms 样本在 1307/1479/1599 ms 达到首段可见状态,并接受 DONE 之前的真实输入;其 DONE 之后的对照拒绝该输入。该诊断不是 CPU 比率校准。每个测量样本仍要求真实输入发生时 FIRST 存在且 DONE 不存在;测量后在 DONE 之后发送的真实按键必须无法通过同一个断言。Host 结算和第 241 个已渲染 turn-tail 仍是完成证据。
 
-[运行 34034524861,job 101490135303](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34034524861/job/101490135303) 记录三个完整浏览器样本,均具有真实输入重叠,并通过 DONE 之后的拒绝对照。最慢分页样本为 843.941625/672.834329/684.461818 ms(中位数 684.461818);首次 Trajectory 样本为 605.788061/367.754027/485.931656 ms(中位数 485.931656)。两者的终点专属托管预期分别为 700 和 500 ms,相同的 1.25× 余量产生 875 和 625 ms 上限。记录中位数对照拒绝历史 650/400 ms 上限,接受这些托管上限,并通过测量判定所用断言拒绝超过各上限一毫秒的值。打开、首段回复、主线程任务、输入及完整壁钟的中位数为 713.910/1486.206/2806.415/947.398/2986.983 ms;仅打开上限根据下文的重复测量重新校准。该运行提供校准数据,不代表基准判定通过;仍需完整的托管重复运行。
+运行 34034524861,job 101490135303 记录三个完整浏览器样本,均具有真实输入重叠,并通过 DONE 之后的拒绝对照。最慢分页样本为 843.941625/672.834329/684.461818 ms(中位数 684.461818);首次 Trajectory 样本为 605.788061/367.754027/485.931656 ms(中位数 485.931656)。两者的终点专属托管预期分别为 700 和 500 ms,相同的 1.25× 余量产生 875 和 625 ms 上限。记录中位数对照拒绝历史 650/400 ms 上限,接受这些托管上限,并通过测量判定所用断言拒绝超过各上限一毫秒的值。打开、首段回复、主线程任务、输入及完整壁钟的中位数为 713.910/1486.206/2806.415/947.398/2986.983 ms;仅打开上限根据下文的重复测量重新校准。该运行提供校准数据,不代表基准判定通过;仍需完整的托管重复运行。
 
-[运行 34036109842,job 101494445658](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34036109842/job/101494445658) 记录打开样本 875.306861/1083.683529/814.700998 ms,中位数为 875.306861 ms,前一次托管运行的中位数为 713.909727 ms。该终点的预期为 900 ms,向上取整较大的重复测量中位数,而非向 875 ms 上限增加微量余量;不变的 1.25× 余量产生 1125 ms 上限。同一个强制断言接受记录中位数,在历史 500 和 875 ms 上限下均拒绝它,并在当前上限下拒绝合成的 1126 ms 值。三个样本均保留真实输入重叠与 DONE 之后的拒绝;其他所有前端中位数均在不变的上限内。此校准不代表 CI 运行通过。
+运行 34036109842,job 101494445658 记录打开样本 875.306861/1083.683529/814.700998 ms,中位数为 875.306861 ms,前一次托管运行的中位数为 713.909727 ms。该终点的预期为 900 ms,向上取整较大的重复测量中位数,而非向 875 ms 上限增加微量余量;不变的 1.25× 余量产生 1125 ms 上限。同一个强制断言接受记录中位数,在历史 500 和 875 ms 上限下均拒绝它,并在当前上限下拒绝合成的 1126 ms 值。三个样本均保留真实输入重叠与 DONE 之后的拒绝;其他所有前端中位数均在不变的上限内。此校准不代表 CI 运行通过。
 
 受控的鼠标重新聚焦延迟等待真实 DONE 标记,不暂停重放:鼠标路径拒绝 DONE 之后的真实输入,而 Enter 提交与纯键盘草稿输入在相同对照下通过全部三个样本。该延迟仅用于诊断。在 arm64 Node 24.19.0 / Chromium 149.0.7827.55 上,不含延迟的三个样本报告首段回复/输入/完整壁钟中位数 288.823/418.868/2567.328 ms,每个样本均满足实际重叠并拒绝 DONE 之后的输入。这证明移除了鼠标操作调度依赖,并不证明某次托管停顿的原因;全部工作负载常量和预算保持固定。
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-07-pwsh-ci-observable-completion.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-07-pwsh-ci-observable-completion.md
-2026-09-07-pwsh-ci-observable-completion.md: ed00ea3f20f24cd152240314f03ee83657eb273d
-2026-09-07-pwsh-ci-observable-completion.zh.md: 8c76d165466821913b17de92c6ac0b3ee9bc06d8
+2026-09-07-pwsh-ci-observable-completion.md: 2cd42296c14602f1f4aad68b8988f84dabbd5479
+2026-09-07-pwsh-ci-observable-completion.zh.md: 9ea83963be7faa274e53b1b9260222f71553e0f4

+ 2 - 2
.agents/notes/implemented/testing/2026-09-07-pwsh-ci-observable-completion.md

@@ -6,9 +6,9 @@ English | [中文](2026-09-07-pwsh-ci-observable-completion.zh.md)
 
 ## Problem
 
-The [hosted coverage job](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033367752/job/101605386802) rejects a persistent PowerShell send because it returns `inferred_idle` rather than `stdin_read`. Output silence is a supported bounded inference, not proof that a command finished. The real-shell test also searches output for text present in the echoed command, which cannot independently prove execution.
+The hosted coverage job (run 34033367752, job 101605386802) rejects a persistent PowerShell send because it returns `inferred_idle` rather than `stdin_read`. Output silence is a supported bounded inference, not proof that a command finished. The real-shell test also searches output for text present in the echoed command, which cannot independently prove execution.
 
-The [snapshot job](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033367752/job/101605386868) rejects both PowerShell scenarios despite successful `PWSH_OK` output. Their fixtures omit the headless profile’s policy events and runtime-context message; their prompt and tool-schema pins also describe an older, smaller composition. Hosts without PowerShell skip these cases and cannot detect that drift.
+The snapshot job (run 34033367752, job 101605386868) rejects both PowerShell scenarios despite successful `PWSH_OK` output. Their fixtures omit the headless profile’s policy events and runtime-context message; their prompt and tool-schema pins also describe an older, smaller composition. Hosts without PowerShell skip these cases and cannot detect that drift.
 
 ## Decision
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-07-pwsh-ci-observable-completion.zh.md

@@ -6,9 +6,9 @@ Status: implemented
 
 ## Problem
 
-[托管 coverage 作业](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033367752/job/101605386802) 因持久 PowerShell send 返回 `inferred_idle` 而非 `stdin_read` 判定失败。输出静默是受支持的有界推断,不是命令完成的证明。真实 shell 测试还在输出中查找被回显命令本身包含的文本,无法独立证明命令执行。
+托管 coverage 作业 (run 34033367752, job 101605386802) 因持久 PowerShell send 返回 `inferred_idle` 而非 `stdin_read` 判定失败。输出静默是受支持的有界推断,不是命令完成的证明。真实 shell 测试还在输出中查找被回显命令本身包含的文本,无法独立证明命令执行。
 
-[快照作业](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033367752/job/101605386868) 在成功输出 `PWSH_OK` 后仍拒绝两个 PowerShell 场景。其 fixture 缺少 headless profile 的策略事件与运行时上下文消息;prompt 和工具 schema pin 也描述了更早、更小的组合。没有 PowerShell 的主机会跳过这些用例,无法发现此类漂移。
+快照作业 (run 34033367752, job 101605386868) 在成功输出 `PWSH_OK` 后仍拒绝两个 PowerShell 场景。其 fixture 缺少 headless profile 的策略事件与运行时上下文消息;prompt 和工具 schema pin 也描述了更早、更小的组合。没有 PowerShell 的主机会跳过这些用例,无法发现此类漂移。
 
 ## Decision
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-07-subagent-teardown-test-budgets.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-07-subagent-teardown-test-budgets.md
-2026-09-07-subagent-teardown-test-budgets.md: 4fe83c421383aa768ffa0d33520407ed8d099d14
-2026-09-07-subagent-teardown-test-budgets.zh.md: 1487351d498c9d0ef9eb6c2e83e47425ac82b9c9
+2026-09-07-subagent-teardown-test-budgets.md: 372c8b311054b11d86b3057415dd98584d94be9f
+2026-09-07-subagent-teardown-test-budgets.zh.md: 8539664280c5fe720f27459fdb2d14c47ad3e780

+ 1 - 1
.agents/notes/implemented/testing/2026-09-07-subagent-teardown-test-budgets.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-07-subagent-teardown-test-budgets.zh.md)
 
 ## Problem
 
-The [Windows coverage run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34085536250/job/101628739668) reports two teardown failures despite granting tests and hooks 90 seconds. The ACP ignored-EOF test races disposal against its own five-second timer. The real Codex test overrides the hook budget with 30 seconds. Neither deadline tests a product latency guarantee. The Codex body has already observed process-tree exit before its hook fails; the log does not identify whether context disposal, HTTP closure, or temporary-directory removal exceeded the hook budget.
+The Windows coverage run (run 34085536250, job 101628739668) reports two teardown failures despite granting tests and hooks 90 seconds. The ACP ignored-EOF test races disposal against its own five-second timer. The real Codex test overrides the hook budget with 30 seconds. Neither deadline tests a product latency guarantee. The Codex body has already observed process-tree exit before its hook fails; the log does not identify whether context disposal, HTTP closure, or temporary-directory removal exceeded the hook budget.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/testing/2026-09-07-subagent-teardown-test-budgets.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-[Windows 覆盖率运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34085536250/job/101628739668) 为测试和钩子提供 90 秒预算,却报告了两个清理失败。ACP 忽略 EOF 测试让清理与自设的五秒定时器竞争。真实 Codex 测试将钩子预算覆盖为 30 秒。这两个期限都不用于验证产品延迟保证。Codex 测试正文在钩子失败前已经观察到进程树退出;日志未指出究竟是上下文释放、HTTP 关闭还是临时目录删除超出了钩子预算。
+Windows 覆盖率运行 (run 34085536250, job 101628739668) 为测试和钩子提供 90 秒预算,却报告了两个清理失败。ACP 忽略 EOF 测试让清理与自设的五秒定时器竞争。真实 Codex 测试将钩子预算覆盖为 30 秒。这两个期限都不用于验证产品延迟保证。Codex 测试正文在钩子失败前已经观察到进程树退出;日志未指出究竟是上下文释放、HTTP 关闭还是临时目录删除超出了钩子预算。
 
 ## 决策
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-08-ci-completion-observations.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-08-ci-completion-observations.md
-2026-09-08-ci-completion-observations.md: 368d3b2f644ca3f7f850dd0ffbd56db19a1dca64
-2026-09-08-ci-completion-observations.zh.md: 1529361fb07b661f7152da9566900e3782cdb67f
+2026-09-08-ci-completion-observations.md: 30145f26fccd1ca7d174d76df738f100827cc32a
+2026-09-08-ci-completion-observations.zh.md: a819d6ad45d80b7fce2f67025629d8d6cf9f8344

+ 1 - 1
.agents/notes/implemented/testing/2026-09-08-ci-completion-observations.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-08-ci-completion-observations.zh.md)
 
 ## Problem
 
-The [reference CI run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34206953049) reports a webhook-created Session absent after a one-second poll and empty PowerShell output before a five-second read deadline. HTTP acceptance, projected UI state, process startup, and durable completion are separate observations. Tests need an explicit completion condition and controls that prevent an intermediate state from satisfying it. The [completion-wait decision](2026-09-08-ci-readiness-and-completion.md) owns those conditions and lane budgets; these fixtures make their ordering and cleanup observable under controlled delays.
+The reference CI run (run 34206953049) reports a webhook-created Session absent after a one-second poll and empty PowerShell output before a five-second read deadline. HTTP acceptance, projected UI state, process startup, and durable completion are separate observations. Tests need an explicit completion condition and controls that prevent an intermediate state from satisfying it. The [completion-wait decision](2026-09-08-ci-readiness-and-completion.md) owns those conditions and lane budgets; these fixtures make their ordering and cleanup observable under controlled delays.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/testing/2026-09-08-ci-completion-observations.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-[参考 CI 运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34206953049)报告:轮询一秒后 webhook 创建的 Session 仍不存在,五秒读取期限内 PowerShell 输出为空。HTTP 接受、UI 投影状态、进程启动和持久化完成是不同的观察。测试需要明确的完成条件,并用对照阻止中间状态满足该条件。[完成等待决策](2026-09-08-ci-readiness-and-completion.zh.md)拥有这些条件与 lane 预算;这些 fixture 通过受控延迟使顺序与清理可观察。
+参考 CI 运行 (run 34206953049)报告:轮询一秒后 webhook 创建的 Session 仍不存在,五秒读取期限内 PowerShell 输出为空。HTTP 接受、UI 投影状态、进程启动和持久化完成是不同的观察。测试需要明确的完成条件,并用对照阻止中间状态满足该条件。[完成等待决策](2026-09-08-ci-readiness-and-completion.zh.md)拥有这些条件与 lane 预算;这些 fixture 通过受控延迟使顺序与清理可观察。
 
 ## 决策
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-08-ci-readiness-and-completion.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-08-ci-readiness-and-completion.md
-2026-09-08-ci-readiness-and-completion.md: 9afc8f7ecc3036acdf9ec2d5c87a50ef74ac9b69
-2026-09-08-ci-readiness-and-completion.zh.md: f3d72276a9160dd040e2b332d8579df98afec8fc
+2026-09-08-ci-readiness-and-completion.md: 1bf4f701f1b66debd982f1b786116e437e35dea4
+2026-09-08-ci-readiness-and-completion.zh.md: ae6d4e0195c66435fba22efae2376f6979ea9f6d

+ 5 - 5
.agents/notes/implemented/testing/2026-09-08-ci-readiness-and-completion.md

@@ -6,15 +6,15 @@ English | [中文](2026-09-08-ci-readiness-and-completion.zh.md)
 
 ## Problem
 
-The [empty master PR run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34206953049) fails while waiting one second for webhook Session creation and five seconds for PowerShell output. Neither test measures a startup latency guarantee. A [separate run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34207864157) shows the same short-budget problem in a desktop worker readiness test and captures a feedback acknowledgement while the composer still holds the submitted command.
+The empty master PR run (run 34206953049) fails while waiting one second for webhook Session creation and five seconds for PowerShell output. Neither test measures a startup latency guarantee. A separate run (run 34207864157) shows the same short-budget problem in a desktop worker readiness test and captures a feedback acknowledgement while the composer still holds the submitted command.
 
-Another [Windows coverage run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34224004885/job/102053583437) reports a null publint child status and an LSP initialization-marker timeout. Their helpers impose five- and three-second limits inside the lane's 90-second test budget. These cases verify publication contents and cancellation behavior rather than cold-start latency.
+Another Windows coverage run (run 34224004885, job 102053583437) reports a null publint child status and an LSP initialization-marker timeout. Their helpers impose five- and three-second limits inside the lane's 90-second test budget. These cases verify publication contents and cancellation behavior rather than cold-start latency.
 
-The [ACP coverage run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34242280527/job/102115221228) exhausts a one-second registry poll after transport failure. Disconnect cleanup includes cancellation, output draining, persistence, and owner disposal; registry removal alone does not establish complete teardown.
+The ACP coverage run (run 34242280527, job 102115221228) exhausts a one-second registry poll after transport failure. Disconnect cleanup includes cancellation, output draining, persistence, and owner disposal; registry removal alone does not establish complete teardown.
 
-A [worker-runtime coverage failure](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34248221544/job/102135631932) exhausts the slow-binding fixture's one-second compute allowance. Concurrent native Windows reproductions exceed that allowance before calling the binding. Worker initialization contributes measured active time; the delayed binding contributes idle time.
+A worker-runtime coverage failure (run 34248221544, job 102135631932) exhausts the slow-binding fixture's one-second compute allowance. Concurrent native Windows reproductions exceed that allowance before calling the binding. Worker initialization contributes measured active time; the delayed binding contributes idle time.
 
-The [Windows coverage run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34324325375/job/102377982193) reports an SDK subprocess exit beyond a fixture's 200 ms confirmation window and an Inspector Worker startup beyond its ten-second default. Neither the protocol-error routing case nor the Cordis tree projection case measures those latency guarantees.
+The Windows coverage run (run 34324325375, job 102377982193) reports an SDK subprocess exit beyond a fixture's 200 ms confirmation window and an Inspector Worker startup beyond its ten-second default. Neither the protocol-error routing case nor the Cordis tree projection case measures those latency guarantees.
 
 ## Decision
 

+ 5 - 5
.agents/notes/implemented/testing/2026-09-08-ci-readiness-and-completion.zh.md

@@ -6,15 +6,15 @@ Status: implemented
 
 ## 问题
 
-[master 空 PR 的运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34206953049)在等待 Webhook Session 创建一秒、等待 PowerShell 输出五秒时失败。两个测试都不衡量启动延迟保证。[另一次运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34207864157)在 Desktop worker 就绪测试中暴露了相同的局部短时限问题,并在输入框仍保留已提交命令时截取了反馈确认。
+master 空 PR 的运行 (run 34206953049)在等待 Webhook Session 创建一秒、等待 PowerShell 输出五秒时失败。两个测试都不衡量启动延迟保证。另一次运行 (run 34207864157)在 Desktop worker 就绪测试中暴露了相同的局部短时限问题,并在输入框仍保留已提交命令时截取了反馈确认。
 
-另一次 [Windows coverage 运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34224004885/job/102053583437)报告了 publint 子进程退出状态为 null,以及 LSP 初始化标记等待超时。对应 helper 在通道的 90 秒测试预算内另设五秒和三秒限制。这些用例验证发布内容与取消行为,不衡量冷启动延迟。
+另一次 Windows coverage 运行 (run 34224004885, job 102053583437)报告了 publint 子进程退出状态为 null,以及 LSP 初始化标记等待超时。对应 helper 在通道的 90 秒测试预算内另设五秒和三秒限制。这些用例验证发布内容与取消行为,不衡量冷启动延迟。
 
-[ACP coverage 运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34242280527/job/102115221228)在传输失败后耗尽一秒的注册表轮询期限。断连清理包含取消、输出排空、持久化和 owner 处置;仅从注册表移除不能证明完整拆卸已经结束。
+ACP coverage 运行 (run 34242280527, job 102115221228)在传输失败后耗尽一秒的注册表轮询期限。断连清理包含取消、输出排空、持久化和 owner 处置;仅从注册表移除不能证明完整拆卸已经结束。
 
-一次 [worker runtime coverage 失败](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34248221544/job/102135631932)耗尽了慢 binding 夹具的一秒计算额度。原生 Windows 并发复现在调用 binding 前已超过该额度。Worker 初始化会累计所测的活跃时间;延迟的 binding 累计空闲时间。
+一次 worker runtime coverage 失败 (run 34248221544, job 102135631932)耗尽了慢 binding 夹具的一秒计算额度。原生 Windows 并发复现在调用 binding 前已超过该额度。Worker 初始化会累计所测的活跃时间;延迟的 binding 累计空闲时间。
 
-[Windows 覆盖率运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34324325375/job/102377982193)报告了 SDK 子进程退出超过测试设置的 200 毫秒确认期限,以及 Inspector Worker 启动超过十秒默认期限。协议错误转发用例和 Cordis 树投影用例都不衡量这些延迟保证。
+Windows 覆盖率运行 (run 34324325375, job 102377982193)报告了 SDK 子进程退出超过测试设置的 200 毫秒确认期限,以及 Inspector Worker 启动超过十秒默认期限。协议错误转发用例和 Cordis 树投影用例都不衡量这些延迟保证。
 
 ## 决策
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-09-user-patch-hmr-test-delivery.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-09-user-patch-hmr-test-delivery.md
-2026-09-09-user-patch-hmr-test-delivery.md: fadd191650411f07c3b8b1f35c94e73869fbc8e2
-2026-09-09-user-patch-hmr-test-delivery.zh.md: b108cf25766ff108d87c429e9f208c2d4cf7afd2
+2026-09-09-user-patch-hmr-test-delivery.md: d31a0feda5aa8c06237a77d9f1df29f9519a4b28
+2026-09-09-user-patch-hmr-test-delivery.zh.md: e959a625699600ac371df91f2f48e0092b7f6f5a

+ 1 - 1
.agents/notes/implemented/testing/2026-09-09-user-patch-hmr-test-delivery.md

@@ -6,7 +6,7 @@ English | [中文](2026-09-09-user-patch-hmr-test-delivery.zh.md)
 
 ## Problem
 
-The [macOS Sandbox run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34238200206/job/102101292119) times out while waiting for the first user-patch addition. Concurrent local reproductions show no filesystem notification reaching HMR. A polling variant also misses a subsequent edit while HMR has no pending refresh. These failures prevent the refresh assertions from exercising the parser, activation, and recovery behavior they own.
+The macOS Sandbox run (run 34238200206, job 102101292119) times out while waiting for the first user-patch addition. Concurrent local reproductions show no filesystem notification reaching HMR. A polling variant also misses a subsequent edit while HMR has no pending refresh. These failures prevent the refresh assertions from exercising the parser, activation, and recovery behavior they own.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/testing/2026-09-09-user-patch-hmr-test-delivery.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-[macOS Sandbox 运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34238200206/job/102101292119) 在等待首次用户 patch 新增时超时。本地并发复现表明,没有文件系统通知到达 HMR。轮询变体也会遗漏后续修改,此时 HMR 没有待执行的刷新。这些失败阻止刷新断言执行其负责验证的解析、激活与回滚行为。
+macOS Sandbox 运行 (run 34238200206, job 102101292119) 在等待首次用户 patch 新增时超时。本地并发复现表明,没有文件系统通知到达 HMR。轮询变体也会遗漏后续修改,此时 HMR 没有待执行的刷新。这些失败阻止刷新断言执行其负责验证的解析、激活与回滚行为。
 
 ## 决策
 

+ 10 - 10
THIRD_PARTY_NOTICES.md

@@ -13,17 +13,17 @@ The complete npm transitive closure, including the Landlock launcher workspace,
 
 The Cordis framework and its foundation libraries are source-vendored into this repository rather than consumed from npm, and republished under the `@deepseek-ai` scope. All are MIT-licensed; each directory preserves its upstream `LICENSE` file. Exact upstream commits and local modifications are recorded in [`vendor/README.md`](vendor/README.md).
 
-| Package | Upstream name | Upstream | License |
+| Package | Upstream name | Source | License |
 | --- | --- | --- | --- |
-| `@deepseek-ai/cosmokit` | `cosmokit` | [github.com/deepseek-harness/cosmokit](https://github.com/deepseek-harness/cosmokit) | MIT |
-| `@deepseek-ai/schemastery` | `schemastery` | [github.com/deepseek-harness/schemastery](https://github.com/deepseek-harness/schemastery) | MIT |
-| `@deepseek-ai/cordis` | `cordis` | [github.com/cordiverse/cordis](https://github.com/cordiverse/cordis) | MIT |
-| `@deepseek-ai/cordis-plugin-loader` | `@cordisjs/plugin-loader` | [github.com/cordiverse/cordis](https://github.com/cordiverse/cordis) | MIT |
-| `@deepseek-ai/cordis-plugin-include` | `@cordisjs/plugin-include` | [github.com/deepseek-harness/cordis](https://github.com/deepseek-harness/cordis) | MIT |
-| `@deepseek-ai/cordis-plugin-group` | `@cordisjs/plugin-group` | [github.com/deepseek-harness/cordis](https://github.com/deepseek-harness/cordis) | MIT |
-| `@deepseek-ai/cordis-plugin-timer` | `@cordisjs/plugin-timer` | [github.com/deepseek-harness/cordis](https://github.com/deepseek-harness/cordis) | MIT |
-| `@deepseek-ai/cordis-plugin-hmr` | `@cordisjs/plugin-hmr` | [github.com/deepseek-harness/cordis](https://github.com/deepseek-harness/cordis) | MIT |
-| `@deepseek-ai/cordis-plugin-logger-console` | `@cordisjs/plugin-logger-console` | [github.com/deepseek-harness/cordis](https://github.com/deepseek-harness/cordis) | MIT |
+| `@deepseek-ai/cosmokit` | `cosmokit` | [vendor/cosmokit](vendor/cosmokit/) | MIT |
+| `@deepseek-ai/schemastery` | `schemastery` | [vendor/schemastery](vendor/schemastery/) | MIT |
+| `@deepseek-ai/cordis` | `cordis` | [vendor/cordis](vendor/cordis/) | MIT |
+| `@deepseek-ai/cordis-plugin-loader` | `@cordisjs/plugin-loader` | [vendor/loader](vendor/loader/) | MIT |
+| `@deepseek-ai/cordis-plugin-include` | `@cordisjs/plugin-include` | [vendor/include](vendor/include/) | MIT |
+| `@deepseek-ai/cordis-plugin-group` | `@cordisjs/plugin-group` | [vendor/group](vendor/group/) | MIT |
+| `@deepseek-ai/cordis-plugin-timer` | `@cordisjs/plugin-timer` | [vendor/timer](vendor/timer/) | MIT |
+| `@deepseek-ai/cordis-plugin-hmr` | `@cordisjs/plugin-hmr` | [vendor/hmr](vendor/hmr/) | MIT |
+| `@deepseek-ai/cordis-plugin-logger-console` | `@cordisjs/plugin-logger-console` | [vendor/logger-console](vendor/logger-console/) | MIT |
 
 ## Runtime npm dependencies
 

+ 1 - 1
apps/cli/config/examples/github-review/cordis.yml

@@ -9,7 +9,7 @@
       name: './github-ready-review-rule.mjs'
       config:
         source: primary-github
-        repository: deepseek-harness/deepseek-harness
+        repository: deepseek-ai/deepseek-harness
         workspacePath: !!js process.env.DSH_GITHUB_REVIEW_WORKSPACE ?? process.cwd()
         agentPreset: standard
         permissionPreset: read-only

+ 1 - 1
apps/cli/tests/fixtures/github-webhook/cordis.yml

@@ -9,7 +9,7 @@
       name: './github-webhook-rule.mjs'
       config:
         source: github-real-e2e
-        repository: deepseek-harness/deepseek-harness
+        repository: deepseek-ai/deepseek-harness
         workspacePath: !!js process.env.DSH_GITHUB_E2E_WORKSPACE
         marker: !!js process.env.DSH_GITHUB_E2E_MARKER
         agentPreset: minimal

+ 2 - 2
apps/cli/tests/github-webhook-real.e2e.ts

@@ -327,10 +327,10 @@ async function sendGitHubDelivery(origin: string): Promise<Response> {
   const body = JSON.stringify({
     action: 'ready_for_review',
     number: 4242,
-    repository: { full_name: 'deepseek-harness/deepseek-harness' },
+    repository: { full_name: 'deepseek-ai/deepseek-harness' },
     pull_request: {
       title: 'Real CLI webhook e2e',
-      html_url: 'https://github.com/deepseek-harness/deepseek-harness/pull/4242',
+      html_url: 'https://github.com/deepseek-ai/deepseek-harness/pull/4242',
       draft: false,
       user: { login: 'octocat' },
       base: { ref: 'master', sha: 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa' },

+ 4 - 4
apps/web/tests/expected/github-ready-review/conversation-expanded.expected.md

@@ -2,16 +2,16 @@
   - treeitem "{{workspace}}" [expanded]:
     - img
     - text: {{workspace}}
-  - treeitem "Review deepseek-harness/deepseek-harness#314 Session actions for Review deepseek-harness/deepseek-harness#314" [selected]:
-    - text: Review deepseek-harness/deepseek-harness#314
-    - button "Session actions for Review deepseek-harness/deepseek-harness#314":
+  - treeitem "Review deepseek-ai/deepseek-harness#314 Session actions for Review deepseek-ai/deepseek-harness#314" [selected]:
+    - text: Review deepseek-ai/deepseek-harness#314
+    - button "Session actions for Review deepseek-ai/deepseek-harness#314":
       - img
 
 ---
 
 - banner:
   - navigation "Session hierarchy":
-    - button "Review deepseek-harness/deepseek-harness#314" [disabled]
+    - button "Review deepseek-ai/deepseek-harness#314" [disabled]
   - img
   - text: Standard mode
   - button "More actions":

+ 4 - 4
apps/web/tests/expected/github-ready-review/conversation.expected.md

@@ -2,16 +2,16 @@
   - treeitem "{{workspace}}" [expanded]:
     - img
     - text: {{workspace}}
-  - treeitem "Review deepseek-harness/deepseek-harness#314 Session actions for Review deepseek-harness/deepseek-harness#314" [selected]:
-    - text: Review deepseek-harness/deepseek-harness#314
-    - button "Session actions for Review deepseek-harness/deepseek-harness#314":
+  - treeitem "Review deepseek-ai/deepseek-harness#314 Session actions for Review deepseek-ai/deepseek-harness#314" [selected]:
+    - text: Review deepseek-ai/deepseek-harness#314
+    - button "Session actions for Review deepseek-ai/deepseek-harness#314":
       - img
 
 ---
 
 - banner:
   - navigation "Session hierarchy":
-    - button "Review deepseek-harness/deepseek-harness#314" [disabled]
+    - button "Review deepseek-ai/deepseek-harness#314" [disabled]
   - img
   - text: Standard mode
   - button "More actions":

+ 3 - 3
apps/web/tests/github-ready-review.e2e.ts

@@ -30,7 +30,7 @@ const EXPANDED_EXPECTED = fileURLToPath(
 const PROVIDER = 'github-webhook-review-test'
 const MODEL = 'reply'
 const SECRET = 'github-webhook-review-secret'
-const TITLE = 'Review deepseek-harness/deepseek-harness#314'
+const TITLE = 'Review deepseek-ai/deepseek-harness#314'
 const REPLY = 'Review complete: no actionable findings.'
 
 /** Deterministic model response for the webhook-created Session. */
@@ -129,10 +129,10 @@ describe.skipIf(MODE === 'record')('web e2e: GitHub ready-for-review', () => {
     const payload = {
       action: 'ready_for_review',
       number: 314,
-      repository: { full_name: 'deepseek-harness/deepseek-harness' },
+      repository: { full_name: 'deepseek-ai/deepseek-harness' },
       pull_request: {
         title: 'Fix session replay',
-        html_url: 'https://github.com/deepseek-harness/deepseek-harness/pull/314',
+        html_url: 'https://github.com/deepseek-ai/deepseek-harness/pull/314',
         draft: false,
         user: { login: 'octocat' },
         base: { ref: 'master', sha: 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa' },

+ 2 - 2
docs/AGENTS.md

@@ -71,6 +71,6 @@ Hunt these in any doc; [dsh-doc](../.agents/skills/dsh-doc/SKILL.md) runs this l
 - Emphasis inflation: bold, CAPS, or "critically" everywhere means nothing stands out. Reserve emphasis for the clause that changes behavior.
 - Spec-speak in `implemented/` Agent Notes: "should", migration plans, acceptance checklists. An implemented Agent Note describes what is, per the [implemented-note instructions](../.agents/notes/implemented/AGENTS.md).
 
-## Cross-reference with machine-checkable links, never free prose
+## Repository references
 
-Link repository references with relative Markdown paths, never bare filenames or Agent Note numbers. `verify-md-links` rejects missing targets and dead `#fragment` anchors.
+Use relative Markdown links for current files and tags or PR numbers for historical references. `verify-md-links` checks local targets. [Reference validation](../scripts/verify-repository-references.ts) rejects actual commit identifiers and disallowed organization URLs in maintained files.

+ 2 - 2
docs/session-format-status.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/session-format-status.md
-session-format-status.md: 9076574b9b12b4f75615ff06939795a8f0051def
-session-format-status.zh.md: 6c230479f8e41441f0c2e9848448a61d772a46b9
+session-format-status.md: 4408aebf1304159596d01064c358e4b8954d4fab
+session-format-status.zh.md: f7969fd382d4d9be0b2401e957dc55ceca156935

+ 3 - 3
docs/session-format-status.md

@@ -30,14 +30,14 @@ latestReleasedVersion: 3
 evidenceTag: dsh-v0.1.5-alpha.1
 ```
 
-Evidence: [published release](https://github.com/deepseek-harness/deepseek-harness/releases/tag/dsh-v0.1.5-alpha.1) and [its tagged writer source](https://github.com/deepseek-harness/deepseek-harness/blob/dsh-v0.1.5-alpha.1/packages/core/session/src/types.ts).
+Evidence: published product tag `dsh-v0.1.5-alpha.1`; tagged writer: `packages/core/session/src/types.ts`.
 
 <a id="updating-the-record"></a>
 ## Updating the record
 
-When a structural writer change is implemented, update the code constant and adjacent catalog together; do not advance this release record before publication. When a product release first publishes a higher Session format, confirm publication and its tagged writer, then advance this record and both evidence links in the same bilingual update. Later product releases carrying the same format do not require changing the record. Never lower it on the development trunk.
+When a structural writer change is implemented, update the code constant and adjacent catalog together; do not advance this release record before publication. When a product release first publishes a higher Session format, confirm publication and its tagged writer, then advance this record and the evidence tag and tagged writer path in the same bilingual update. Later product releases carrying the same format do not require changing the record. Never lower it on the development trunk.
 
-The [documentation-standard test](../scripts/doc-standard.spec.ts) checks record structure, bilingual equality, evidence-link consistency, and that the documented release does not exceed the checkout writer. This keyless check does not query GitHub or prove that the record is up to date; publication verification remains part of the release update.
+The [documentation-standard test](../scripts/doc-standard.spec.ts) checks record structure, bilingual equality, evidence-tag and writer-path consistency, and that the documented release does not exceed the checkout writer. This keyless check does not query GitHub or prove that the record is up to date; publication verification remains part of the release update.
 
 Use “current format” and “next adjacent version” for general behavior. Keep explicit numbers for fixed migration inputs and outputs, wire schemas, historical evidence, and tests of those particular versions. The [format-version cookbook](cookbook/adding-a-session-format-version.md) uses N for the verified latest released format and N+1 for its successor.
 

+ 3 - 3
docs/session-format-status.zh.md

@@ -30,14 +30,14 @@ latestReleasedVersion: 3
 evidenceTag: dsh-v0.1.5-alpha.1
 ```
 
-证据:[已发布产品版本](https://github.com/deepseek-harness/deepseek-harness/releases/tag/dsh-v0.1.5-alpha.1)及[对应标签的写入器源码](https://github.com/deepseek-harness/deepseek-harness/blob/dsh-v0.1.5-alpha.1/packages/core/session/src/types.ts)
+证据:已发布产品标签 `dsh-v0.1.5-alpha.1`;该标签的写入器路径:`packages/core/session/src/types.ts`
 
 <a id="updating-the-record"></a>
 ## 更新记录
 
-实现结构性写入器变更时,一起更新代码常量与相邻迁移目录;不要在产品发布前推进此发布记录。当产品首次发布更高的 Session 格式时,确认发布事实及对应标签的写入器,然后在同一次双语更新中推进本记录与两个证据链接。后续携带相同格式的产品发布无需改变此记录。开发主干上的记录绝不降低。
+实现结构性写入器变更时,一起更新代码常量与相邻迁移目录;不要在产品发布前推进此发布记录。当产品首次发布更高的 Session 格式时,确认发布事实及对应标签的写入器,然后在同一次双语更新中推进本记录与证据标签及该标签的写入器路径。后续携带相同格式的产品发布无需改变此记录。开发主干上的记录绝不降低。
 
-[文档标准测试](../scripts/doc-standard.spec.ts)检查记录结构、双语一致性、证据链接一致性,以及文档中的已发布版本不高于工作区写入器。这个无密钥检查不会查询 GitHub,也不能证明记录是最新的;核实发布事实仍属于发布更新的一部分。
+[文档标准测试](../scripts/doc-standard.spec.ts)检查记录结构、双语一致性、证据标签及写入器路径一致性,以及文档中的已发布版本不高于工作区写入器。这个无密钥检查不会查询 GitHub,也不能证明记录是最新的;核实发布事实仍属于发布更新的一部分。
 
 一般行为使用“当前格式”和“下一条相邻版本”等表述。固定迁移的输入与输出、协议 schema、历史证据及针对特定版本的测试保留明确版本号。[格式版本实操手册](cookbook/adding-a-session-format-version.zh.md)用 N 表示已核实的最新发布格式,用 N+1 表示其后继版本。
 

+ 2 - 2
docs/user/develop/basic/publish.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/user/develop/basic/publish.md
-publish.md: 58abad728d5f684a72d855b77821a26a725de53f
-publish.zh.md: dbb49c0cd0e5ca26e42d17d94acb5f1601a27d15
+publish.md: 17b0cb80f9cf385773ae9e1f4d7061ec1df9c1a2
+publish.zh.md: 3cc65af58d34028a6bff03455451153e2c7f439b

+ 1 - 1
docs/user/develop/basic/publish.md

@@ -160,7 +160,7 @@ dsh plugin --profile demo add github:you/hello-plugin
 
 But a git install fetches **sources, not built artifacts**: nothing runs your `build` script, so a TypeScript package arrives without its `lib/` output and fails to load. Two things must happen, one on each side:
 
-- **The author** ships a `prepare` script — pnpm runs it after a git install — that builds the published entry points from source, self-contained: it must not assume dev-only context such as a sibling monorepo checkout. [turtle-ui](https://github.com/deepseek-harness/turtle-ui) is a working example: its `prepare` runs a dedicated tsdown config that transpiles `src/` without project references or type checking.
+- **The author** ships a `prepare` script — pnpm runs it after a git install — that builds the published entry points from source, self-contained: it must not assume dev-only context such as a sibling monorepo checkout. A dedicated tsdown config can transpile `src/` without project references or type checking.
 - **The user** allowlists the build. pnpm ≥10 refuses to run a git dependency's `prepare` script until it is explicitly allowed, so the first `add` fails; `dsh` points at the fix — copy the exact package key pnpm printed into the profile's `pnpm-workspace.yaml`:
 
   ```yaml

+ 1 - 1
docs/user/develop/basic/publish.zh.md

@@ -160,7 +160,7 @@ dsh plugin --profile demo add github:you/hello-plugin
 
 但 git 安装拉取的是**源码,不是构建产物**:没有任何环节运行你的 `build` 脚本,因此 TypeScript 包到手时没有 `lib/` 输出,加载会失败。必须两边各做一件事:
 
-- **作者**提供一个 `prepare` 脚本——pnpm 在 git 安装后运行它——从源码构建出发布入口,且必须自包含:不能假设仅开发环境才有的上下文,例如旁边有一份 monorepo checkout。[turtle-ui](https://github.com/deepseek-harness/turtle-ui) 是一个可用的例子:它的 `prepare` 运行一份专用的 tsdown 配置,直接转译 `src/`,不用项目引用,也不做类型检查。
+- **作者**提供一个 `prepare` 脚本——pnpm 在 git 安装后运行它——从源码构建出发布入口,且必须自包含:不能假设仅开发环境才有的上下文,例如旁边有一份 monorepo checkout。专用的 tsdown 配置可以直接转译 `src/`,不用项目引用,也不做类型检查。
 - **用户**为构建授权。pnpm ≥10 在得到显式允许之前拒绝运行 git 依赖的 `prepare` 脚本,所以第一次 `add` 会失败;`dsh` 会指出修法——把 pnpm 打印的确切包键复制进该 profile 的 `pnpm-workspace.yaml`:
 
   ```yaml

+ 2 - 0
native/system/docs/packaging.md

@@ -21,4 +21,6 @@ The `./landlock-run` API stays importable without native payloads and reports un
 
 Platform tarballs use npm pack to preserve the launcher's executable bit. The entry uses pnpm pack to convert workspace dependency versions. Prepack rejects missing or undeclared payloads, invalid ELF/Mach-O architecture or type, addons without Node-API exports, and launchers without executable permission.
 
+Maintained manifests identify the public source repository. When `GITHUB_REPOSITORY` is set, the release packer copies packages and prepack scripts into a temporary workspace and sets the copied manifests' repository URL from that workflow repository and `GITHUB_SERVER_URL` (default `https://github.com`). This preserves npm trusted publishing's repository identity without changing source manifests. GitHub Actions requires a valid repository identity; local packing without workflow context retains the public source URL. Staging is removed after packing or a prepack failure, and publication consumes the resulting tarballs unchanged.
+
 The installed-artifact rehearsal verifies concrete dependency versions and absence of installation hooks, performs an offline npm install from local tarballs, and compares installed bytes with build outputs. It then proves flock contention/close release and probes the installed Landlock launcher; real confinement remains required on enforcing CI kernels.

+ 2 - 2
native/system/package.json

@@ -11,11 +11,11 @@
     "build:native": "tsx ./scripts/build.ts",
     "build:test-oracle": "node ./scripts/build-test-oracle.mjs",
     "typecheck": "tsc --noEmit && tsc -b --dry",
-    "test": "node ./test/entry.test.js && node ./test/launcher.test.js && node --test ./test/flock.test.js ./test/package-matrix.test.js",
+    "test": "node ./test/entry.test.js && node ./test/launcher.test.js && node --test ./test/flock.test.js ./test/package-matrix.test.js ./test/release-packing.test.js",
     "test:entry": "node ./test/entry.test.js",
     "test:launcher": "node ./test/launcher.test.js",
     "test:flock": "node --test ./test/flock.test.js",
-    "test:packaging": "node --test ./test/package-matrix.test.js",
+    "test:packaging": "node --test ./test/package-matrix.test.js ./test/release-packing.test.js",
     "gha:matrix": "node ./scripts/github-matrix.mjs",
     "release:bump": "node ./scripts/bump-release.mjs",
     "release:commit": "node ./scripts/commit-release.mjs",

+ 1 - 1
native/system/packages/darwin-arm64/package.json

@@ -4,7 +4,7 @@
   "description": "Prebuilt POSIX flock Node-API binding for macOS arm64",
   "repository": {
     "type": "git",
-    "url": "git+https://github.com/deepseek-harness/deepseek-harness.git",
+    "url": "git+https://github.com/deepseek-ai/deepseek-harness.git",
     "directory": "native/system/packages/darwin-arm64"
   },
   "os": [

+ 1 - 1
native/system/packages/darwin-x64/package.json

@@ -4,7 +4,7 @@
   "description": "Prebuilt POSIX flock Node-API binding for macOS x64",
   "repository": {
     "type": "git",
-    "url": "git+https://github.com/deepseek-harness/deepseek-harness.git",
+    "url": "git+https://github.com/deepseek-ai/deepseek-harness.git",
     "directory": "native/system/packages/darwin-x64"
   },
   "os": [

+ 1 - 1
native/system/packages/entry/package.json

@@ -5,7 +5,7 @@
   "description": "Prebuilt system primitives: a Linux Landlock launcher and asynchronous POSIX flock through stable Node-API",
   "repository": {
     "type": "git",
-    "url": "git+https://github.com/deepseek-harness/deepseek-harness.git",
+    "url": "git+https://github.com/deepseek-ai/deepseek-harness.git",
     "directory": "native/system/packages/entry"
   },
   "exports": {

+ 1 - 1
native/system/packages/linux-arm64/package.json

@@ -4,7 +4,7 @@
   "description": "Linux arm64 system binaries: static Landlock launcher and glibc/musl Node-API flock addons",
   "repository": {
     "type": "git",
-    "url": "git+https://github.com/deepseek-harness/deepseek-harness.git",
+    "url": "git+https://github.com/deepseek-ai/deepseek-harness.git",
     "directory": "native/system/packages/linux-arm64"
   },
   "os": [

+ 1 - 1
native/system/packages/linux-x64/package.json

@@ -4,7 +4,7 @@
   "description": "Linux x64 system binaries: static Landlock launcher and glibc/musl Node-API flock addons",
   "repository": {
     "type": "git",
-    "url": "git+https://github.com/deepseek-harness/deepseek-harness.git",
+    "url": "git+https://github.com/deepseek-ai/deepseek-harness.git",
     "directory": "native/system/packages/linux-x64"
   },
   "os": [

+ 72 - 28
native/system/scripts/pack-release.mjs

@@ -10,9 +10,12 @@
  * The flag packs only THIS host's platform package plus the entries — for
  * per-architecture CI legs, where the other architecture's binary does not
  * exist (the exact refusal its prepack gate exists for).
+ * Workflow repository metadata is projected only into disposable pack inputs;
+ * source manifests retain the public source home. See docs/packaging.md.
  */
 
 import fs from 'node:fs';
+import os from 'node:os';
 import path from 'node:path';
 import { spawnSync } from 'node:child_process';
 import { entryDirs, platformDirs, readJson, root } from './repo.mjs';
@@ -26,14 +29,14 @@ function hostPlatformDirs() {
   return platformDirs().filter((dir) => readJson(path.join(root, dir, 'prebuilds.json')).platform === hostPlatform);
 }
 
-function run(command, args) {
+function run(command, args, cwd) {
   const result = spawnSync(command, args, {
-    cwd: root,
+    cwd,
     stdio: 'inherit',
   });
   if (result.error) throw result.error;
   if (result.status !== 0) {
-    process.exit(result.status ?? 1);
+    throw new Error(`${command} failed (status=${result.status}, signal=${result.signal})`);
   }
 }
 
@@ -44,33 +47,74 @@ function tarballName(manifest) {
   return `${manifest.name}-${manifest.version}.tgz`;
 }
 
-fs.rmSync(destination, { recursive: true, force: true });
-fs.mkdirSync(destination, { recursive: true });
-
-const dirs = [...(currentPlatformOnly ? hostPlatformDirs() : platformDirs()), ...entryDirs()];
-const platformSet = new Set(platformDirs());
-const publishOrder = [];
-for (const dir of dirs) {
-  const manifest = readJson(path.join(root, dir, 'package.json'));
-  // Platform packages are packed with npm: pnpm pack (observed on 11.7.0)
-  // normalizes file modes and STRIPS the executable bit, which ships a
-  // launcher no consumer can spawn; npm pack preserves it. Platform packages
-  // have no dependencies by construction, so they need none of pnpm's
-  // workspace-protocol conversion — the entry packages do, and carry no
-  // executables, so they keep pnpm pack.
-  if (platformSet.has(dir)) {
-    run('npm', ['pack', `./${dir}`, '--pack-destination', destination]);
-  } else {
-    run('pnpm', ['--dir', dir, 'pack', '--pack-destination', destination]);
+/** Repository identity npm verifies against the workflow's OIDC claims. */
+function workflowRepositoryUrl() {
+  const repository = process.env.GITHUB_REPOSITORY;
+  if (repository === undefined && process.env.GITHUB_ACTIONS !== 'true') return undefined;
+  if (repository === undefined || !/^[a-zA-Z0-9_.-]+\/[a-zA-Z0-9_.-]+$/.test(repository)) {
+    throw new Error('GITHUB_REPOSITORY must identify the workflow owner/repository');
+  }
+  const server = new URL(process.env.GITHUB_SERVER_URL || 'https://github.com');
+  if (server.protocol !== 'https:' || server.pathname !== '/' || server.search || server.hash
+    || server.username || server.password) {
+    throw new Error('GITHUB_SERVER_URL must be an HTTPS origin');
   }
+  return `git+${server.origin}/${repository}.git`;
+}
 
-  const tarball = tarballName(manifest);
-  const tarballPath = path.join(destination, tarball);
-  if (!fs.existsSync(tarballPath)) {
-    throw new Error(`expected pack output not found: ${tarballPath}`);
+/** Copy pack inputs while keeping source manifests and completed tarballs untouched. */
+function stagePackages(staging, repositoryUrl) {
+  for (const directory of ['packages', 'scripts']) {
+    fs.cpSync(path.join(root, directory), path.join(staging, directory), {
+      recursive: true,
+      verbatimSymlinks: true,
+    });
+  }
+  fs.copyFileSync(path.join(root, 'package.json'), path.join(staging, 'package.json'));
+  fs.writeFileSync(path.join(staging, 'pnpm-workspace.yaml'), 'packages:\n  - packages/*\n');
+  for (const dir of [...platformDirs(), ...entryDirs()]) {
+    const manifestPath = path.join(staging, dir, 'package.json');
+    const manifest = readJson(manifestPath);
+    manifest.repository = { ...manifest.repository, url: repositoryUrl };
+    fs.writeFileSync(manifestPath, `${JSON.stringify(manifest, null, 2)}\n`);
   }
-  publishOrder.push(tarball);
 }
 
-fs.writeFileSync(path.join(destination, 'publish-order.txt'), `${publishOrder.join('\n')}\n`);
-console.log(`Packed ${publishOrder.length} packages into ${path.relative(root, destination)}`);
+const repositoryUrl = workflowRepositoryUrl();
+const staging = repositoryUrl === undefined ? undefined : fs.mkdtempSync(path.join(os.tmpdir(), 'native-system-pack-'));
+try {
+  if (staging !== undefined) stagePackages(staging, repositoryUrl);
+  const packRoot = staging ?? root;
+  fs.rmSync(destination, { recursive: true, force: true });
+  fs.mkdirSync(destination, { recursive: true });
+
+  const dirs = [...(currentPlatformOnly ? hostPlatformDirs() : platformDirs()), ...entryDirs()];
+  const platformSet = new Set(platformDirs());
+  const publishOrder = [];
+  for (const dir of dirs) {
+    const manifest = readJson(path.join(root, dir, 'package.json'));
+    // Platform packages are packed with npm: pnpm pack (observed on 11.7.0)
+    // normalizes file modes and STRIPS the executable bit, which ships a
+    // launcher no consumer can spawn; npm pack preserves it. Platform packages
+    // have no dependencies by construction, so they need none of pnpm's
+    // workspace-protocol conversion — the entry packages do, and carry no
+    // executables, so they keep pnpm pack.
+    if (platformSet.has(dir)) {
+      run('npm', ['pack', `./${dir}`, '--pack-destination', destination], packRoot);
+    } else {
+      run('pnpm', ['--dir', dir, 'pack', '--pack-destination', destination], packRoot);
+    }
+
+    const tarball = tarballName(manifest);
+    const tarballPath = path.join(destination, tarball);
+    if (!fs.existsSync(tarballPath)) {
+      throw new Error(`expected pack output not found: ${tarballPath}`);
+    }
+    publishOrder.push(tarball);
+  }
+
+  fs.writeFileSync(path.join(destination, 'publish-order.txt'), `${publishOrder.join('\n')}\n`);
+  console.log(`Packed ${publishOrder.length} packages into ${path.relative(root, destination)}`);
+} finally {
+  if (staging !== undefined) fs.rmSync(staging, { recursive: true, force: true });
+}

+ 134 - 0
native/system/test/release-packing.test.js

@@ -0,0 +1,134 @@
+import assert from 'node:assert/strict';
+import fs from 'node:fs';
+import os from 'node:os';
+import path from 'node:path';
+import { spawnSync } from 'node:child_process';
+import { fileURLToPath } from 'node:url';
+import { test } from 'node:test';
+
+const publicRepository = 'git+https://github.com/deepseek-ai/deepseek-harness.git';
+
+/** The real packer and prepack scripts operate on isolated, format-valid pack inputs. */
+function fixture(t) {
+  const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'native-packing-test-'));
+  t.after(() => fs.rmSync(dir, { recursive: true, force: true }));
+  const scratch = path.join(dir, 'scratch');
+  fs.mkdirSync(scratch);
+  fs.cpSync(fileURLToPath(new URL('../scripts/', import.meta.url)), path.join(dir, 'scripts'), { recursive: true });
+  const writeJson = (file, value) => {
+    fs.mkdirSync(path.dirname(path.join(dir, file)), { recursive: true });
+    fs.writeFileSync(path.join(dir, file), `${JSON.stringify(value, null, 2)}\n`);
+  };
+  const repository = (name) => ({ type: 'git', url: publicRepository, directory: `native/system/packages/${name}` });
+  writeJson('package.json', { name: 'native-packing-fixture', private: true, version: '1.2.3', packageManager: 'pnpm@11.7.0' });
+  fs.writeFileSync(path.join(dir, 'pnpm-workspace.yaml'), 'packages:\n  - packages/*\n');
+  writeJson('packages/linux-x64/package.json', {
+    name: '@fixture/native-linux-x64', version: '1.2.3', repository: repository('linux-x64'),
+    os: ['linux'], cpu: ['x64'], files: ['bin/', 'prebuilds.json'],
+    scripts: { prepack: 'node ../../scripts/verify-launcher-binary.mjs' },
+  });
+  writeJson('packages/linux-x64/prebuilds.json', {
+    platform: 'linux-x64', binaries: [{ tool: 'landlock-run', kind: 'static-musl', path: 'bin/landlock-run' }],
+  });
+  const binary = Buffer.alloc(64);
+  binary.writeUInt32LE(0x464c457f, 0);
+  binary[4] = 2;
+  binary[5] = 1;
+  binary.writeUInt16LE(2, 16);
+  binary.writeUInt16LE(62, 18);
+  fs.mkdirSync(path.join(dir, 'packages/linux-x64/bin'));
+  fs.writeFileSync(path.join(dir, 'packages/linux-x64/bin/landlock-run'), binary, { mode: 0o755 });
+  writeJson('packages/entry/package.json', {
+    name: '@fixture/native', version: '1.2.3', type: 'module', repository: repository('entry'),
+    exports: { '.': './lib/index.js' }, files: ['lib/'],
+    optionalDependencies: { '@fixture/native-linux-x64': 'workspace:*' },
+    scripts: { prepack: 'node ../../scripts/verify-entry-lib.mjs' },
+  });
+  fs.mkdirSync(path.join(dir, 'packages/entry/lib'));
+  fs.writeFileSync(path.join(dir, 'packages/entry/lib/index.js'), 'export const packed = true;\n');
+  fs.mkdirSync(path.join(dir, 'packages/entry/node_modules/@fixture'), { recursive: true });
+  fs.symlinkSync('../../../linux-x64', path.join(dir, 'packages/entry/node_modules/@fixture/native-linux-x64'), 'junction');
+  const manifests = ['package.json', 'packages/entry/package.json', 'packages/linux-x64/package.json'];
+  const original = manifests.map(file => fs.readFileSync(path.join(dir, file), 'utf8'));
+  const env = Object.fromEntries(Object.entries(process.env)
+    .filter(([key]) => !/KEY|TOKEN|SECRET|PASSWORD|^GITHUB_|^NODE_PATH$/i.test(key)));
+  Object.assign(env, { TMPDIR: scratch, TMP: scratch, TEMP: scratch, npm_config_cache: path.join(dir, 'npm-cache') });
+  return {
+    dir, binary,
+    pack: (extra = {}) => spawnSync(process.execPath, ['scripts/pack-release.mjs', 'output'], {
+      cwd: dir, env: { ...env, ...extra }, encoding: 'utf8', timeout: 120_000,
+    }),
+    unchanged() {
+      assert.deepEqual(manifests.map(file => fs.readFileSync(path.join(dir, file), 'utf8')), original);
+      assert.deepEqual(fs.readdirSync(scratch).filter(name => name.startsWith('native-system-pack-')), []);
+    },
+  };
+}
+
+function completed(result) {
+  assert.equal(result.error, undefined);
+  assert.equal(result.signal, null);
+  assert.equal(result.status, 0, result.stdout + result.stderr);
+}
+
+function tarFile(dir, tarball, file) {
+  const result = spawnSync('tar', ['-xOf', path.join(dir, 'output', tarball), `package/${file}`], {
+    timeout: 120_000,
+  });
+  completed(result);
+  return result.stdout;
+}
+
+for (const workflow of [false, true]) {
+  test(`packing ${workflow ? 'workflow' : 'local'} artifacts preserves source manifests and native payloads`, { timeout: 180_000 }, (t) => {
+    const f = fixture(t);
+    const result = f.pack(workflow ? {
+      GITHUB_ACTIONS: 'true', GITHUB_REPOSITORY: 'fixture-owner/native-runtime', GITHUB_SERVER_URL: 'https://github.example.com',
+    } : {});
+    completed(result);
+    f.unchanged();
+    const files = fs.readFileSync(path.join(f.dir, 'output/publish-order.txt'), 'utf8').trim().split('\n');
+    assert.deepEqual(files, ['fixture-native-linux-x64-1.2.3.tgz', 'fixture-native-1.2.3.tgz']);
+    const manifests = files.map(file => JSON.parse(tarFile(f.dir, file, 'package.json')));
+    for (const manifest of manifests) {
+      assert.equal(manifest.repository.url, workflow ? 'git+https://github.example.com/fixture-owner/native-runtime.git' : publicRepository);
+      assert.equal(manifest.repository.type, 'git');
+    }
+    assert.equal(manifests[0].repository.directory, 'native/system/packages/linux-x64');
+    assert.equal(manifests[1].repository.directory, 'native/system/packages/entry');
+    assert.equal(manifests[1].optionalDependencies['@fixture/native-linux-x64'], '1.2.3');
+    assert.deepEqual(tarFile(f.dir, files[0], 'bin/landlock-run'), f.binary);
+    const extracted = path.join(f.dir, 'extracted');
+    fs.mkdirSync(extracted);
+    completed(spawnSync('tar', ['-xzf', path.join(f.dir, 'output', files[0]), '-C', extracted], { timeout: 120_000 }));
+    if (process.platform !== 'win32') assert.notEqual(fs.statSync(path.join(extracted, 'package/bin/landlock-run')).mode & 0o111, 0);
+  });
+}
+
+test('invalid workflow identity fails before deleting existing pack output', (t) => {
+  const f = fixture(t);
+  fs.mkdirSync(path.join(f.dir, 'output'));
+  fs.writeFileSync(path.join(f.dir, 'output/retained'), 'previous artifact');
+  for (const env of [{ GITHUB_ACTIONS: 'true' }, { GITHUB_REPOSITORY: 'owner/repo/extra' }, {
+    GITHUB_REPOSITORY: 'owner/repo', GITHUB_SERVER_URL: 'https://github.example.com/path',
+  }]) {
+    const result = f.pack(env);
+    assert.equal(result.error, undefined);
+    assert.equal(result.signal, null);
+    assert.equal(result.status, 1);
+    assert.match(result.stderr, /GITHUB_(?:REPOSITORY|SERVER_URL)/);
+    assert.equal(fs.readFileSync(path.join(f.dir, 'output/retained'), 'utf8'), 'previous artifact');
+    f.unchanged();
+  }
+});
+
+test('a prepack rejection cleans staging and leaves source manifests unchanged', { timeout: 180_000 }, (t) => {
+  const f = fixture(t);
+  fs.rmSync(path.join(f.dir, 'packages/linux-x64/bin/landlock-run'));
+  const result = f.pack({ GITHUB_REPOSITORY: 'fixture-owner/native-runtime' });
+  assert.equal(result.error, undefined);
+  assert.equal(result.signal, null);
+  assert.equal(result.status, 1);
+  assert.match(result.stderr, /missing bin\/landlock-run/);
+  f.unchanged();
+});

+ 1 - 0
package.json

@@ -96,6 +96,7 @@
     "verify-md-links": "tsx scripts/verify-md-links.ts",
     "verify-doc-site-fragments": "tsx scripts/verify-doc-site-fragments.ts",
     "verify-public-repository-links": "tsx scripts/verify-public-repository-links.ts",
+    "verify-repository-references": "tsx scripts/verify-repository-references.ts",
     "verify-concrete-terms": "tsx scripts/verify-concrete-terms.ts",
     "verify-doc-refs": "tsx scripts/verify-doc-refs.ts",
     "verify-subsystem-pages": "tsx scripts/verify-subsystem-pages.ts",

+ 3 - 8
scripts/check-workspace-constraints.ts

@@ -44,12 +44,7 @@ const publicNativePackages = new Set([
 const publicationSourceAllowlist: Readonly<Record<string, readonly string[]>> = {
   '@deepseek-ai/node-addon-system': ['src/main.c', 'src/flock.c'],
 }
-const repositoryUrl = 'git+https://github.com/deepseek-harness/deepseek-harness.git'
-/**
- * Source home the published packages point consumers at. It differs from
- * {@link repositoryUrl}, which the Landlock packages keep because npm resolves
- * their trusted publishing against the repository that runs the workflow.
- */
+/** Public source home recorded in maintained package manifests. */
 const publishedRepositoryUrl = 'git+https://github.com/deepseek-ai/deepseek-harness.git'
 /** Packages that participate in the experimental policy. */
 const experimentalPackageDirectory = /^packages\/experimental\/[^/]+$/
@@ -351,9 +346,9 @@ export function checkWorkspaceManifest({ dir, manifest }: WorkspaceManifest): st
     }
     const expectedDirectory = dir
     if (manifest.repository?.type !== 'git'
-      || manifest.repository.url !== repositoryUrl
+      || manifest.repository.url !== publishedRepositoryUrl
       || manifest.repository.directory !== expectedDirectory) {
-      errors.push(`${label}: published Landlock package repository must use ${repositoryUrl} with directory ${expectedDirectory} for trusted publishing`)
+      errors.push(`${label}: published Landlock package repository must use ${publishedRepositoryUrl} with directory ${expectedDirectory}`)
     }
   } else if (isReleaseMemberDirectory(dir)) {
     // Release members state that they are publishable: npm refuses a private

+ 37 - 42
scripts/doc-standard.spec.ts

@@ -155,7 +155,7 @@ interface SessionFormatRelease {
   evidenceTag: string
 }
 
-/** Validate the release record and evidence links; throw on malformed or inconsistent input. */
+/** Validate release metadata and tagged writer evidence; throw on malformed or inconsistent input. */
 function validateSessionFormatRelease(source: string, currentWriterVersion: number): SessionFormatRelease {
   const normalized = source.replaceAll('\r\n', '\n')
   const openings = [...normalized.matchAll(/^```yaml session-format-release[ \t]*$/gmu)]
@@ -182,34 +182,28 @@ function validateSessionFormatRelease(source: string, currentWriterVersion: numb
     || !/^dsh-v\d+\.\d+\.\d+(?:-[\dA-Za-z]+(?:[.-][\dA-Za-z]+)*)?(?:\+[\dA-Za-z]+(?:[.-][\dA-Za-z]+)*)?$/u.test(evidenceTag)) {
     throw new Error('evidenceTag must be a non-empty dsh-v version tag without URL delimiters')
   }
-  const repository = 'https://github.com/deepseek-harness/deepseek-harness'
-  for (const link of [
-    `${repository}/releases/tag/${evidenceTag}`,
-    `${repository}/blob/${evidenceTag}/packages/core/session/src/types.ts`,
-  ]) {
-    if (!normalized.includes(`](${link})`)) throw new Error(`Missing matching evidence link: ${link}`)
-  }
+  const hasEvidence = normalized.split('\n').some(line =>
+    line.includes(`\`${evidenceTag}\``) && line.includes('`packages/core/session/src/types.ts`'))
+  if (!hasEvidence) throw new Error('Missing matching evidence tag and tagged writer path')
   return { latestReleasedVersion, evidenceTag }
 }
 
-function sessionFormatReleaseFixture(): { record: SessionFormatRelease; body: string; links: string; source: string } {
+function sessionFormatReleaseFixture(): { record: SessionFormatRelease; body: string; evidence: string; source: string } {
   const record = validateSessionFormatRelease(
     readFileSync(resolve(root, 'docs/session-format-status.md'), 'utf8'),
     readCurrentSessionFormatVersion(root),
   )
   const body = `latestReleasedVersion: ${record.latestReleasedVersion}\nevidenceTag: ${record.evidenceTag}`
-  const repository = 'https://github.com/deepseek-harness/deepseek-harness'
-  const links = `[release](${repository}/releases/tag/${record.evidenceTag})\n`
-    + `[source](${repository}/blob/${record.evidenceTag}/packages/core/session/src/types.ts)`
-  return { record, body, links, source: releaseDocument(body, links) }
+  const evidence = `Evidence: published product tag \`${record.evidenceTag}\`; tagged writer: \`packages/core/session/src/types.ts\`.`
+  return { record, body, evidence, source: releaseDocument(body, evidence) }
 }
 
-function releaseDocument(body: string, links: string): string {
-  return `\`\`\`yaml session-format-release\n${body}\n\`\`\`\n\n${links}\n`
+function releaseDocument(body: string, evidence: string): string {
+  return `\`\`\`yaml session-format-release\n${body}\n\`\`\`\n\n${evidence}\n`
 }
 
 describe('Session format release authority', () => {
-  it('keeps the bilingual release records equal and consistent with the writer and evidence links', () => {
+  it('keeps bilingual release metadata consistent with the writer and tagged evidence', () => {
     const records = ['docs/session-format-status.md', 'docs/session-format-status.zh.md'].map(file =>
       validateSessionFormatRelease(readFileSync(resolve(root, file), 'utf8'), readCurrentSessionFormatVersion(root)),
     )
@@ -217,33 +211,33 @@ describe('Session format release authority', () => {
   })
 
   it('accepts a released writer and a newer development writer, including format zero', () => {
-    const { record, body, links, source } = sessionFormatReleaseFixture()
+    const { record, body, evidence, source } = sessionFormatReleaseFixture()
     expect(validateSessionFormatRelease(source, record.latestReleasedVersion)).toEqual(record)
     expect(validateSessionFormatRelease(source, record.latestReleasedVersion + 1)).toEqual(record)
-    const zero = releaseDocument(body.replace(`latestReleasedVersion: ${record.latestReleasedVersion}`, 'latestReleasedVersion: 0'), links)
+    const zero = releaseDocument(body.replace(`latestReleasedVersion: ${record.latestReleasedVersion}`, 'latestReleasedVersion: 0'), evidence)
     expect(validateSessionFormatRelease(zero, 0)).toEqual({ ...record, latestReleasedVersion: 0 })
   })
 
   it('rejects missing, duplicated, unclosed, and malformed release records', () => {
-    const { record, body, links, source } = sessionFormatReleaseFixture()
+    const { record, body, evidence, source } = sessionFormatReleaseFixture()
     for (const invalid of [
-      links,
+      evidence,
       source + source,
       source + '\n```yaml session-format-release\n',
       `\`\`\`yaml session-format-release\n${body}`,
-      releaseDocument('[unterminated', links),
-      releaseDocument('', links),
-      releaseDocument('null', links),
-      releaseDocument('scalar', links),
-      releaseDocument(`- latestReleasedVersion: ${record.latestReleasedVersion}`, links),
-      releaseDocument(`${body}\n---\n${body}`, links),
+      releaseDocument('[unterminated', evidence),
+      releaseDocument('', evidence),
+      releaseDocument('null', evidence),
+      releaseDocument('scalar', evidence),
+      releaseDocument(`- latestReleasedVersion: ${record.latestReleasedVersion}`, evidence),
+      releaseDocument(`${body}\n---\n${body}`, evidence),
     ]) {
       expect(() => validateSessionFormatRelease(invalid, record.latestReleasedVersion), invalid).toThrow()
     }
   })
 
   it('rejects missing, duplicate, and extra record fields', () => {
-    const { record, body, links } = sessionFormatReleaseFixture()
+    const { record, body, evidence } = sessionFormatReleaseFixture()
     for (const invalid of [
       '{}',
       `evidenceTag: ${record.evidenceTag}`,
@@ -252,45 +246,46 @@ describe('Session format release authority', () => {
       `${body}\nevidenceTag: ${record.evidenceTag}`,
       `${body}\nreleased: true`,
     ]) {
-      expect(() => validateSessionFormatRelease(releaseDocument(invalid, links), record.latestReleasedVersion), invalid).toThrow()
+      expect(() => validateSessionFormatRelease(releaseDocument(invalid, evidence), record.latestReleasedVersion), invalid).toThrow()
     }
   })
 
   it('rejects invalid released versions and releases beyond the current writer', () => {
-    const { record, links } = sessionFormatReleaseFixture()
+    const { record, evidence } = sessionFormatReleaseFixture()
     for (const value of ['-1', '1.5', String(Number.MAX_SAFE_INTEGER + 1), '.inf', '.nan', 'null', 'true', '"0"']) {
-      const source = releaseDocument(`latestReleasedVersion: ${value}\nevidenceTag: ${record.evidenceTag}`, links)
+      const source = releaseDocument(`latestReleasedVersion: ${value}\nevidenceTag: ${record.evidenceTag}`, evidence)
       expect(() => validateSessionFormatRelease(source, Number.MAX_SAFE_INTEGER), value).toThrow('non-negative safe integer')
     }
     const writer = readCurrentSessionFormatVersion(root)
-    const future = releaseDocument(`latestReleasedVersion: ${writer + 1}\nevidenceTag: ${record.evidenceTag}`, links)
+    const future = releaseDocument(`latestReleasedVersion: ${writer + 1}\nevidenceTag: ${record.evidenceTag}`, evidence)
     expect(() => validateSessionFormatRelease(future, writer)).toThrow('must not exceed the current writer')
   })
 
   it('rejects empty, malformed, and URL-injecting evidence tags', () => {
-    const { record, links } = sessionFormatReleaseFixture()
+    const { record, evidence } = sessionFormatReleaseFixture()
     for (const tag of [
       null, true, 1, '', ' ', 'dsh-v', record.evidenceTag.replace('dsh-v', 'v'),
       `${record.evidenceTag}/other`, `${record.evidenceTag}?query`, `${record.evidenceTag}#fragment`,
       `${record.evidenceTag}%2Fother`, `${record.evidenceTag})`, `${record.evidenceTag}\n`,
     ]) {
-      const source = releaseDocument(`latestReleasedVersion: ${record.latestReleasedVersion}\nevidenceTag: ${JSON.stringify(tag)}`, links)
+      const source = releaseDocument(`latestReleasedVersion: ${record.latestReleasedVersion}\nevidenceTag: ${JSON.stringify(tag)}`, evidence)
       expect(() => validateSessionFormatRelease(source, record.latestReleasedVersion), String(tag)).toThrow('dsh-v version tag')
     }
   })
 
-  it('rejects absent or mismatched release and tagged-source links', () => {
-    const { record, body, links } = sessionFormatReleaseFixture()
+  it('rejects absent or mismatched release tags and tagged writer paths', () => {
+    const { record, body, evidence } = sessionFormatReleaseFixture()
     for (const invalid of [
       '',
-      links.replace(`/releases/tag/${record.evidenceTag}`, `/releases/tag/${record.evidenceTag}-other`),
-      links.replace(`/blob/${record.evidenceTag}/`, '/blob/main/'),
-      links.replace('/packages/core/session/src/types.ts', '/packages/core/session/src/other.ts'),
-      links.replaceAll('github.com', 'example.com'),
-      links.replace(`${record.evidenceTag})`, `${record.evidenceTag}?query)`),
-      links.replace('types.ts)', 'types.ts#fragment)'),
+      evidence.replace(record.evidenceTag, `${record.evidenceTag}-other`),
+      evidence.replace(record.evidenceTag, 'main'),
+      evidence.replace('packages/core/session/src/types.ts', 'packages/core/session/src/other.ts'),
+      evidence.replace(record.evidenceTag, `${record.evidenceTag}?query`),
+      evidence.replace('types.ts', 'types.ts#fragment'),
+      evidence.replace('; tagged writer:', ';\n tagged writer:'),
     ]) {
-      expect(() => validateSessionFormatRelease(releaseDocument(body, invalid), record.latestReleasedVersion), invalid).toThrow('Missing matching evidence link')
+      expect(() => validateSessionFormatRelease(releaseDocument(body, invalid), record.latestReleasedVersion), invalid)
+        .toThrow('Missing matching evidence tag and tagged writer path')
     }
   })
 })

+ 6 - 6
scripts/gen-third-party-notices.ts

@@ -443,7 +443,7 @@ export function parseVendoredRows(text: string): VendoredRow[] {
  * disclosed, so a row that stops matching the table format is a hard error
  * rather than a package that quietly vanishes from the notices.
  */
-function collectVendored(): VendoredRow[] {
+function collectVendored(): (VendoredRow & { sourceDirectory: string })[] {
   const rows = parseVendoredRows(readFileSync(resolve(root, 'vendor/README.md'), 'utf8'))
   const onDisk = new Map<string, string>()
   for (const entry of readdirSync(resolve(root, 'vendor'), { withFileTypes: true })) {
@@ -457,15 +457,15 @@ function collectVendored(): VendoredRow[] {
   if (missing.length > 0) {
     throw new Error(`gen-third-party-notices: vendor/README.md has no manifest-table row for ${missing.join(', ')}; its table format changed or the sync is incomplete.`)
   }
-  for (const row of rows) {
+  return rows.map((row) => {
     const dir = onDisk.get(row.npmName)
     if (dir === undefined) throw new Error(`gen-third-party-notices: vendored package ${row.npmName} from vendor/README.md has no vendor/ directory.`)
     const license = readManifest(`vendor/${dir}/package.json`).license
     if (license !== 'MIT') {
       throw new Error(`gen-third-party-notices: vendored ${row.npmName} declares license ${JSON.stringify(license)}; the vendored section assumes MIT throughout.`)
     }
-  }
-  return rows
+    return { ...row, sourceDirectory: `vendor/${dir}` }
+  })
 }
 
 /** Whether a parsed TOML value is a table rather than an array or scalar. */
@@ -725,9 +725,9 @@ The complete npm transitive closure, including the Landlock launcher workspace,
 
 The Cordis framework and its foundation libraries are source-vendored into this repository rather than consumed from npm, and republished under the \`@deepseek-ai\` scope. All are MIT-licensed; each directory preserves its upstream \`LICENSE\` file. Exact upstream commits and local modifications are recorded in [\`vendor/README.md\`](vendor/README.md).
 
-| Package | Upstream name | Upstream | License |
+| Package | Upstream name | Source | License |
 | --- | --- | --- | --- |
-${vendored.map(row => `| \`${row.npmName}\` | \`${row.upstreamName}\` | [${row.upstream.replace('https://', '')}](${row.upstream}) | MIT |`).join('\n')}
+${vendored.map(row => `| \`${row.npmName}\` | \`${row.upstreamName}\` | [${row.sourceDirectory}](${row.sourceDirectory}/) | MIT |`).join('\n')}
 
 ## Runtime npm dependencies
 

+ 14 - 0
scripts/run-gates.spec.ts

@@ -194,6 +194,20 @@ describe('gate graph validation', () => {
     expect(ids).toContain('concrete-terms')
   })
 
+  it('checks all maintained repository references locally and in CI', () => {
+    const { scripts } = JSON.parse(readFileSync(new URL('../package.json', import.meta.url), 'utf8')) as {
+      scripts: Record<string, string>
+    }
+    for (const mode of ['doc-sync', 'doc-quick', 'ci-static'] as const) {
+      const gates = withPnpmEntrypoint(() => gatesForMode(mode))
+      expect(gates).toContainEqual(expect.objectContaining({
+        id: 'repository-references',
+        displayCommand: 'pnpm run verify-repository-references',
+      }))
+    }
+    expect(scripts['verify-repository-references']).toBe('tsx scripts/verify-repository-references.ts')
+  })
+
   it('keeps package-group subsystem ownership in the documentation gate', () => {
     const ids = withPnpmEntrypoint(() => gatesForMode('doc-sync').map(subject => subject.id))
 

Bu fark içinde çok fazla dosya değişikliği olduğu için bazı dosyalar gösterilmiyor