Browse Source

Merge branch 'master' into fix/environment-prompt-suffix

Tianyi Cui 3 weeks ago
parent
commit
beb23febfd
25 changed files with 287 additions and 34 deletions
  1. 6 0
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.i18n.yaml
  2. 27 0
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md
  3. 27 0
      .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.zh.md
  4. 2 2
      .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.i18n.yaml
  5. 1 1
      .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.md
  6. 1 1
      .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.zh.md
  7. 2 2
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml
  8. 5 3
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
  9. 5 3
      .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md
  10. 6 0
      .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.i18n.yaml
  11. 26 0
      .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.md
  12. 26 0
      .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.zh.md
  13. 4 7
      .github/workflows/ci.yml
  14. 26 3
      benchmarks/session-open/session-open.bench.ts
  15. 2 2
      docs/development.i18n.yaml
  16. 1 1
      docs/development.md
  17. 1 1
      docs/development.zh.md
  18. 2 2
      python/sdk-runtime/README.i18n.yaml
  19. 1 1
      python/sdk-runtime/README.md
  20. 1 1
      python/sdk-runtime/README.zh.md
  21. 5 1
      python/sdk-runtime/src/deepseek_harness_runtime/__init__.py
  22. 53 1
      python/sdk/tests/test_runtime_resolution.py
  23. 15 0
      python/sdk/tests/test_smoke_model.py
  24. 41 2
      scripts/ci-workflow.spec.ts
  25. 1 0
      scripts/smoke-python-runtime.py

+ 6 - 0
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md
+2026-09-06-windows-python-console-spawn-wait.md: 92443bcf8a6e4e5609dc469efa4ebd1d82ab127f
+2026-09-06-windows-python-console-spawn-wait.zh.md: dba2f324b29955580fc11e7cea7a0525a8bc8c86

+ 27 - 0
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.md

@@ -0,0 +1,27 @@
+# Agent Note: Wait for the Windows Python console runtime
+
+Status: implemented
+
+English | [中文](2026-09-06-windows-python-console-spawn-wait.zh.md)
+
+## Problem
+
+The installed Python `dsh.exe` console command intermittently exits with Windows access violation `0xc0000005` before initializing a profile. Its smoke assertion omitted the process status and reported only empty streams. A [native faulthandler probe](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34030851888) captures the fault in Python 3.10 `os._execvpe`, called by the runtime console entry, rather than in the bundled Node executable. Direct executable controls pass.
+
+## Decision
+
+The [Python console entry](../../../../python/sdk-runtime/src/deepseek_harness_runtime/__init__.py) uses `subprocess.run` on Windows, inherits standard streams and environment, waits for runtime completion, and exits with the runtime status. POSIX retains `os.execvpe` process replacement. Windows CRT exec is not POSIX process replacement; the explicit spawn-and-wait path avoids the observed native exec operation.
+
+The [installed-wheel smoke](../../../../scripts/smoke-python-runtime.py) reports decimal and unsigned 32-bit hexadecimal status alongside captured streams when profile installation fails. This preserves the distinction between ordinary command failure and native process exceptions.
+
+## Alternatives considered
+
+**Disable Node compile caching.** Not selected: cache environment changes correlated with early probes, but cold-cache controls also passed and Python faulthandler locates the actual fault at the native exec call. Cache configuration remains unchanged.
+
+**Retry or bypass the installed console command.** Rejected because either masks the shipped command failure instead of repairing its process launch. The keyless installed-wheel assertion remains required.
+
+## Consequences
+
+Windows keeps a Python parent until the runtime exits; it no longer depends on CRT overlay behavior. The standard synchronous subprocess implementation owns waiting and interruption cleanup. No custom process-tree manager or global host setting is added.
+
+[Runtime-resolution tests](../../../../python/sdk/tests/test_runtime_resolution.py) retain POSIX forwarding and cover Windows argument/environment forwarding, statuses 0/37/513, real child completion, Unicode streams and arguments with spaces. Native Windows owns the wide exit-status case because POSIX truncates process statuses to eight bits. The [native fixed-count comparison](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34031142773) passes all four patched launches with compile caching enabled; all four unpatched controls also pass in that batch, so it is not a same-batch reproduction. Full installed-wheel CI must validate the final artifact separately from local branch-level tests.

+ 27 - 0
.agents/notes/implemented/bug-fix/2026-09-06-windows-python-console-spawn-wait.zh.md

@@ -0,0 +1,27 @@
+# Agent Note: 等待 Windows Python 控制台运行时
+
+Status: implemented
+
+[English](2026-09-06-windows-python-console-spawn-wait.md) | 中文
+
+## 问题
+
+Python 安装的 `dsh.exe` 控制台命令会在初始化 profile 前间歇性地以 Windows 访问冲突 `0xc0000005` 退出。其冒烟断言遗漏进程状态,只报告空标准流。[原生 faulthandler 探测](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34030851888) 将故障定位在运行时控制台入口调用的 Python 3.10 `os._execvpe`,而非打包的 Node 可执行文件。直接启动可执行文件的对照通过。
+
+## 决策
+
+[Python 控制台入口](../../../../python/sdk-runtime/src/deepseek_harness_runtime/__init__.py) 在 Windows 上使用 `subprocess.run`,继承标准流与环境,等待运行时结束,再以运行时状态退出。POSIX 保留 `os.execvpe` 进程替换。Windows CRT exec 并非 POSIX 进程替换;显式启动并等待的路径避开观测到的原生 exec 操作。
+
+[安装后 wheel 冒烟测试](../../../../scripts/smoke-python-runtime.py) 在 profile 安装失败时,同时报告十进制、无符号 32 位十六进制状态与捕获的标准流。这保留普通命令失败和原生进程异常的区别。
+
+## 已考虑的替代方案
+
+**禁用 Node 编译缓存。** 未采用:早期探测中缓存环境变化与结果相关,但冷缓存对照也能通过,且 Python faulthandler 将实际故障定位在原生 exec 调用。缓存配置保持不变。
+
+**重试或绕过已安装的控制台命令。** 拒绝,因为二者都会掩盖已发布命令的失败,而不是修复进程启动。keyless 安装后 wheel 断言仍为必需检查。
+
+## 后果
+
+Windows 保留 Python 父进程直到运行时退出,不再依赖 CRT overlay 行为。标准同步子进程实现负责等待和中断清理。不添加自定义进程树管理器或全局主机设置。
+
+[运行时解析测试](../../../../python/sdk/tests/test_runtime_resolution.py) 保留 POSIX 转发验证,并覆盖 Windows 参数/环境转发、状态 0/37/513、真实子进程完成、Unicode 标准流和带空格的参数。宽退出状态由原生 Windows 验证,因为 POSIX 会将进程状态截断为八位。[原生固定次数对照](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34031142773) 中,启用编译缓存的四次修复后启动全部通过;该批次四次未修复对照也全部通过,因此它不是同批次复现。完整安装后 wheel CI 必须独立于本地分支级测试,验证最终产物。

+ 2 - 2
.agents/notes/implemented/process/2026-07-26-ci-failover-runbook.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/process/2026-07-26-ci-failover-runbook.md
-2026-07-26-ci-failover-runbook.md: f5711fd7e9c32f7ca555e06bb59b67e2129745d6
-2026-07-26-ci-failover-runbook.zh.md: 05803ff3ad8cee57c2b78a25f2940157e243cb64
+2026-07-26-ci-failover-runbook.md: d5c12492671941c45cf3ccab255dd76fb53773bd
+2026-07-26-ci-failover-runbook.zh.md: b42a14dc9a3e4f40c59b27633752f6969173777e

+ 1 - 1
.agents/notes/implemented/process/2026-07-26-ci-failover-runbook.md

@@ -6,7 +6,7 @@ English | [中文](2026-07-26-ci-failover-runbook.zh.md)
 
 ## Problem
 
-The three required Linux worker jobs in [CI](../../../../.github/workflows/ci.yml) (`node 24 / static`, `node 24 / coverage`, `node 24 / snapshots and artifacts`) run on the hosted enterprise 32-core pools; the required verdict job that aggregates them (`all checks passed`) runs on standard `ubuntu-latest`; the [native Windows jobs](2026-08-08-native-windows-pull-request-ci.md) run on the hosted `dsh-windows-2025-16core` larger runner. When the enterprise pools degrade — jobs queue indefinitely or the enterprise labels vanish — every open pull request becomes unmergeable, and the ordinary recovery of merging a fix is itself deadlocked behind the very required checks that cannot run. **Scope: two independent switches, one per platform.** `DSH_CI_FAILOVER_LINUX` recovers an enterprise Linux-pool outage (the three required Linux workers plus the `all checks passed` verdict); `DSH_CI_FAILOVER_WINDOWS` recovers a hosted Windows-pool outage (the native Windows jobs). A Linux-pool outage need not retarget Windows jobs and vice versa. The verdict's other required dependencies (`node-compat`, `python-sdk`, `windows`) stay on standard hosted runners by design (the portable boundary); in a broader GitHub-hosted capacity failure that also takes out the standard pools, those dependencies still block `all checks passed`. An outage therefore needs a switch any responder with repository write access can throw without merging anything.
+The three required Linux worker jobs in [CI](../../../../.github/workflows/ci.yml) (`node 24 / static`, `node 24 / coverage`, `node 24 / snapshots and artifacts`) run on the hosted enterprise 32-core pools; the required verdict job that aggregates them (`all checks passed`) runs on standard `ubuntu-latest`; the [native Windows jobs](2026-08-08-native-windows-pull-request-ci.md) run on the hosted `dsh-windows-2025-16core` larger runner. When the enterprise pools degrade — jobs queue indefinitely or the enterprise labels vanish — every open pull request becomes unmergeable, and the ordinary recovery of merging a fix is itself deadlocked behind the very required checks that cannot run. **Scope: two independent switches, one per platform.** `DSH_CI_FAILOVER_LINUX` recovers an enterprise Linux-pool outage (the three required Linux workers plus the `all checks passed` verdict); `DSH_CI_FAILOVER_WINDOWS` recovers a hosted Windows-pool outage (the native Windows jobs). A Linux-pool outage need not retarget Windows jobs and vice versa. The verdict's other required dependencies (`node-24-bench`, `node-compat`, `python-sdk`, `windows`) stay on standard hosted runners by design (the portable boundary); in a broader GitHub-hosted capacity failure that also takes out the standard pools, those dependencies still block `all checks passed`. An outage therefore needs a switch any responder with repository write access can throw without merging anything.
 
 ## Decision
 

+ 1 - 1
.agents/notes/implemented/process/2026-07-26-ci-failover-runbook.zh.md

@@ -6,7 +6,7 @@ Status: implemented
 
 ## 问题
 
-[CI](../../../../.github/workflows/ci.yml) 中三个必需的 Linux 工作作业(`node 24 / static`、`node 24 / coverage`、`node 24 / snapshots and artifacts`)运行在托管的企业级 32 核池上;聚合它们的必需判定作业(`all checks passed`)运行在标准 `ubuntu-latest` 上;[原生 Windows 作业](2026-08-08-native-windows-pull-request-ci.zh.md)运行在托管的 `dsh-windows-2025-16core` 大型运行器上。当企业池发生故障——作业无限排队或企业标签消失——所有开启的拉取请求都无法合并,而"合并一个修复"这一常规恢复手段本身正被那些无法运行的必需检查死锁。**适用范围:两个独立开关,每个平台一个。**`DSH_CI_FAILOVER_LINUX` 恢复企业级 Linux 池故障(三个必需的 Linux 工作作业加 `all checks passed` 判定作业);`DSH_CI_FAILOVER_WINDOWS` 恢复托管 Windows 池故障(原生 Windows 作业)。Linux 池故障无需重定向 Windows 作业,反之亦然。判定作业的其余必需依赖(`node-compat`、`python-sdk`、`windows`)按设计留在标准托管运行器上(可移植边界);若更大范围的 GitHub 托管容量故障连标准池一并击倒,这些依赖仍会阻塞 `all checks passed`。因此故障需要一个任何具备仓库写权限的响应者都能在不合并任何代码的情况下触发的开关。
+[CI](../../../../.github/workflows/ci.yml) 中三个必需的 Linux 工作作业(`node 24 / static`、`node 24 / coverage`、`node 24 / snapshots and artifacts`)运行在托管的企业级 32 核池上;聚合它们的必需判定作业(`all checks passed`)运行在标准 `ubuntu-latest` 上;[原生 Windows 作业](2026-08-08-native-windows-pull-request-ci.zh.md)运行在托管的 `dsh-windows-2025-16core` 大型运行器上。当企业池发生故障——作业无限排队或企业标签消失——所有开启的拉取请求都无法合并,而"合并一个修复"这一常规恢复手段本身正被那些无法运行的必需检查死锁。**适用范围:两个独立开关,每个平台一个。**`DSH_CI_FAILOVER_LINUX` 恢复企业级 Linux 池故障(三个必需的 Linux 工作作业加 `all checks passed` 判定作业);`DSH_CI_FAILOVER_WINDOWS` 恢复托管 Windows 池故障(原生 Windows 作业)。Linux 池故障无需重定向 Windows 作业,反之亦然。判定作业的其余必需依赖(`node-24-bench`、`node-compat`、`python-sdk`、`windows`)按设计留在标准托管运行器上(可移植边界);若更大范围的 GitHub 托管容量故障连标准池一并击倒,这些依赖仍会阻塞 `all checks passed`。因此故障需要一个任何具备仓库写权限的响应者都能在不合并任何代码的情况下触发的开关。
 
 ## 决策
 

+ 2 - 2
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
-2026-09-04-session-open-performance-gate.md: 30e65eb52dea48426cf87cf91053a9133ab42763
-2026-09-04-session-open-performance-gate.zh.md: 3e202c966ea4764e135c79e1d238af707dca1f8d
+2026-09-04-session-open-performance-gate.md: 2820c9d7d0e5b7d9382c7f8d6540154440175f26
+2026-09-04-session-open-performance-gate.zh.md: 965b9035074504870bcb2f1ca8166962c264d75a

+ 5 - 3
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md

@@ -12,7 +12,7 @@ Measuring only `SessionPersistence.open()` does not stably describe the result f
 
 ## Decision
 
-Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench`. The private `@deepseek-ai/dsh-benchmarks` workspace owns benchmark-only dependencies. The command first builds workspace libraries and dedicated workers under `benchmarks/.dsh-build/`, then invokes `vitest.bench.config.ts`. The job runs the benchmark lane alone; Vitest runs one file at a time and only prepares input, starts measurement children, aggregates results, and enforces budgets. Every timed CPU path executes compiled JavaScript under plain Node with `NODE_OPTIONS` removed and no TypeScript loader; bare workspace imports therefore resolve from `benchmarks/node_modules` through package exports to built `lib/` entries.
+Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench`. The private `@deepseek-ai/dsh-benchmarks` workspace owns benchmark-only dependencies. The command first builds workspace libraries and dedicated workers under `benchmarks/.dsh-build/`, then invokes `vitest.bench.config.ts`. The [standard hosted runner decision](2026-09-06-standard-hosted-benchmark-runner.md) owns runner selection and the outer job timeout. The job runs the benchmark lane alone; Vitest runs one file at a time and only prepares input, starts measurement children, aggregates results, and enforces budgets. Every timed CPU path executes compiled JavaScript under plain Node with `NODE_OPTIONS` removed and no TypeScript loader; bare workspace imports therefore resolve from `benchmarks/node_modules` through package exports to built `lib/` entries.
 
 Required performance gates live under top-level `benchmarks/`, grouped by measured user path rather than package ownership. Host files use `*.bench.ts`, Client-face files use `*.bench.client.ts`, and scenario-specific workers and fixtures stay beside their benchmark without a benchmark suffix. Package-local `.perf.ts` files remain non-gating diagnostics; `scripts/` owns orchestration rather than benchmark cases.
 
@@ -37,7 +37,7 @@ Normal-heap mode performs a fixed pair of explicit garbage collections after Hos
 
 The performance gate does not duplicate semantic assertions owned by functional tests; it requires only that the target call completes and reaches its measured endpoint. The Client-fold benchmark continues to use the real `ConversationNodeAssembler` and every Chat Definition, and requires both the large window's absolute time and its scaling relative to the small window to remain below fixed budgets.
 
-Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
+Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Except for current-generation `open`, source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. Current-generation `open` uses a directly measured standard-runner expectation of 50 ms with only the 1.25× headroom, rounded up to a 63 ms budget. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
 
 ## Calibration evidence
 
@@ -54,12 +54,14 @@ Five-sample medians on the same Node 24 reference machine establish the positive
 
 The pre-stack implementation keeps V0 as its current format, so first open does not change its on-disk representation; its native V0 first-history and Agent-resume measurements therefore apply to both lifecycle rows.
 
+The [standard two-CPU run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384/job/101461539961) at `ca3ffe95dac2c55eefeb16ed9b61067bbd19ee90` uses Node 24.20.0 x64 and Ubuntu image `20260831.293.1`. Its five current-generation `open` samples are 49.2, 47.4, 49.1, 48.6, and 48.1 ms: median 48.6 ms, maximum 49.2 ms. The rounded 50 ms CI expectation gives a 63 ms limit without reapplying the 2× machine scale. The log identifies two available CPUs but not their model; it does not isolate hardware from the Node-version change. This is endpoint-specific runner calibration, not evidence of an application optimization or a new reference-machine measurement. Every other benchmark passes its existing budget. Deterministic controls reject the observed median at the historical 30 ms limit, accept it at 63 ms, reject a synthetic 75 ms reopen median, and reject a synthetic 4,000 ms first-open duration at its unchanged 550 ms limit. These controls verify budget enforcement, not a measured new regression.
+
 The calibrated source budgets are:
 
 | Measurement | Reference expectation | CI budget |
 |---|---:|---:|
 | First-open `open` | 220 ms | 550 ms |
-| Current-generation `open` | 12 ms | 30 ms |
+| Current-generation `open` | 12 ms (historical reference; CI expectation: 50 ms) | 63 ms |
 | Complete read | 8 ms | 20 ms |
 | Session restore | 24 ms | 60 ms |
 | Projection | 14 ms | 35 ms |

+ 5 - 3
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md

@@ -12,7 +12,7 @@ Session format v2 的推出改变了两条成本随模型输出增长的路径
 
 ## 决定
 
-Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench`。私有 `@deepseek-ai/dsh-benchmarks` workspace 拥有 benchmark 专属依赖。该命令先构建 workspace library 和 `benchmarks/.dsh-build/` 下的专用 worker,再调用 `vitest.bench.config.ts`。该 job 单独运行 benchmark lane;Vitest 逐文件运行,只负责准备输入、启动测量子进程、汇总结果和执行预算断言。每条被计时的 CPU 路径都以纯 Node 执行编译后的 JavaScript,并移除 `NODE_OPTIONS` 且不加载 TypeScript runtime;workspace 裸导入因此从 `benchmarks/node_modules` 通过 package exports 解析到构建后的 `lib/` 入口。
+Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench`。私有 `@deepseek-ai/dsh-benchmarks` workspace 拥有 benchmark 专属依赖。该命令先构建 workspace library 和 `benchmarks/.dsh-build/` 下的专用 worker,再调用 `vitest.bench.config.ts`。[标准托管运行器决策](2026-09-06-standard-hosted-benchmark-runner.zh.md)拥有运行器选择及外层 job 超时。该 job 单独运行 benchmark lane;Vitest 逐文件运行,只负责准备输入、启动测量子进程、汇总结果和执行预算断言。每条被计时的 CPU 路径都以纯 Node 执行编译后的 JavaScript,并移除 `NODE_OPTIONS` 且不加载 TypeScript runtime;workspace 裸导入因此从 `benchmarks/node_modules` 通过 package exports 解析到构建后的 `lib/` 入口。
 
 必需性能 gate 位于顶层 `benchmarks/`,按被测用户路径而非 package 归属组织。Host 文件使用 `*.bench.ts`,Client 面文件使用 `*.bench.client.ts`,场景专属 worker 与 fixture 留在对应 benchmark 旁且不带 benchmark 后缀。包内 `.perf.ts` 文件仍是非门禁诊断;`scripts/` 负责编排而不承载 benchmark case。
 
@@ -37,7 +37,7 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 性能 gate 不重复功能测试的内容断言,只要求目标调用完成并到达对应的可观察终点。Client fold benchmark 继续使用真实 `ConversationNodeAssembler` 与全部 Chat Definition,要求大窗口的绝对时间和相对小窗口的缩放比均低于固定预算。
 
-预算按各测量终点分别校准。两次 Node 24.19 x64 CI 运行的中位数最大相差 5.2%;其 CPU 密集型壁钟时间是 Node 24.18 arm64 参考运行的 1.95–2.06 倍。源码常量记录参考机器上的预期耗时;`ciTimeBudget()` 将其乘以实测的 2 倍 CI 时间系数和 1.25 倍波动余量。GC 后增量堆与 Client fold 缩放预算不属于壁钟时间,因此只使用 1.25 倍余量。128 MB 完成性检查仍是独立的瞬时分配限制。由此得到的 first-open 时间上限、受限堆检查与 Client fold 上限都会拒绝已知退化。栈前参考提交固定为 `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5`,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
+预算按各测量终点分别校准。两次 Node 24.19 x64 CI 运行的中位数最大相差 5.2%;其 CPU 密集型壁钟时间是 Node 24.18 arm64 参考运行的 1.95–2.06 倍。除当前 generation `open` 外,源码常量记录参考机器上的预期耗时;`ciTimeBudget()` 将其乘以实测的 2 倍 CI 时间系数和 1.25 倍波动余量。当前 generation `open` 使用标准运行器直接测得的 50 ms 预期值,仅乘以 1.25 倍余量,向上取整得到 63 ms 预算。GC 后增量堆与 Client fold 缩放预算不属于壁钟时间,因此只使用 1.25 倍余量。128 MB 完成性检查仍是独立的瞬时分配限制。由此得到的 first-open 时间上限、受限堆检查与 Client fold 上限都会拒绝已知退化。栈前参考提交固定为 `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5`,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
 
 ## 校准证据
 
@@ -54,12 +54,14 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 栈前实现以 V0 作为当前格式,因此 first open 不改变磁盘表示;它的原生 V0 首屏历史与 Agent resume 测量同时适用于两个生命周期行。
 
+`ca3ffe95dac2c55eefeb16ed9b61067bbd19ee90` 上的[标准双 CPU 运行](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384/job/101461539961)使用 Node 24.20.0 x64 和 Ubuntu 镜像 `20260831.293.1`。当前 generation `open` 的五次样本为 49.2、47.4、49.1、48.6 和 48.1 ms:中位数 48.6 ms,最大值 49.2 ms。取整后的 50 ms CI 预期值给出 63 ms 上限,不重复乘以 2 倍机器系数。日志标明两个可用 CPU,但未记录型号;它无法区分硬件变化与 Node 版本变化的影响。这是端点专属的运行器校准,不是应用优化或参考机器新测量的证据。其他每项 benchmark 均通过既有预算。确定性正反例在历史 30 ms 上限下拒绝实测中位数,在 63 ms 下接受它,拒绝合成的 75 ms reopen 中位数,并以未改变的 550 ms 上限拒绝合成的 4,000 ms 首次打开耗时。这些正反例验证预算执行,不代表测得新的退化。
+
 校准后的源码预算如下:
 
 | 测量项 | 参考机预期 | CI 预算 |
 |---|---:|---:|
 | First-open `open` | 220 ms | 550 ms |
-| 当前 generation `open` | 12 ms | 30 ms |
+| 当前 generation `open` | 12 ms(历史参考值;CI 预期值:50 ms) | 63 ms |
 | 完整 read | 8 ms | 20 ms |
 | Session restore | 24 ms | 60 ms |
 | Projection | 14 ms | 35 ms |

+ 6 - 0
.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.md
+2026-09-06-standard-hosted-benchmark-runner.md: af95af6ee8cef1128a5475863ea7a1aaf66f30b9
+2026-09-06-standard-hosted-benchmark-runner.zh.md: 47095a5b2e91318c2f3ec5ac3cc9f6df3bce04d6

+ 26 - 0
.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.md

@@ -0,0 +1,26 @@
+# Agent Note: Standard hosted runner for required benchmarks
+
+Status: implemented
+
+English | [中文](2026-09-06-standard-hosted-benchmark-runner.zh.md)
+
+## Problem
+
+Wall-clock performance checks need an isolated execution lane and a consistent runner class. Routing them through the enterprise Linux failover switch makes their measurements depend on either larger hosted capacity or a shared self-hosted VM, while also consuming capacity needed by parallel correctness checks.
+
+## Decision
+
+The required benchmark job in [ci.yml](../../../../.github/workflows/ci.yml) uses the standard GitHub-hosted `ubuntu-24.04` runner independently of Linux failover. It always attempts to restore the pnpm store cache and retains a standalone benchmark lane. The complete job has a 15-minute timeout covering setup, installation, builds, and measurements. This bounds infrastructure execution, not an individual performance assertion.
+
+The [Session performance decision](2026-09-04-session-open-performance-gate.md) continues to own workloads, timing and memory budgets, worker isolation, and calibration. Only current-generation `open` uses an endpoint-specific 50 ms standard-runner expectation with the existing 1.25× headroom, giving a 63 ms limit. All other performance budgets and the worker, test, and hook deadlines remain unchanged. Successful raw measurements remain in the Actions log through step-local `DSH_GATE_VERBOSE=1`. The hardware-comparison workflows retain their deliberately different runner sizes.
+
+## Alternatives considered
+
+- Enterprise or shared self-hosted routing retains more build capacity but ties the measurement environment to unrelated failover operations.
+- Increasing performance thresholds without endpoint measurements conflates a bounded CI execution with a regression allowance. Threshold changes require measured calibration and positive and negative controls.
+
+## Consequences
+
+A standard runner trades parallel build capacity for a fixed measurement class without removing the required verdict. Cache misses and runner variation can still affect total duration. Each runner change needs an actual hosted benchmark run before its job timeout is treated as validated; local workflow assertions alone cannot establish execution time.
+
+The owning [workflow tests](../../../../scripts/ci-workflow.spec.ts) pin runner routing, unconditional cache restoration, required status, and the job timeout. Negative controls reject failover routing, a cache condition, and the former 30-minute job bound.

+ 26 - 0
.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.zh.md

@@ -0,0 +1,26 @@
+# Agent Note: 必需 benchmark 使用标准托管运行器
+
+Status: implemented
+
+[English](2026-09-06-standard-hosted-benchmark-runner.md) | 中文
+
+## 问题
+
+壁钟性能检查需要独立执行的 lane 和一致的运行器类别。通过企业 Linux 故障转移开关路由这些检查,会让测量取决于大型托管运行器或共享自托管虚拟机,同时占用并行正确性检查所需的容量。
+
+## 决定
+
+[ci.yml](../../../../.github/workflows/ci.yml) 中的必需 benchmark job 使用标准 GitHub 托管 `ubuntu-24.04` 运行器,不受 Linux 故障转移影响。它始终尝试恢复 pnpm 存储缓存,并保留独立的 benchmark lane。整个 job 的超时为 15 分钟,覆盖准备、安装、构建和测量。这限制的是基础设施执行时间,而非单项性能断言。
+
+[Session 性能决策](2026-09-04-session-open-performance-gate.zh.md) 继续拥有工作负载、时间和内存预算、worker 隔离及校准。仅当前 generation `open` 使用端点专属的 50 ms 标准运行器预期值,乘以既有 1.25 倍余量后得到 63 ms 上限。其他性能预算以及 worker、测试和钩子的截止时间均保持不变。步骤级 `DSH_GATE_VERBOSE=1` 使成功运行的原始测量保留在 Actions 日志中。硬件比较工作流保留有意设置的不同运行器规格。
+
+## 考虑过的替代方案
+
+- 企业或共享自托管路由保留更多构建容量,但使测量环境受无关故障转移操作影响。
+- 没有端点测量就提高性能阈值,会混淆有界 CI 执行与退化容许量。阈值调整需要实测校准及正反例。
+
+## 后果
+
+标准运行器以并行构建容量换取固定测量类别,不移除必需判定。缓存未命中和运行器波动仍会影响总耗时。每次更换运行器都需要实际托管 benchmark 运行,才能认定 job 超时经过验证;本地工作流断言无法单独证明执行耗时。
+
+所属[工作流测试](../../../../scripts/ci-workflow.spec.ts) 固定运行器路由、无条件缓存恢复、必需状态及 job 超时。反例验证拒绝故障转移路由、缓存条件和原来的 30 分钟 job 上限。

+ 4 - 7
.github/workflows/ci.yml

@@ -158,15 +158,11 @@ jobs:
 
   node-24-bench:
     if: github.event_name == 'pull_request'
-    runs-on: >-
-      ${{ vars.DSH_CI_FAILOVER_LINUX == 'selfhosted'
-          && github.event.pull_request.user.login != 'dependabot[bot]'
-          && fromJSON('["self-hosted", "linux", "x64", "vm-backup"]')
-          || 'dsh-ubuntu-24-04-16core' }}
+    runs-on: ubuntu-24.04
     name: node 24 / benchmarks
     # Wall-clock budgets need an otherwise idle runner, so this job runs the
     # benchmark lane alone instead of joining a concurrent gate aggregate.
-    timeout-minutes: 30
+    timeout-minutes: 15
     steps:
       - uses: actions/checkout@v6
         with:
@@ -189,7 +185,6 @@ jobs:
           echo "path=$store_path" >> "$GITHUB_OUTPUT"
 
       - uses: actions/cache/restore@v4
-        if: vars.DSH_CI_FAILOVER_LINUX != 'selfhosted' || github.event.pull_request.user.login == 'dependabot[bot]'
         with:
           path: ${{ steps.pnpm-store.outputs.path }}
           key: ${{ runner.os }}-node-${{ env.PRIMARY_NODE_VERSION }}-pnpm-${{ hashFiles('pnpm-lock.yaml') }}
@@ -200,6 +195,8 @@ jobs:
         run: pnpm install --frozen-lockfile
 
       - name: Run performance benchmarks
+        env:
+          DSH_GATE_VERBOSE: '1'
         run: pnpm run check:ci:bench
 
   node-24-consumers:

+ 26 - 3
benchmarks/session-open/session-open.bench.ts

@@ -45,7 +45,6 @@ const SOURCE_GENERATION_BY_ACCESS = {
 /** Expected durations on the reference machine before CI scaling and variance headroom. */
 const EXPECTED_MS = {
   migrationOpen: 220,
-  reopenOpen: 12,
   read: 8,
   sessionRestore: 24,
   projection: 14,
@@ -56,7 +55,9 @@ const EXPECTED_MS = {
 } as const
 
 const MIGRATION_OPEN_BUDGET_MS = ciTimeBudget(EXPECTED_MS.migrationOpen)
-const REOPEN_OPEN_BUDGET_MS = ciTimeBudget(EXPECTED_MS.reopenOpen)
+/** Standard two-CPU CI reopen samples span 47.4–49.2 ms; 50 ms is the rounded expectation. */
+const EXPECTED_REOPEN_CI_MS = 50
+const REOPEN_OPEN_BUDGET_MS = Math.ceil(EXPECTED_REOPEN_CI_MS * PERFORMANCE_BUDGET_HEADROOM)
 const READ_BUDGET_MS = ciTimeBudget(EXPECTED_MS.read)
 const SESSION_RESTORE_BUDGET_MS = ciTimeBudget(EXPECTED_MS.sessionRestore)
 const PROJECTION_BUDGET_MS = ciTimeBudget(EXPECTED_MS.projection)
@@ -270,6 +271,28 @@ const ACCESS_BENCHMARKS: readonly AccessBenchmarkSpec[] = [
   },
 ]
 
+function expectOpenWithinBudget(value: number, budget: number): void {
+  expect(value).toBeLessThanOrEqual(budget)
+}
+
+describe('standard hosted reopen calibration', () => {
+  it('accepts the recorded two-CPU samples that exceed the historical budget', () => {
+    const recordedMedian = median([49.2, 47.4, 49.1, 48.6, 48.1])
+
+    expect(recordedMedian).toBe(48.6)
+    expect(() => expectOpenWithinBudget(recordedMedian, ciTimeBudget(12))).toThrow()
+    expectOpenWithinBudget(recordedMedian, REOPEN_OPEN_BUDGET_MS)
+    expect(REOPEN_OPEN_BUDGET_MS).toBe(63)
+  })
+
+  it('rejects synthetic reopen and multi-second first-open regressions', () => {
+    const regressionMedian = median([74, 75, 76, 75, 74])
+    expect(() => expectOpenWithinBudget(regressionMedian, REOPEN_OPEN_BUDGET_MS)).toThrow()
+    expect(MIGRATION_OPEN_BUDGET_MS).toBe(550)
+    expect(() => expectOpenWithinBudget(4_000, MIGRATION_OPEN_BUDGET_MS)).toThrow()
+  })
+})
+
 describe('opening a large Session for first open and post-upgrade reopen', () => {
   const suite = new SessionOpenBenchmarkSuite()
 
@@ -291,7 +314,7 @@ describe('opening a large Session for first open and post-upgrade reopen', () =>
             projection: PROJECTION_BUDGET_MS,
           },
         }))
-        expect(result.openMs.median).toBeLessThanOrEqual(access.openBudgetMs)
+        expectOpenWithinBudget(result.openMs.median, access.openBudgetMs)
         expect(result.readMs.median).toBeLessThanOrEqual(READ_BUDGET_MS)
         expect(result.sessionRestoreMs.median).toBeLessThanOrEqual(SESSION_RESTORE_BUDGET_MS)
         expect(result.projectionMs.median).toBeLessThanOrEqual(PROJECTION_BUDGET_MS)

+ 2 - 2
docs/development.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/development.md
-development.md: a57c99d606a73cb938f339e080ab0ba05913902a
-development.zh.md: 5439fec59cb7c245d655393690bf6c848cdf20fa
+development.md: 37028720ff2487a81caecb8b2e2c6e5bc26b2df5
+development.zh.md: 01e3e88605cb4f9ea0b14df4a0099f436c48c5e9

+ 1 - 1
docs/development.md

@@ -120,7 +120,7 @@ Contributors can opt into the comprehensive local gate set with `pnpm run check:
 
 ### CI gates
 
-The keyless [CI workflow](../.github/workflows/ci.yml) groups independent gates into broad lanes and runs a smaller compatibility signal across supported Node versions. Artifact consumers wait for one build within their lane. The separate real-API workflow runs `pnpm run test:e2e` with its configured worker bound. See [scripts/run-gates.ts](../scripts/run-gates.ts) and the workflow files for the current gate and job inventory.
+The keyless [CI workflow](../.github/workflows/ci.yml) groups independent gates into broad lanes and runs a smaller compatibility signal across supported Node versions. Artifact consumers wait for one build within their lane. Required benchmarks run separately on standard GitHub-hosted Linux; the [benchmark runner decision](../.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.md) owns routing and the job timeout. The separate real-API workflow runs `pnpm run test:e2e` with its configured worker bound. See [scripts/run-gates.ts](../scripts/run-gates.ts) and the workflow files for the current gate and job inventory.
 
 The credential-free dsh dependency-layout and dsh/vendor pack rehearsals use the existing Linux self-hosted pool only when `DSH_CI_FAILOVER_LINUX=selfhosted` and the event is a trusted master push or same-repository, non-fork, non-Dependabot pull request. All other cases, including manual dispatch, use `ubuntu-24.04`; manual publication stays hosted. See the [release rehearsal runner decision](../.agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.md) for persistent-store isolation and fallback limits.
 

+ 1 - 1
docs/development.zh.md

@@ -124,7 +124,7 @@ vendor manifest 守卫检查 `vendor/*/src` 下的改动是否连同对应的 `v
 
 ### CI 门禁
 
-keyless [CI 工作流](../.github/workflows/ci.yml) 将独立门禁分组到若干宽粒度 lane,并在受支持的 Node 版本上运行一组较小的兼容性检查。产物消费方在各自 lane 内等待一次 build。单独的真实 API 工作流按其配置的 worker 上限运行 `pnpm run test:e2e`。当前门禁和 job 清单以 [scripts/run-gates.ts](../scripts/run-gates.ts) 和工作流文件为准。
+keyless [CI 工作流](../.github/workflows/ci.yml) 将独立门禁分组到若干宽粒度 lane,并在受支持的 Node 版本上运行一组较小的兼容性检查。产物消费方在各自 lane 内等待一次 build。必需 benchmark 在标准 GitHub 托管 Linux 上独立运行;[benchmark 运行器决策](../.agents/notes/implemented/testing/2026-09-06-standard-hosted-benchmark-runner.zh.md)拥有路由及 job 超时。单独的真实 API 工作流按其配置的 worker 上限运行 `pnpm run test:e2e`。当前门禁和 job 清单以 [scripts/run-gates.ts](../scripts/run-gates.ts) 和工作流文件为准。
 
 不带凭据的 dsh 依赖布局检查与 dsh/vendor 打包演练仅在 `DSH_CI_FAILOVER_LINUX=selfhosted`,且事件为受信任的 master 推送或同仓库、非 fork、非 Dependabot 拉取请求时使用现有 Linux 自托管池。其余情况(包括手动触发)均使用 `ubuntu-24.04`;手动发布仍使用托管运行器。持久化存储隔离与回退限制见[发布演练运行器决策](../.agents/notes/implemented/process/2026-09-06-release-rehearsal-selfhosted.zh.md)。
 

+ 2 - 2
python/sdk-runtime/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write python/sdk-runtime/README.md
-README.md: 050ae85d9b0a38b82a3c66a84c3d8f34e137be6c
-README.zh.md: 7066d7224752294c25b58cfe8fb6a94524a2013d
+README.md: fb7478f305015e60dafd23861ab6fd91f10757f3
+README.zh.md: 7c663aba8d387fb7b1048afe28d50faee7350303

+ 1 - 1
python/sdk-runtime/README.md

@@ -19,7 +19,7 @@ Both carriers execute the same `dsh` grammar and shipped profiles, including the
 - `bundled_package_dir() -> Path` returns the installed module-data root and verifies its release metadata.
 - `bundled_runtime_path() -> Path` returns the current platform executable and verifies required sidecars.
 - `resolve_bundled_launch_args(mode=None) -> tuple[str, ...]` returns the executable argv by default. Explicit `mode="node"` or `DSH_RUNTIME_MODE=node` selects the repo-only Node carrier.
-- `main()` implements the installed `dsh` console command and rejects an absent or blank `DSH_HOME` before replacing the Python process.
+- `main()` implements the installed `dsh` console command and rejects an absent or blank `DSH_HOME`. On Windows it waits for the bundled process with inherited standard streams and forwards its exit status; on POSIX it replaces the Python process.
 
 Unsupported platforms and missing executables or sidecars raise `FileNotFoundError` with the build and installation routes. Unknown runtime modes raise `ValueError`.
 

+ 1 - 1
python/sdk-runtime/README.zh.md

@@ -19,7 +19,7 @@ Wheel 会安装 `dsh` 控制台命令和 `deepseek_harness_runtime` Python 模
 - `bundled_package_dir() -> Path` 返回已安装模块数据根目录,并校验发布元数据。
 - `bundled_runtime_path() -> Path` 返回当前平台可执行程序,并校验必需伴随文件。
 - `resolve_bundled_launch_args(mode=None) -> tuple[str, ...]` 默认返回可执行程序 argv。显式 `mode="node"` 或 `DSH_RUNTIME_MODE=node` 会选择仅限仓库使用的 Node 载体。
-- `main()` 实现已安装的 `dsh` 控制台命令,并在替换 Python 进程前拒绝缺失或空白的 `DSH_HOME`。
+- `main()` 实现已安装的 `dsh` 控制台命令,并拒绝缺失或空白的 `DSH_HOME`。在 Windows 上,它让打包进程继承标准流,等待其结束并转发退出状态;在 POSIX 上,它替换 Python 进程。
 
 不支持的平台以及缺失的可执行程序或伴随文件会抛出 `FileNotFoundError`,并指出构建与安装路径。未知运行时模式会抛出 `ValueError`。
 

+ 5 - 1
python/sdk-runtime/src/deepseek_harness_runtime/__init__.py

@@ -24,6 +24,7 @@ from __future__ import annotations
 import os
 import platform
 import shutil
+import subprocess
 import sys
 from pathlib import Path
 
@@ -156,7 +157,7 @@ def _node_launch_args() -> tuple[str, str]:
 
 
 def main() -> None:
-    """Execute the bundled dsh CLI with an explicitly selected Harness home."""
+    """Launch the CLI with explicit DSH_HOME; wait on Windows, replace the process on POSIX."""
     if not os.environ.get("DSH_HOME", "").strip():
         print(
             "dsh: the Python runtime command requires an explicit DSH_HOME; "
@@ -165,6 +166,9 @@ def main() -> None:
         )
         raise SystemExit(2)
     argv = (*resolve_bundled_launch_args(), *sys.argv[1:])
+    if sys.platform == "win32":
+        # Windows CRT exec does not replace the process; wait and preserve the runtime status.
+        raise SystemExit(subprocess.run(argv, env=os.environ).returncode)
     os.execvpe(argv[0], argv, os.environ)
 
 

+ 53 - 1
python/sdk/tests/test_runtime_resolution.py

@@ -2,7 +2,11 @@
 
 from __future__ import annotations
 
+import os
+import subprocess
+import sys
 from pathlib import Path
+from types import SimpleNamespace
 
 import deepseek_harness_runtime as runtime
 import pytest
@@ -129,7 +133,7 @@ def test_python_dsh_command_executes_the_bundled_cli(
     called: dict[str, object] = {}
     monkeypatch.setenv("DSH_HOME", "/explicit/home")
     monkeypatch.setattr(runtime, "resolve_bundled_launch_args", lambda: ("/runtime",))
-    monkeypatch.setattr(runtime.sys, "argv", ["dsh", "plugin", "--profile", "sdk", "list"])
+    monkeypatch.setattr(runtime, "sys", SimpleNamespace(platform="linux", argv=["dsh", "plugin", "--profile", "sdk", "list"]))
 
     def execvpe(file: str, args: tuple[str, ...], env: dict[str, str]) -> None:
         called.update(file=file, args=args, home=env.get("DSH_HOME"))
@@ -143,3 +147,51 @@ def test_python_dsh_command_executes_the_bundled_cli(
         "args": ("/runtime", "plugin", "--profile", "sdk", "list"),
         "home": "/explicit/home",
     }
+
+
+@pytest.mark.parametrize("returncode", [0, 37, 513])
+def test_windows_console_waits_and_forwards_runtime_status(monkeypatch: pytest.MonkeyPatch, returncode: int) -> None:
+    monkeypatch.setenv("DSH_HOME", "/explicit/home")
+    monkeypatch.setattr(runtime, "sys", SimpleNamespace(platform="win32", argv=["dsh", "plugin", "argument with spaces", "中文"]))
+    monkeypatch.setattr(runtime, "resolve_bundled_launch_args", lambda: ("runtime.exe",))
+    called = []
+
+    def run(args: tuple[str, ...], **kwargs: object) -> subprocess.CompletedProcess[str]:
+        called.append((args, kwargs))
+        return subprocess.CompletedProcess(args, returncode)
+
+    def forbidden_exec(*args: object) -> None:
+        pytest.fail("Windows console must wait instead of entering CRT exec")
+
+    monkeypatch.setattr(subprocess, "run", run)
+    monkeypatch.setattr(runtime.os, "execvpe", forbidden_exec)
+    with pytest.raises(SystemExit) as result:
+        main()
+    assert result.value.code == returncode
+    assert called == [(("runtime.exe", "plugin", "argument with spaces", "中文"), {"env": os.environ})]
+
+
+@pytest.mark.parametrize("returncode", [0, 37, pytest.param(513, marks=pytest.mark.skipif(sys.platform != "win32", reason="POSIX truncates process exit codes to eight bits"))])
+def test_windows_console_branch_preserves_real_child_io_and_completion(tmp_path: Path, returncode: int) -> None:
+    child = tmp_path / "child with spaces.py"
+    sentinel = tmp_path / "finished"
+    child.write_text(
+        "import pathlib,sys\n"
+        "assert sys.argv[1] == 'argument with spaces'\n"
+        "assert sys.argv[2] == '中文'\n"
+        "print('stdout-中文', flush=True)\n"
+        "print('stderr-中文', file=sys.stderr, flush=True)\n"
+        f"pathlib.Path({str(sentinel)!r}).write_text('done')\n"
+        f"raise SystemExit({returncode})\n", encoding="utf-8",
+    )
+    driver = (
+        "import deepseek_harness_runtime as runtime; from types import SimpleNamespace; "
+        f"runtime.sys = SimpleNamespace(platform='win32', argv=['dsh', 'argument with spaces', '中文']); "
+        f"runtime.resolve_bundled_launch_args = lambda: ({sys.executable!r}, {str(child)!r}); runtime.main()"
+    )
+    result = subprocess.run([sys.executable, "-c", driver], capture_output=True, text=True, encoding="utf-8",
+                            env={**os.environ, "DSH_HOME": str(tmp_path), "PYTHONIOENCODING": "utf-8"}, timeout=15)
+    assert result.returncode == returncode, result.stderr
+    assert result.stdout == "stdout-中文\n"
+    assert result.stderr == "stderr-中文\n"
+    assert sentinel.read_text() == "done"

+ 15 - 0
python/sdk/tests/test_smoke_model.py

@@ -1,6 +1,7 @@
 from __future__ import annotations
 
 import runpy
+import subprocess
 from pathlib import Path
 
 import pytest
@@ -262,3 +263,17 @@ def test_snapshot_generation_filename_must_match_header(tmp_path: Path) -> None:
 
     with pytest.raises(AssertionError, match="filename declares Session format v1"):
         SMOKE["selected_snapshot_session_files"](tmp_path)
+
+
+@pytest.mark.parametrize("returncode", [1, -1073741819, 3221225477])
+def test_profile_plugin_failure_reports_native_exit_status(monkeypatch: pytest.MonkeyPatch, returncode: int) -> None:
+    def failed_install(*args: object, **kwargs: object) -> subprocess.CompletedProcess[str]:
+        return subprocess.CompletedProcess(args=[], returncode=returncode, stdout="", stderr="")
+
+    monkeypatch.setattr(subprocess, "run", failed_install)
+    with pytest.raises(AssertionError) as error:
+        SMOKE["smoke_sdk_profile_plugin"]("http://127.0.0.1:1")
+    message = str(error.value)
+    assert f"returncode={returncode}" in message
+    assert f"0x{returncode & 0xffffffff:08x}" in message
+    assert "stdout='' stderr=''" in message

+ 41 - 2
scripts/ci-workflow.spec.ts

@@ -235,10 +235,10 @@ describe('CI workflow', () => {
     expect(aggregate.needs).not.toContain('windows-observational')
     expect(aggregate.needs).not.toContain('serial-windows')
 
-    // Linux failover is a separate switch: the four required Linux workers
+    // Linux failover is a separate switch: the three enterprise Linux workers
     // and the verdict job resolve their pool through DSH_CI_FAILOVER_LINUX,
     // never the Windows switch.
-    for (const [jobName, job] of [['node-24', node24], ['node-24-coverage', node24Coverage], ['node-24-bench', node24Bench], ['node-24-consumers', node24Consumers]] as const) {
+    for (const [jobName, job] of [['node-24', node24], ['node-24-coverage', node24Coverage], ['node-24-consumers', node24Consumers]] as const) {
       expect(typeof job['runs-on']).toBe('string')
       expect(job['runs-on'], `${jobName} runs-on must use the Linux failover switch`).toContain('DSH_CI_FAILOVER_LINUX')
       expect(job['runs-on'], `${jobName} runs-on must not use the Windows failover switch`).not.toContain('DSH_CI_FAILOVER_WINDOWS')
@@ -269,6 +269,45 @@ describe('CI workflow', () => {
     expect(windowsObservational.env).not.toMatchObject({ DSH_GATE_FAIL_FAST: '1' })
   })
 
+  it('runs required benchmarks on standard hosted Linux independently of failover', () => {
+    const workflow = loadWorkflow('.github/workflows/ci.yml')
+    const benchmark = workflowJob(workflow, 'node-24-bench')
+    const aggregate = workflowJob(workflow, 'all-checks-passed')
+
+    expect(benchmark['runs-on']).toBe('ubuntu-24.04')
+    expect(benchmark.if).toBe("github.event_name == 'pull_request'")
+    expect(benchmark.needs).toBeUndefined()
+    expect(benchmark['continue-on-error']).toBeUndefined()
+    expect(benchmark.env).toBeUndefined()
+    expect(aggregate.needs).toContain('node-24-bench')
+  })
+
+  it('always restores the hosted benchmark pnpm cache', () => {
+    const benchmark = workflowJob(loadWorkflow('.github/workflows/ci.yml'), 'node-24-bench')
+    if (!Array.isArray(benchmark.steps)) throw new TypeError('benchmark job must define steps')
+    const caches = benchmark.steps.filter(step => isRecord(step) && step.uses === 'actions/cache/restore@v4')
+
+    expect(caches).toHaveLength(1)
+    expect(caches[0]).not.toHaveProperty('if')
+    expect(caches[0]).toMatchObject({
+      with: {
+        path: '${{ steps.pnpm-store.outputs.path }}',
+        key: "${{ runner.os }}-node-${{ env.PRIMARY_NODE_VERSION }}-pnpm-${{ hashFiles('pnpm-lock.yaml') }}",
+      },
+    })
+  })
+
+  it('bounds the complete benchmark job to fifteen minutes', () => {
+    const benchmark = workflowJob(loadWorkflow('.github/workflows/ci.yml'), 'node-24-bench')
+
+    expect(benchmark['timeout-minutes']).toBe(15)
+    expect(benchmark.steps).toContainEqual({
+      name: 'Run performance benchmarks',
+      env: { DSH_GATE_VERBOSE: '1' },
+      run: 'pnpm run check:ci:bench',
+    })
+  })
+
   it('gives the Wine Host TypeScript compile the repository heap budget', () => {
     const wineGates = readFileSync(resolve(root, 'scripts/wine-windows-gates.sh'), 'utf8')
 

+ 1 - 0
scripts/smoke-python-runtime.py

@@ -1227,6 +1227,7 @@ def smoke_sdk_profile_plugin(base_url: str) -> None:
         if installed.returncode != 0:
             raise AssertionError(
                 f"Python-installed dsh could not add the external profile plugin: "
+                f"returncode={installed.returncode} (0x{installed.returncode & 0xffffffff:08x}) "
                 f"stdout={installed.stdout!r} stderr={installed.stderr!r}"
             )
         manifest = json.loads((dsh_home / "profiles" / "sdk" / "package.json").read_text())