Przeglądaj źródła

test(perf): gate opening large Sessions with synthetic CI benchmarks

Add a benchmark lane (`vitest.bench.config.ts`, `pnpm run test:bench`,
gate mode `ci-bench`) and a required `node 24 / benchmarks` CI job that
runs it alone. Benchmarks synthesize their input in-process from fixed
parameters and fail on documented budgets:

- `open-generation.bench.ts`: a 200-turn released-v0 log with 500 text
  and 125 reasoning deltas per reply (127,400 events, ~2.8 MB) encoded
  through the frozen v0 codec; the migrating first `open()` must finish
  within 2,000 ms in a child process capped at 128 MB of old space, and a
  fresh process must open the published current generation within 500 ms.
- `conversation-fold.bench.client.ts`: 200 replies whose compact streams
  hold 2,000 text + 500 reasoning deltas each, folded through every Chat
  Definition by the real assembler; the fold must finish within 150 ms
  and stay within 3x the fold of the same window with 100 deltas per
  reply.

On this commit both gates fail: the migration exhausts the 128 MB heap
(4.8 s and 696 MB peak RSS without the cap; the pre-stack decode of the
same bytes took 34 ms and 168 MB) and the fold scales 11x with the delta
count. The stacked fixes bring both paths to O(records).
Tianyi Cui 4 tygodni temu
rodzic
commit
2395f112d8

+ 6 - 0
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
+2026-09-04-session-open-performance-gate.md: 6115522ba676cb4caae93926e265c13df78aad74
+2026-09-04-session-open-performance-gate.zh.md: d7d71662cc6592bf3df0420028a5d8a3329c7a40

+ 38 - 0
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md

@@ -0,0 +1,38 @@
+# Agent Note: Required CI performance gate for opening large Sessions
+
+Status: implemented
+
+English | [中文](2026-09-04-session-open-performance-gate.zh.md)
+
+## Problem
+
+The Session format v2 rollout changed two paths whose cost scales with model output: the JSONL backend migrates and publishes a released-v0 log on its first `open()`, and the Client folds each settled reply's embedded compact stream. Neither path had an executed performance check, so a first open that grew from about 35 ms to about 5 s on a 127,400-event synthetic log (and from about 0.3 s to 26 s on a 575,000-chunk real log, with peak RSS of 2.7 GB and heap exhaustion under a 512 MB limit) and a Client fold that grew linearly with streamed deltas instead of compact records both reached master unnoticed. Unit tests use small logs, the coverage gate measures lines, and the existing `test:web:perf` inventory is a manual diagnostic outside CI.
+
+## Decision
+
+Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench` → `vitest.bench.config.ts`, which collects `packages/*/*/tests/**/*.bench.ts` and `*.bench.client.ts` and runs one file at a time. The job runs the benchmark lane alone, on the same runner selector and failover switch as the other required Linux workers, and joins the `all checks passed` verdict.
+
+Every benchmark synthesizes its input in-process from fixed parameters: numbered prompts, counter tokens, fixed timestamps. Recorded Sessions are never used because they carry user content, differ between machines, and drift as fixtures are re-recorded. Each benchmark documents its budget beside the constant that enforces it, and budgets follow three rules: a wall-clock budget sits a small multiple above the intended cost and well below the regression it guards; a memory budget runs the measured path in a child Node process under a fixed `--max-old-space-size`, so an allocation regression fails as an out-of-memory exit regardless of the runner's physical memory; and a scaling assertion compares two sizes of the same workload so a complexity regression fails on any host speed.
+
+The first two gates cover the two regressed paths:
+
+| Benchmark | Workload | Gates |
+|---|---|---|
+| `packages/session/session-persistence-jsonl/tests/open-generation.bench.ts` | 200 turns × (500 text + 125 reasoning deltas) = 127,400 released-v0 events, about 2.8 MB, encoded through the frozen v0 codec with packed rows | migrating first `open()` ≤ 2,000 ms under a 128 MB heap; fresh-process open of the published current generation ≤ 500 ms; minimum of three attempts |
+| `packages/client/ui-chat/tests/conversation-fold.bench.client.ts` | 200 replies whose compact streams hold 2,000 text + 500 reasoning deltas each (500,000 deltas in 1,600 records), folded through every Chat Definition by the real `ConversationNodeAssembler` | large fold ≤ 150 ms; large fold ≤ 3× the fold of the same window with 100 deltas per reply |
+
+Measured on the reference machine at the commit that introduced the gate, the migration benchmark exhausted the 128 MB heap and the fold benchmark scaled 11× between the small and large windows, so both gates fail on the regressed code and pass once the paths do O(records) work.
+
+## Alternatives considered
+
+**Extend the manual `test:web:perf` inventory.** Rejected: it stays outside CI by design, measures a simplified fold rather than the registered Definitions, and asserts nothing.
+
+**Time-only budgets.** Rejected: a single absolute budget either fails on slower runners or passes a regression on faster ones; the heap cap and the scaling ratio give host-independent verdicts, and the wall-clock budget remains as the timeout that the user-visible symptom is about.
+
+**Benchmark the real recorded corpus.** Rejected: corpus fixtures are small by policy, recorded material must not become a benchmark input, and their re-recording would silently move the baseline.
+
+**Run the benchmarks inside an existing gate aggregate.** Rejected: aggregates run gates concurrently on one runner, so wall-clock measurements would inherit the neighbours' CPU load.
+
+## Consequences
+
+Every pull request pays one more required Linux job of a few minutes, dominated by install time rather than the benchmarks themselves. A change that makes first open or the Client fold slower than its budget, heavier than its heap limit, or proportional to streamed deltas fails in the PR that introduces it, with the measured numbers printed in the job log. A budget change is a reviewed edit of the constant and its rationale comment, never an environment override, and a new benchmark must state which owner-visible path and which regression class it guards. The gate does not measure browser rendering, network transfer, or real recorded Sessions; those remain covered by the manual `test:web:perf` inventory and by review.

+ 38 - 0
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md

@@ -0,0 +1,38 @@
+# Agent Note: 打开大型 Session 的必需 CI 性能 gate
+
+Status: implemented
+
+[English](2026-09-04-session-open-performance-gate.md) | 中文
+
+## 问题
+
+Session format v2 的推出改变了两条成本随模型输出增长的路径:JSONL backend 在首次 `open()` 时迁移并发布 released-v0 log,Client 则 fold 每个已结算回复中嵌入的紧凑 stream。两条路径都没有可执行的性能检查,因此首次打开在 127,400 事件的合成 log 上从约 35 ms 增长到约 5 s(在 575,000 chunk 的真实 log 上从约 0.3 s 增长到 26 s,峰值 RSS 2.7 GB,并在 512 MB 堆限制下耗尽堆),以及 Client fold 随流式 delta 数而不是紧凑记录数线性增长,都未被察觉地进入了 master。单元测试使用小 log,coverage gate 只度量行数,现有的 `test:web:perf` 清单是 CI 之外的手动诊断。
+
+## 决定
+
+Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench` → `vitest.bench.config.ts`,后者收集 `packages/*/*/tests/**/*.bench.ts` 与 `*.bench.client.ts` 并逐文件运行。该 job 单独运行基准 lane,与其他必需 Linux worker 使用同一 runner 选择器和 failover 开关,并加入 `all checks passed` 判定。
+
+每个基准都在进程内按固定参数合成输入:编号的 prompt、计数 token、固定时间戳。绝不使用录制的 Session,因为它们携带用户内容、在不同机器上不同,并随 fixture 重新录制而漂移。每个基准在强制执行预算的常量旁记录其预算,预算遵循三条规则:壁钟预算取目标成本的小倍数并远低于所防护的回归;内存预算把被测路径放在固定 `--max-old-space-size` 的子 Node 进程中运行,使分配回归无论 runner 物理内存多大都以 out-of-memory 退出失败;缩放断言比较同一负载的两个规模,使复杂度回归在任何主机速度下都失败。
+
+前两个 gate 覆盖两条回归路径:
+
+| 基准 | 负载 | Gate |
+|---|---|---|
+| `packages/session/session-persistence-jsonl/tests/open-generation.bench.ts` | 200 轮 ×(500 text + 125 reasoning delta)= 127,400 个 released-v0 事件,约 2.8 MB,经冻结 v0 codec 以 packed row 编码 | 迁移的首次 `open()` 在 128 MB 堆下 ≤ 2,000 ms;新进程打开已发布 current generation ≤ 500 ms;三次尝试取最小值 |
+| `packages/client/ui-chat/tests/conversation-fold.bench.client.ts` | 200 个回复,每个紧凑 stream 含 2,000 text + 500 reasoning delta(1,600 条记录中 500,000 个 delta),由真实 `ConversationNodeAssembler` 经全部 Chat Definition fold | 大窗口 fold ≤ 150 ms;大窗口 fold ≤ 每回复 100 delta 的同一窗口的 3 倍 |
+
+在引入该 gate 的提交上于参考机器测得:迁移基准耗尽 128 MB 堆,fold 基准在小窗口与大窗口之间缩放 11 倍,因此两个 gate 都在回归代码上失败,并在两条路径改为 O(records) 工作后通过。
+
+## 考虑过的替代方案
+
+**扩展手动 `test:web:perf` 清单。** 拒绝:它有意留在 CI 之外,测量的是简化 fold 而非已注册的 Definition,且不做任何断言。
+
+**只用时间预算。** 拒绝:单一绝对预算要么在较慢的 runner 上失败,要么在较快的 runner 上放过回归;堆上限与缩放比给出与主机无关的判定,壁钟预算则作为用户可见症状所对应的超时保留。
+
+**用真实录制语料做基准。** 拒绝:语料 fixture 按策略保持小体量,录制材料不得成为基准输入,且其重新录制会静默移动基线。
+
+**把基准放进现有 gate 聚合中运行。** 拒绝:聚合在一个 runner 上并发运行各 gate,壁钟测量会继承邻居的 CPU 负载。
+
+## 后果
+
+每个 pull request 多付出一个几分钟的必需 Linux job,其时间主要花在安装而不是基准本身。让首次打开或 Client fold 慢于预算、重于堆限制或与流式 delta 数成正比的改动,会在引入它的 PR 中失败,并把测得的数字打印在 job 日志里。预算变更是对常量及其理由注释的受评审编辑,绝不是环境变量覆盖;新增基准必须说明它防护哪条 owner 可见路径和哪类回归。该 gate 不测量浏览器渲染、网络传输或真实录制的 Session;这些仍由手动 `test:web:perf` 清单和评审覆盖。

+ 47 - 1
.github/workflows/ci.yml

@@ -156,6 +156,52 @@ jobs:
       - name: Run exhaustive coverage
         run: pnpm run check:ci:coverage
 
+  node-24-bench:
+    if: github.event_name == 'pull_request'
+    runs-on: >-
+      ${{ vars.DSH_CI_FAILOVER_LINUX == 'selfhosted'
+          && github.event.pull_request.user.login != 'dependabot[bot]'
+          && fromJSON('["self-hosted", "linux", "x64", "vm-backup"]')
+          || 'dsh-ubuntu-24-04-16core' }}
+    name: node 24 / benchmarks
+    # Wall-clock budgets need an otherwise idle runner, so this job runs the
+    # benchmark lane alone instead of joining a concurrent gate aggregate.
+    timeout-minutes: 30
+    steps:
+      - uses: actions/checkout@v6
+        with:
+          persist-credentials: false
+
+      - uses: pnpm/action-setup@v4
+        with:
+          dest: ${{ runner.temp }}/setup-pnpm-${{ github.run_id }}-${{ github.run_attempt }}
+
+      - uses: actions/setup-node@v6
+        with:
+          node-version: ${{ env.PRIMARY_NODE_VERSION }}
+
+      - name: Configure pnpm store path
+        id: pnpm-store
+        run: |
+          store_root="$HOME/.local/share/pnpm/store"
+          echo "PNPM_CONFIG_STORE_DIR=$store_root" >> "$GITHUB_ENV"
+          store_path=$(PNPM_CONFIG_STORE_DIR="$store_root" pnpm store path --silent)
+          echo "path=$store_path" >> "$GITHUB_OUTPUT"
+
+      - uses: actions/cache/restore@v4
+        if: vars.DSH_CI_FAILOVER_LINUX != 'selfhosted' || github.event.pull_request.user.login == 'dependabot[bot]'
+        with:
+          path: ${{ steps.pnpm-store.outputs.path }}
+          key: ${{ runner.os }}-node-${{ env.PRIMARY_NODE_VERSION }}-pnpm-${{ hashFiles('pnpm-lock.yaml') }}
+          restore-keys: |
+            ${{ runner.os }}-node-${{ env.PRIMARY_NODE_VERSION }}-pnpm-
+
+      - name: Install (immutable)
+        run: pnpm install --frozen-lockfile
+
+      - name: Run performance benchmarks
+        run: pnpm run check:ci:bench
+
   node-24-consumers:
     if: github.event_name == 'pull_request'
     runs-on: >-
@@ -665,7 +711,7 @@ jobs:
           && github.event.pull_request.user.login != 'dependabot[bot]'
           && fromJSON('["self-hosted", "linux", "x64", "vm-backup"]')
           || 'ubuntu-latest' }}
-    needs: [node-24, node-24-coverage, node-24-consumers, node-compat, python-sdk, python-runtime, windows, windows-build, windows-native-tests]
+    needs: [node-24, node-24-coverage, node-24-bench, node-24-consumers, node-compat, python-sdk, python-runtime, windows, windows-build, windows-native-tests]
     if: always() && github.event_name == 'pull_request'
     steps:
       - name: Fail if any needed job did not succeed

+ 2 - 2
docs/testing.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/testing.md
-testing.md: c1c75c911b27ddaf9e8c340eef3a5cf394a8d82c
-testing.zh.md: ff84d439a53a0b7bc7a6607857a0c0303c2aadff
+testing.md: 3b9e22278821bff6c47ff4291223c6738e0e8fe7
+testing.zh.md: 2674eef23b9eab07f0a09aae688e68dd7ea9c32a

+ 3 - 2
docs/testing.md

@@ -10,14 +10,15 @@ How this repo tests, tier by tier, and the rules that keep a green suite meaning
 - **Coverage gate** (`pnpm run test:coverage`): the gating run, per-file 100% on `packages/*/*/src`. An uncovered line is often dead code the gate flags for deletion, not a missing test to bolt on. Line coverage is necessary, never sufficient — it proves lines ran, not that the feature works as shipped. Per-file 100% on `packages/shell/pwsh-local/src` needs a real `pwsh`: without one its executor suites self-skip and `vitest.config.ts` exempts the file so pwsh-less hosts stay green, while CI runners ship pwsh and enforce the full bar.
 - **Real-API e2e** (`pnpm run test:e2e`): with-key tests against live provider APIs — the DeepSeek model plus provider-specific smokes that gate on their own keys (`EXA_API_KEY`, `PERPLEXITY_API_KEY`, …); each suite self-skips without its key so keyless CI stays green ([real-API e2e Agent Note](../.agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.md)).
 - **Owner-local expected output** (`pnpm run test:expected`): keyless assembled CLI/process expectations without a recorded-session round trip. Drivers use `*.expected.e2e.ts` beside `tests/expected/`; CI runs built exports. Package/script expectations use `test`, while browser expectations use `test:web`.
+- **Performance benchmarks** (`pnpm run test:bench`; required Linux PR gate `node 24 / benchmarks`): `*.bench.ts` and Client-face `*.bench.client.ts` files under `packages/*/*/tests/` synthesize input from fixed parameters, never recorded material, and fail on a documented wall-clock budget, heap limit, or scaling ratio ([rules and current gates](../.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md)).
 - **Snapshot** (`pnpm run test:snapshot`): a top-level scenario's highest recorded parent generation supplies user input and model replay, then serves as the expected persisted result. Parent filenames are `session[.vN].jsonl`; child roles are `session.<ordinal>[.vN].jsonl`; v0 omits `.v0`, positive versions require lowercase `.vN`, and each filename must agree with its header. Process scenarios start through `dsh`: headless owns one-shot behavior, the SDK owns persistent control, ACP owns automation-protocol behavior, and Web retains browser/ARIA evidence beside the same Session. `snapshot.yml` declares the profile, composition/header class, recording policy, exceptional replay or input metadata, and workspace facts. Typed tokens preserve parent/child identity relationships; only header pins own prompt/schema sidecars. A mutating scenario independently compares the complete `workspace.expected/` tree, which record and refresh never rewrite. Use `test:snapshot:record` when a model transcript changes and `test:snapshot:refresh` when replay input remains valid; review every resulting diff.
 - **Web browser snapshot** (`pnpm run test:web`; required Linux PR gate): Chromium compares session-driven output under `snapshots/web/` and UI-only output under `apps/web/tests/expected/`. CI forces read-only `DSH_SNAPSHOT=replay`, never writing expected outputs; record/refresh stay local and every diff is reviewed ([web e2e lane](../.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.md), [CI gate decision](../.agents/notes/implemented/testing/2026-07-30-web-browser-snapshot-ci-gate.md)). `test:web` builds first for plugin CSS.
 
-Session fixtures retain headers and payloads but omit body sequence/time envelopes; replay synthesizes them. Replay, record, and refresh select each parent/child role's highest generation. Current v2 uses `.v2`, one row per event, and embedded compact Assistant streams; retained v0 (suffixless) and v1 (`.v1`) may keep canonical packed rows for migration coverage. [The migrator](../scripts/migrate-packed-session-fixtures.ts) rewrites older historical layouts.
+Session fixtures retain headers and payloads but omit body sequence/time envelopes; replay synthesizes them and, like record and refresh, selects each role's highest generation. Retained v0 and v1 generations may keep packed rows for migration coverage; [the migrator](../scripts/migrate-packed-session-fixtures.ts) rewrites older layouts.
 
 ## How specs execute
 
-Forked workers run several spec files at once, the coverage gate splits into concurrent partitions beside the other gates in its job, and the self-hosted runners share one host and one volume. Only the process is isolated: ports, predictable paths, external namespaces, and inherited children are not. Own each acquired resource through its teardown, and read a spec that passes only when it runs alone as a defect in the spec rather than an unstable runner. [dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) owns the allocation, restoration, synchronization, timeout-budget, platform, and teardown rules; its [flake diagnosis workflow](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md) classifies an existing probabilistic failure.
+Forked workers run several spec files at once, the coverage gate splits into concurrent partitions beside the other gates in its job, and the self-hosted runners share one host and one volume. Only the process is isolated: ports, predictable paths, external namespaces, and inherited children are not. Own each acquired resource through its teardown; a spec that passes only when run alone is defective, not the runner. [dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) owns the allocation, restoration, synchronization, timeout-budget, platform, and teardown rules; its [flake diagnosis workflow](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md) classifies an existing probabilistic failure.
 
 ## The with-key policy: inference is cheap here
 

+ 3 - 2
docs/testing.zh.md

@@ -10,14 +10,15 @@
 - **覆盖率门禁**(`pnpm run test:coverage`):门禁级运行,对 `packages/*/*/src` 按文件 100% 覆盖。未覆盖的行往往是门禁正确标记出的死代码(应删除),而非需要补写的测试。行覆盖率是必要条件,但永远不是充分条件:它证明行被执行过,不证明功能按交付预期工作。`packages/shell/pwsh-local/src` 的按文件 100% 覆盖需要真实的 `pwsh`:缺少它时其执行器套件会自动跳过,`vitest.config.ts` 会豁免该文件以使无 pwsh 的主机保持绿色,而 CI runner 自带 pwsh,仍按完整标准执行门禁。
 - **真实 API e2e**(`pnpm run test:e2e`):带密钥测试调用真实提供方 API,包括 DeepSeek 模型以及各提供方特有的冒烟测试;这些测试各自由自己的密钥控制(`EXA_API_KEY`、`PERPLEXITY_API_KEY` 等),缺少密钥时套件会自动跳过,使 keyless CI 保持绿色([真实 API e2e Agent Note](../.agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.zh.md))。
 - **所属位置的预期输出**(`pnpm run test:expected`):无录制会话往返的无密钥组装 CLI/进程预期。驱动使用 `*.expected.e2e.ts`,并与 `tests/expected/` 同属一处;CI 针对构建产物运行。包/脚本预期使用 `test`,浏览器预期使用 `test:web`。
+- **性能基准**(`pnpm run test:bench`;必需的 Linux PR gate `node 24 / benchmarks`):`packages/*/*/tests/` 下的 `*.bench.ts` 与 Client 面 `*.bench.client.ts` 文件按固定参数合成输入,绝不含录制材料,并在超出已记录的壁钟预算、堆限制或缩放比时失败([规则与当前 gate](../.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md))。
 - **快照**(`pnpm run test:snapshot`):顶层场景数值最高的已录制 parent generation 同时提供用户输入和模型回放,并作为持久化结果的预期值。parent 文件名是 `session[.vN].jsonl`;child 角色使用 `session.<ordinal>[.vN].jsonl`;v0 省略 `.v0`,正版本必须使用小写 `.vN`,且每个文件名必须与其 header 一致。进程级场景都通过 `dsh` 启动:headless 负责一次性行为,SDK 负责持久控制,ACP 负责自动化协议行为,Web 在同一 Session 旁保留浏览器与 ARIA 证据。`snapshot.yml` 声明 profile、组合与请求头类别、录制策略、例外回放或输入元数据以及 workspace 事实。带类型的 token 保留父子身份关系;只有请求头 pin 拥有 prompt/schema sidecar。变更 workspace 的场景会独立比较完整的 `workspace.expected/` 目录,record 与 refresh 绝不改写该目录。当模型 transcript(文本记录)变化时使用 `test:snapshot:record`,回放输入仍有效时使用 `test:snapshot:refresh`;请审查所有结果差异。
 - **Web 浏览器快照**(`pnpm run test:web`;必需的 Linux PR(Pull Request)门禁):Chromium 比较 `snapshots/web/` 下由会话驱动的输出,以及 `apps/web/tests/expected/` 下仅含 UI 的输出。CI 强制只读的 `DSH_SNAPSHOT=replay`,绝不写入预期输出;record/refresh 留在本地,每处 diff 都须评审([web e2e 车道](../.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.zh.md)、[CI 门禁决策](../.agents/notes/implemented/testing/2026-07-30-web-browser-snapshot-ci-gate.zh.md))。`test:web` 会先构建以交付插件 CSS。
 
-Session fixture 保留 header 与 payload,但省略正文 seq/time envelope;replay 会合成这些 envelope。Replay、record 与 refresh 会选择每个 parent/child 角色的最高 generation。当前 v2 使用 `.v2`、每个事件一行,并嵌入紧凑 Assistant stream;保留的 v0(无后缀)与 v1(`.v1`)可以为迁移覆盖保留规范 packed row。[迁移器](../scripts/migrate-packed-session-fixtures.ts)会改写更旧的历史布局。
+Session fixture 保留 header 与 payload,但省略正文 seq/time envelope;replay 会合成这些 envelope,并与 record、refresh 一样选择每个角色的最高 generation。保留的 v0 与 v1 generation 可以为迁移覆盖保留 packed row;[迁移器](../scripts/migrate-packed-session-fixtures.ts)会改写更旧的布局。
 
 ## spec 如何被执行
 
-fork 出的 worker 会同时运行多个 spec 文件,coverage gate 会拆成并发的 partition,与同一个 job 中的其它 gate 并排运行,而自托管 runner 共用同一台宿主机和同一个卷。被隔离的只有进程:端口、可预测路径、外部命名空间和继承而来的子进程都不隔离。为每个占用的资源负责到它的 teardown,并把「只有单独运行时才通过」的 spec 读作该 spec 的缺陷,而不是 runner 不稳定。[dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) 负责资源分配、状态恢复、同步、超时预算、平台差异与 teardown 规则;它的 [flake 诊断流程](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md)用于归类已经存在的概率性失败。
+fork 出的 worker 会同时运行多个 spec 文件,coverage gate 会拆成并发的 partition,与同一个 job 中的其它 gate 并排运行,而自托管 runner 共用同一台宿主机和同一个卷。被隔离的只有进程:端口、可预测路径、外部命名空间和继承而来的子进程都不隔离。为每个占用的资源负责到它的 teardown;只有单独运行时才通过的 spec 是缺陷,而不是 runner 不稳定。[dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) 负责资源分配、状态恢复、同步、超时预算、平台差异与 teardown 规则;它的 [flake 诊断流程](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md)用于归类已经存在的概率性失败。
 
 ## 带密钥策略:推理(inference)在这里很便宜
 

+ 2 - 0
package.json

@@ -36,6 +36,7 @@
     "test:coverage": "vitest run --coverage",
     "test:coverage:partitioned": "tsx scripts/run-coverage-partitions.ts",
     "test:e2e": "vitest run --config vitest.e2e.config.ts",
+    "test:bench": "vitest run --config vitest.bench.config.ts",
     "test:expected": "vitest run --config vitest.expected.config.ts",
     "test:expected:refresh": "DSH_SNAPSHOT=refresh vitest run --config vitest.expected.config.ts",
     "test:issue-management": "node .github/issue-management/policy.test.mjs",
@@ -59,6 +60,7 @@
     "check:ci:static": "tsx scripts/run-gates.ts ci-static",
     "check:ci:lint:contracts-ready": "tsx scripts/run-gates.ts ci-lint-contracts-ready",
     "check:ci:coverage": "tsx scripts/run-gates.ts ci-coverage",
+    "check:ci:bench": "tsx scripts/run-gates.ts ci-bench",
     "check:ci:snapshot": "tsx scripts/run-gates.ts ci-snapshot",
     "check:ci:artifacts": "tsx scripts/run-gates.ts ci-artifacts",
     "check:ci:consumers": "tsx scripts/run-gates.ts ci-consumers",

+ 217 - 0
packages/client/ui-chat/tests/conversation-fold.bench.client.ts

@@ -0,0 +1,217 @@
+/**
+ * Performance gate for the cold Client fold of a large Session format v2
+ * history window: every registered Chat Definition runs over a synthesized
+ * window in which each assistant reply embeds its compact stream. The gate
+ * bounds the wall time and requires the fold to scale with the number of
+ * compact stream records rather than with the number of streamed deltas.
+ */
+
+import { describe, expect, it } from 'vitest'
+import { AssistantStreamAccumulator } from '@deepseek-ai/dsh-llm/assistant-stream'
+import type { StreamChunk } from '@deepseek-ai/dsh-llm'
+import type { SessionEvent } from '@deepseek-ai/dsh-session/types'
+import type { ChatSnapshot } from '@deepseek-ai/dsh-client-ui-chat/client'
+import type { SessionEventLikeEntry } from '@deepseek-ai/dsh-api-session-controller/client'
+import {
+  ConversationNodeAssembler,
+  inspectRequestPrompt,
+  type ConversationNodeDefinition,
+  type ConversationViewDefinition,
+} from '@deepseek-ai/dsh-client-ui-conversation/client'
+import { assistantDefinition } from '../src/client/conversation-nodes/assistant.ts'
+import { chatViewDefinition } from '../src/client/conversation-nodes/chat-snapshot-builder.ts'
+import { commandDefinition } from '../src/client/conversation-nodes/command.ts'
+import { compactionDefinition } from '../src/client/conversation-nodes/compaction.ts'
+import { unknownFallbackDefinition } from '../src/client/conversation-nodes/fallback.ts'
+import { nextStepInboxDefinition } from '../src/client/conversation-nodes/inbox.ts'
+import { messageDefinition } from '../src/client/conversation-nodes/message.ts'
+import { requestPromptDefinition } from '../src/client/conversation-nodes/request-prompt.ts'
+import { retryDefinition } from '../src/client/conversation-nodes/retry.ts'
+import { toolDefinition } from '../src/client/conversation-nodes/tool.ts'
+import { turnErrorDefinition } from '../src/client/conversation-nodes/turn-error.ts'
+import { turnMaxTokensDefinition } from '../src/client/conversation-nodes/turn-max-tokens.ts'
+import { turnProcessDefinition } from '../src/client/conversation-nodes/turn-process.ts'
+import { turnTailDefinition } from '../src/client/conversation-nodes/turn-tail.ts'
+
+/** Replies in the folded window; each carries one reasoning block and one text block. */
+const TURNS = 200
+
+/** Text deltas per reply in the large workload; the reply also streams `deltas / 4` reasoning deltas. */
+const LARGE_DELTAS = 2_000
+
+/** Text deltas per reply in the small workload used as the scaling reference. */
+const SMALL_DELTAS = 100
+
+/**
+ * Wall-clock budget for folding the large window (200 replies, 500,000 streamed
+ * deltas compacted into 800 stream records). The pre-stack fold processed the
+ * equivalent packed chunk rows in a few milliseconds; the budget leaves room
+ * for the complete Definition set and slower CI hosts while staying below the
+ * per-delta replay that needed hundreds of milliseconds for this window.
+ */
+const LARGE_FOLD_BUDGET_MS = 150
+
+/**
+ * Maximum ratio between folding the large and the small window. Both windows
+ * hold the same number of events and compact records, so a fold that scales
+ * with records stays near 1; a fold that replays every delta grows with the
+ * 20× delta count.
+ */
+const MAX_DELTA_SCALING = 3
+
+/** Attempts per workload; the gate compares minima so scheduler noise only adds. */
+const ATTEMPTS = 3
+
+const TIME_ZERO = 1_700_000_000_000
+
+class BenchEventDefinitions {
+  readonly definitions: readonly ConversationNodeDefinition[] = [
+    nextStepInboxDefinition,
+    messageDefinition,
+    requestPromptDefinition(inspectRequestPrompt),
+    assistantDefinition,
+    turnProcessDefinition,
+    toolDefinition,
+    commandDefinition,
+    compactionDefinition,
+    retryDefinition,
+    turnErrorDefinition,
+    turnMaxTokensDefinition,
+    turnTailDefinition,
+  ]
+
+  entries(): readonly ConversationNodeDefinition[] {
+    return this.definitions
+  }
+
+  fallbackEntry(): ConversationNodeDefinition {
+    return unknownFallbackDefinition
+  }
+}
+
+class BenchViewDefinitions {
+  entries(): readonly ConversationViewDefinition[] {
+    return [chatViewDefinition]
+  }
+}
+
+function entry(seq: number, type: string, data: unknown, extra: Record<string, unknown> = {}): SessionEventLikeEntry {
+  return {
+    type: 'event',
+    event: { seq, time: TIME_ZERO + seq, type, data, ...extra } as unknown as SessionEvent,
+  }
+}
+
+/**
+ * Synthesize one v2 history window: `turns` completed replies whose compact
+ * streams are accumulated from `deltas` text deltas and `deltas / 4`
+ * reasoning deltas each.
+ */
+function synthesizeWindow(turns: number, deltas: number): { readonly entries: readonly SessionEventLikeEntry[]; readonly records: number } {
+  const entries: SessionEventLikeEntry[] = []
+  let seq = 0
+  let records = 0
+  const push = (type: string, data: unknown, extra: Record<string, unknown> = {}): void => {
+    entries.push(entry(seq, type, data, extra))
+    seq += 1
+  }
+  const reasoningDeltas = Math.floor(deltas / 4)
+  for (let turn = 1; turn <= turns; turn += 1) {
+    push('turn/start', { turn })
+    push('user/message', {
+      id: `user-${String(turn)}`,
+      role: 'user',
+      content: [{ type: 'text', text: `prompt ${String(turn)}` }],
+      source: { kind: 'user' },
+    }, { surfaceOp: 'append' })
+    push('step/start', { turn, step: 1 })
+    const accumulator = new AssistantStreamAccumulator()
+    let time = TIME_ZERO + seq * 1_000
+    const stream = (chunk: StreamChunk): void => {
+      accumulator.push({ time, chunk })
+      time += 1
+    }
+    stream({ type: 'block-start', index: 0, blockType: 'reasoning' })
+    let reasoning = ''
+    for (let index = 0; index < reasoningDeltas; index += 1) {
+      const delta = `r${String(index)} `
+      reasoning += delta
+      stream({ type: 'reasoning-delta', index: 0, text: delta })
+    }
+    stream({ type: 'block-end', index: 0, block: { type: 'reasoning', text: reasoning } })
+    stream({ type: 'block-start', index: 1, blockType: 'text' })
+    let text = ''
+    for (let index = 0; index < deltas; index += 1) {
+      const delta = `w${String(index)} `
+      text += delta
+      stream({ type: 'text-delta', index: 1, text: delta })
+    }
+    stream({ type: 'block-end', index: 1, block: { type: 'text', text } })
+    const usage = { inputTokens: 100, outputTokens: deltas }
+    stream({ type: 'usage', usage })
+    stream({ type: 'finish', reason: { kind: 'stop' } })
+    const snapshot = accumulator.snapshot()
+    records += snapshot.length
+    push('assistant/message', {
+      turn,
+      step: 1,
+      message: {
+        id: `assistant-${String(turn)}`,
+        role: 'assistant',
+        content: [{ type: 'reasoning', text: reasoning }, { type: 'text', text }],
+        source: { kind: 'model', provider: 'bench', model: 'bench' },
+      },
+      usage,
+      stream: snapshot,
+    }, { surfaceOp: 'append' })
+    push('step/end', { turn, step: 1 })
+    push('turn/end', { turn, reason: { kind: 'completed' } })
+  }
+  return { entries, records }
+}
+
+function foldOnce(entries: readonly SessionEventLikeEntry[]): { readonly ms: number; readonly nodes: number } {
+  const started = performance.now()
+  const assembler = new ConversationNodeAssembler(new BenchEventDefinitions(), new BenchViewDefinitions())
+  assembler.replaceWindow(entries, false)
+  assembler.activateTarget('chat')
+  const snapshot = assembler.snapshot('chat') as ChatSnapshot | undefined
+  return { ms: performance.now() - started, nodes: snapshot?.order.length ?? 0 }
+}
+
+function bestOf(entries: readonly SessionEventLikeEntry[]): { readonly ms: number; readonly nodes: number } {
+  let best = foldOnce(entries)
+  for (let attempt = 1; attempt < ATTEMPTS; attempt += 1) {
+    const next = foldOnce(entries)
+    if (next.ms < best.ms) best = next
+  }
+  return best
+}
+
+describe('cold Chat fold of a large v2 history window', () => {
+  it(`folds ${String(TURNS)} replies with ${String(LARGE_DELTAS)} deltas each within ${String(LARGE_FOLD_BUDGET_MS)} ms and scales with compact records`, () => {
+    const small = synthesizeWindow(TURNS, SMALL_DELTAS)
+    const large = synthesizeWindow(TURNS, LARGE_DELTAS)
+    expect(large.entries.length).toBe(small.entries.length)
+    expect(large.records).toBe(small.records)
+
+    const smallFold = bestOf(small.entries)
+    const largeFold = bestOf(large.entries)
+    const scaling = largeFold.ms / Math.max(smallFold.ms, 1)
+    console.log(JSON.stringify({
+      benchmark: 'conversation-fold/large-window',
+      events: large.entries.length,
+      compactRecords: large.records,
+      streamedDeltas: TURNS * (LARGE_DELTAS + Math.floor(LARGE_DELTAS / 4)),
+      chatNodes: largeFold.nodes,
+      smallFoldMs: Math.round(smallFold.ms * 10) / 10,
+      largeFoldMs: Math.round(largeFold.ms * 10) / 10,
+      scaling: Math.round(scaling * 100) / 100,
+      budgetMs: LARGE_FOLD_BUDGET_MS,
+      maxScaling: MAX_DELTA_SCALING,
+    }))
+    expect(largeFold.nodes).toBeGreaterThan(0)
+    expect(largeFold.ms).toBeLessThanOrEqual(LARGE_FOLD_BUDGET_MS)
+    expect(scaling).toBeLessThanOrEqual(MAX_DELTA_SCALING)
+  })
+})

+ 153 - 0
packages/session/session-persistence-jsonl/tests/open-generation.bench.ts

@@ -0,0 +1,153 @@
+/**
+ * Performance gate for opening a large released-v0 Session log through the
+ * JSONL backend: the first `open()` migrates and publishes the current
+ * generation; later opens decode the published generation. Both run in child
+ * processes under a fixed heap limit so an allocation regression fails as an
+ * out-of-memory exit instead of passing on a machine with more memory.
+ */
+
+import { spawn } from 'node:child_process'
+import { copyFile, mkdir, mkdtemp, readdir, rm } from 'node:fs/promises'
+import { tmpdir } from 'node:os'
+import { join } from 'node:path'
+import { afterAll, beforeAll, describe, expect, it } from 'vitest'
+import type { OpenGenerationWorkerReport } from './open-generation.bench.worker.ts'
+import { SYNTHETIC_SESSION_DIRECTORY, writeSyntheticReleasedV0Log } from './synthetic-released-v0-log.ts'
+
+/** 200 turns × (500 text + 125 reasoning deltas): 127,400 released-v0 events in about 2.8 MB of JSONL. */
+const SHAPE = { turns: 200, textDeltas: 500 } as const
+
+/**
+ * Wall-clock budget for the migrating first `open()`. The pre-stack backend
+ * decoded the same bytes in about 35 ms on the reference machine; a whole
+ * artifact migration that validates, transforms, publishes, and re-reads the
+ * log costs a small multiple of that, and the budget leaves a further
+ * multiple for slower CI hosts while staying far below the ~5 s that the
+ * repeated-snapshot implementation needed.
+ */
+const MIGRATION_BUDGET_MS = 2_000
+
+/**
+ * Old-space limit for the migrating child process. Pre-stack decoding of the
+ * same log completed under 128 MB; the repeated-snapshot migration exhausted
+ * that heap. Holding the limit fixed keeps the gate independent of the
+ * runner's physical memory.
+ */
+const MIGRATION_HEAP_LIMIT_MB = 128
+
+/** Wall-clock budget for a fresh process opening the already published current generation. */
+const STEADY_OPEN_BUDGET_MS = 500
+
+/** Attempts per measurement; the gate compares the minimum so scheduler noise only adds. */
+const ATTEMPTS = 3
+
+const WORKER = join(import.meta.dirname, 'open-generation.bench.worker.ts')
+
+interface WorkerRun {
+  readonly report: OpenGenerationWorkerReport | undefined
+  readonly exitCode: number | null
+  readonly stderr: string
+}
+
+function runWorker(root: string, mode: 'migrate' | 'steady', heapLimitMb: number): Promise<WorkerRun> {
+  return new Promise((resolve, reject) => {
+    const child = spawn(process.execPath, [
+      `--max-old-space-size=${String(heapLimitMb)}`,
+      '--import',
+      'tsx/esm',
+      WORKER,
+      root,
+      mode,
+    ], { cwd: process.cwd(), stdio: ['ignore', 'pipe', 'pipe'] })
+    let stdout = ''
+    let stderr = ''
+    child.stdout.setEncoding('utf8').on('data', (chunk: string) => { stdout += chunk })
+    child.stderr.setEncoding('utf8').on('data', (chunk: string) => { stderr += chunk })
+    child.once('error', reject)
+    child.once('close', (exitCode) => {
+      const line = stdout.trim().split('\n').at(-1)
+      let report: OpenGenerationWorkerReport | undefined
+      if (exitCode === 0 && line !== undefined && line.startsWith('{')) {
+        report = JSON.parse(line) as OpenGenerationWorkerReport
+      }
+      resolve({ report, exitCode, stderr })
+    })
+  })
+}
+
+function requireReport(run: WorkerRun, label: string): OpenGenerationWorkerReport {
+  if (run.report === undefined) {
+    const lines = run.stderr.trim().split('\n')
+    const fatal = lines.filter(line => /FATAL ERROR|heap limit|out of memory/i.test(line))
+    const detail = (fatal.length > 0 ? fatal : lines.slice(-8)).join('\n')
+    throw new Error(`${label} exited with ${String(run.exitCode)} under --max-old-space-size=${String(MIGRATION_HEAP_LIMIT_MB)}:\n${detail}`)
+  }
+  return run.report
+}
+
+describe('opening a large released-v0 Session log', () => {
+  let scratch: string
+  let sourcePath: string
+  let sourceBytes = 0
+  let sourceEvents = 0
+  const migratedRoots: string[] = []
+
+  beforeAll(async () => {
+    scratch = await mkdtemp(join(tmpdir(), 'dsh-open-generation-bench-'))
+    const written = await writeSyntheticReleasedV0Log(join(scratch, 'source'), SHAPE)
+    sourcePath = written.path
+    sourceBytes = written.bytes
+    sourceEvents = written.events
+  })
+
+  afterAll(async () => {
+    await rm(scratch, { recursive: true, force: true })
+  })
+
+  it(`migrates ${String(SHAPE.turns)} turns of streamed replies within ${String(MIGRATION_BUDGET_MS)} ms under a ${String(MIGRATION_HEAP_LIMIT_MB)} MB heap`, async () => {
+    const reports: OpenGenerationWorkerReport[] = []
+    for (let attempt = 0; attempt < ATTEMPTS; attempt += 1) {
+      const root = join(scratch, `migrate-${String(attempt)}`)
+      await mkdir(join(root, SYNTHETIC_SESSION_DIRECTORY), { recursive: true })
+      await copyFile(sourcePath, join(root, SYNTHETIC_SESSION_DIRECTORY, 'session.jsonl'))
+      reports.push(requireReport(await runWorker(root, 'migrate', MIGRATION_HEAP_LIMIT_MB), `migration attempt ${String(attempt)}`))
+      migratedRoots.push(root)
+      const files = (await readdir(join(root, SYNTHETIC_SESSION_DIRECTORY))).sort()
+      expect(files).toEqual(['session.jsonl', 'session.v2.jsonl'])
+    }
+    const openMs = Math.min(...reports.map(report => report.openMs))
+    const parseMs = Math.min(...reports.map(report => report.parseMs))
+    console.log(JSON.stringify({
+      benchmark: 'open-generation/migrate',
+      sourceBytes,
+      sourceEvents,
+      currentEvents: reports[0]?.events,
+      openMs: Math.round(openMs),
+      parseMs: Math.round(parseMs),
+      readMs: Math.round(Math.min(...reports.map(report => report.readMs))),
+      heapUsedMb: Math.round(Math.max(...reports.map(report => report.heapUsedMb))),
+      heapLimitMb: MIGRATION_HEAP_LIMIT_MB,
+      budgetMs: MIGRATION_BUDGET_MS,
+    }))
+    expect(reports.every(report => report.headerVersion === 2)).toBe(true)
+    expect(openMs).toBeLessThanOrEqual(MIGRATION_BUDGET_MS)
+  })
+
+  it(`opens the published current generation within ${String(STEADY_OPEN_BUDGET_MS)} ms`, async () => {
+    expect(migratedRoots.length, 'a published current generation from the migration benchmark').toBeGreaterThan(0)
+    const reports: OpenGenerationWorkerReport[] = []
+    for (let attempt = 0; attempt < ATTEMPTS; attempt += 1) {
+      const root = migratedRoots[attempt % migratedRoots.length] as string
+      reports.push(requireReport(await runWorker(root, 'steady', MIGRATION_HEAP_LIMIT_MB), `steady attempt ${String(attempt)}`))
+    }
+    const openMs = Math.min(...reports.map(report => report.openMs))
+    console.log(JSON.stringify({
+      benchmark: 'open-generation/steady',
+      openMs: Math.round(openMs),
+      readMs: Math.round(Math.min(...reports.map(report => report.readMs))),
+      heapUsedMb: Math.round(Math.max(...reports.map(report => report.heapUsedMb))),
+      budgetMs: STEADY_OPEN_BUDGET_MS,
+    }))
+    expect(openMs).toBeLessThanOrEqual(STEADY_OPEN_BUDGET_MS)
+  })
+})

+ 63 - 0
packages/session/session-persistence-jsonl/tests/open-generation.bench.worker.ts

@@ -0,0 +1,63 @@
+/**
+ * Child-process worker for the open-generation benchmark: opens one Session
+ * through the JSONL backend under the caller's heap limit and reports timings
+ * as one JSON line. Arguments: `<root> <mode>` where mode is `migrate`
+ * (release-v0 source only) or `steady` (published current generation).
+ */
+
+import { Context } from '@deepseek-ai/cordis'
+import { readFile } from 'node:fs/promises'
+import { join } from 'node:path'
+import { SESSION_FORMAT_VERSION, SessionId } from '@deepseek-ai/dsh-session'
+import JsonlSessionPersistence from '@deepseek-ai/dsh-session-persistence-jsonl'
+import { SYNTHETIC_SESSION_DIRECTORY, SYNTHETIC_SESSION_ID } from './synthetic-released-v0-log.ts'
+
+/** Timings printed by the worker. */
+export interface OpenGenerationWorkerReport {
+  readonly mode: 'migrate' | 'steady'
+  /** `open()` wall time; for `migrate` this includes publishing the current generation. */
+  readonly openMs: number
+  /** `read()` of the complete current event list after `open()`. */
+  readonly readMs: number
+  /** `JSON.parse` of every source line, as the pure parsing floor of the same bytes. */
+  readonly parseMs: number
+  readonly events: number
+  readonly headerVersion: number
+  readonly heapUsedMb: number
+}
+
+const [root, mode] = process.argv.slice(2)
+if (root === undefined || (mode !== 'migrate' && mode !== 'steady')) {
+  throw new Error('usage: open-generation.bench.worker.ts <root> migrate|steady')
+}
+
+const sourceText = await readFile(join(root, SYNTHETIC_SESSION_DIRECTORY, 'session.jsonl'), 'utf8')
+const parseStarted = performance.now()
+for (const line of sourceText.split('\n')) {
+  if (line.length > 0) JSON.parse(line)
+}
+const parseMs = performance.now() - parseStarted
+
+const ctx = new Context()
+await ctx.plugin(JsonlSessionPersistence, { root, compression: 'none' })
+const openStarted = performance.now()
+const handle = await ctx.sessionPersistence.open(SessionId(SYNTHETIC_SESSION_ID), 'read')
+const openMs = performance.now() - openStarted
+const readStarted = performance.now()
+const events = await handle.read()
+const readMs = performance.now() - readStarted
+await handle.close()
+if (handle.header.version !== SESSION_FORMAT_VERSION) {
+  throw new Error(`expected current format v${SESSION_FORMAT_VERSION}, opened v${handle.header.version}`)
+}
+const report: OpenGenerationWorkerReport = {
+  mode,
+  openMs,
+  readMs,
+  parseMs,
+  events: events.length,
+  headerVersion: handle.header.version,
+  heapUsedMb: process.memoryUsage().heapUsed / 1_048_576,
+}
+process.stdout.write(`${JSON.stringify(report)}\n`)
+process.exit(0)

+ 142 - 0
packages/session/session-persistence-jsonl/tests/synthetic-released-v0-log.ts

@@ -0,0 +1,142 @@
+/**
+ * Deterministic released-v0 Session log synthesized from fixed parameters.
+ * The content is generated in-process (numbered prompts, counters, and
+ * repeated tokens) so the benchmark input carries no recorded material.
+ */
+
+import { mkdir, writeFile } from 'node:fs/promises'
+import { join } from 'node:path'
+import { releasedV0SessionFormatCodec } from '@deepseek-ai/dsh-session-format-v0-to-v1'
+import type { SessionFormatEvent } from '@deepseek-ai/dsh-session-format'
+
+/** Fixed workload parameters; every count below is derived from them. */
+export interface SyntheticV0LogShape {
+  /** Completed turns, each with one user prompt and one streamed assistant reply. */
+  readonly turns: number
+  /** `text-delta` chunks per reply; the reply also streams `textDeltas / 4` reasoning deltas. */
+  readonly textDeltas: number
+}
+
+/** Session id and cwd used by every synthesized log. */
+export const SYNTHETIC_SESSION_ID = 'bench-session'
+export const SYNTHETIC_SESSION_CWD = '/bench'
+
+/** Physical directory of the synthesized log below one JSONL root (project slug + session segment). */
+export const SYNTHETIC_SESSION_DIRECTORY = join('--bench--', SYNTHETIC_SESSION_ID)
+
+const TIME_ZERO = 1_700_000_000_000
+
+interface SyntheticEvent {
+  readonly type: string
+  readonly seq: number
+  readonly time: number
+  readonly data: unknown
+  readonly sourceEventSeqs?: readonly number[]
+  readonly surfaceOp?: 'append'
+}
+
+/**
+ * Build the logical released-v0 events for one shape.
+ * @param shape - fixed workload parameters.
+ * @returns dense events in log order.
+ */
+export function synthesizeReleasedV0Events(shape: SyntheticV0LogShape): readonly SyntheticEvent[] {
+  const events: SyntheticEvent[] = []
+  let seq = 0
+  let time = TIME_ZERO
+  const push = (type: string, data: unknown, extra: Partial<SyntheticEvent> = {}): number => {
+    events.push({ type, seq, time, data, ...extra })
+    seq += 1
+    time += 1
+    return seq - 1
+  }
+  const reasoningDeltas = Math.floor(shape.textDeltas / 4)
+  for (let turn = 1; turn <= shape.turns; turn += 1) {
+    push('turn/start', { turn })
+    push('user/message', {
+      id: `user-${turn}`,
+      role: 'user',
+      content: [{ type: 'text', text: `prompt ${turn}` }],
+      source: { kind: 'user' },
+    }, { surfaceOp: 'append' })
+    push('step/start', { turn, step: 1 })
+    const chunkSeqs: number[] = []
+    const chunk = (value: unknown): void => {
+      chunkSeqs.push(push('assistant/chunk', { turn, step: 1, chunk: value }))
+    }
+    chunk({ type: 'block-start', index: 0, blockType: 'reasoning' })
+    let reasoning = ''
+    for (let index = 0; index < reasoningDeltas; index += 1) {
+      const delta = `r${index} `
+      reasoning += delta
+      chunk({ type: 'reasoning-delta', index: 0, text: delta })
+    }
+    chunk({ type: 'block-end', index: 0, block: { type: 'reasoning', text: reasoning } })
+    chunk({ type: 'block-start', index: 1, blockType: 'text' })
+    let text = ''
+    for (let index = 0; index < shape.textDeltas; index += 1) {
+      const delta = `w${index} `
+      text += delta
+      chunk({ type: 'text-delta', index: 1, text: delta })
+    }
+    chunk({ type: 'block-end', index: 1, block: { type: 'text', text } })
+    const usage = { inputTokens: 100, outputTokens: shape.textDeltas }
+    chunk({ type: 'usage', usage })
+    chunk({ type: 'finish', reason: { kind: 'stop' } })
+    push('assistant/message', {
+      turn,
+      step: 1,
+      message: {
+        id: `assistant-${turn}`,
+        role: 'assistant',
+        content: [{ type: 'reasoning', text: reasoning }, { type: 'text', text }],
+        source: { kind: 'model', provider: 'bench', model: 'bench' },
+      },
+      usage,
+    }, { sourceEventSeqs: chunkSeqs, surfaceOp: 'append' })
+    push('step/end', { turn, step: 1 })
+    push('turn/end', { turn, reason: { kind: 'completed' } })
+  }
+  return events
+}
+
+/**
+ * Encode one shape as the released-v0 physical JSONL text (packed chunk rows).
+ * @param shape - fixed workload parameters.
+ * @returns the complete file text plus the logical event count.
+ */
+export function synthesizeReleasedV0LogText(shape: SyntheticV0LogShape): { readonly text: string; readonly events: number } {
+  const events = synthesizeReleasedV0Events(shape)
+  const header = {
+    version: 0,
+    id: SYNTHETIC_SESSION_ID,
+    createdAt: TIME_ZERO,
+    cwd: SYNTHETIC_SESSION_CWD,
+    isSeeded: false,
+    delegationDepth: 0,
+  }
+  const encoded = releasedV0SessionFormatCodec.encodeArtifact(
+    { header, inheritedEventCount: 0, events: events as unknown as readonly SessionFormatEvent[] },
+    { packChunks: true },
+  )
+  const lines = [JSON.stringify(encoded.header), ...encoded.rows.map(row => JSON.stringify(row))]
+  return { text: `${lines.join('\n')}\n`, events: events.length }
+}
+
+/**
+ * Write the synthesized raw v0 log where the JSONL backend expects it.
+ * @param root - JSONL persistence root directory.
+ * @param shape - fixed workload parameters.
+ * @returns the written path, byte length, and logical event count.
+ */
+export async function writeSyntheticReleasedV0Log(
+  root: string,
+  shape: SyntheticV0LogShape,
+): Promise<{ readonly path: string; readonly bytes: number; readonly events: number }> {
+  const { text, events } = synthesizeReleasedV0LogText(shape)
+  const directory = join(root, SYNTHETIC_SESSION_DIRECTORY)
+  await mkdir(directory, { recursive: true })
+  const path = join(directory, 'session.jsonl')
+  await writeFile(path, text)
+  return { path, bytes: Buffer.byteLength(text), events }
+}

+ 10 - 3
scripts/ci-workflow.spec.ts

@@ -66,13 +66,14 @@ describe('CI workflow', () => {
       || !isRecord(workflow.jobs['windows-observational'])
       || !isRecord(workflow.jobs['node-24'])
       || !isRecord(workflow.jobs['node-24-coverage'])
+      || !isRecord(workflow.jobs['node-24-bench'])
       || !isRecord(workflow.jobs['node-24-consumers'])
       || !isRecord(workflow.jobs['node-compat'])
       || !isRecord(workflow.jobs['all-checks-passed'])
       || !isRecord(masterWorkflow.jobs)
       || !isRecord(masterWorkflow.jobs['wine-apt-cache'])
       || !isRecord(masterWorkflow.jobs['serial-windows'])) {
-      throw new TypeError('CI workflow must define windows, windows-build, windows-coverage, windows-native-tests, windows-observational, node-24, node-24-coverage, node-24-consumers, node-compat, and all-checks-passed; ci-master must define wine-apt-cache and serial-windows')
+      throw new TypeError('CI workflow must define windows, windows-build, windows-coverage, windows-native-tests, windows-observational, node-24, node-24-coverage, node-24-bench, node-24-consumers, node-compat, and all-checks-passed; ci-master must define wine-apt-cache and serial-windows')
     }
 
     const windows = workflow.jobs.windows
@@ -84,6 +85,7 @@ describe('CI workflow', () => {
     const serialWindows = masterWorkflow.jobs['serial-windows']
     const node24 = workflow.jobs['node-24']
     const node24Coverage = workflow.jobs['node-24-coverage']
+    const node24Bench = workflow.jobs['node-24-bench']
     const node24Consumers = workflow.jobs['node-24-consumers']
     const nodeCompat = workflow.jobs['node-compat']
     const aggregate = workflow.jobs['all-checks-passed']
@@ -223,15 +225,20 @@ describe('CI workflow', () => {
     // half-close tests are stabilized; observational stays out too.
     expect(aggregate.needs).toContain('windows')
     expect(aggregate.needs).toContain('windows-build')
+    // The benchmark lane is a required verdict input and runs alone so its
+    // wall-clock budgets never share a runner with a concurrent aggregate.
+    expect(aggregate.needs).toContain('node-24-bench')
+    expect(node24Bench.name).toBe('node 24 / benchmarks')
+    expect(node24Bench.env).toBeUndefined()
     expect(aggregate.needs).not.toContain('windows-coverage')
     expect(aggregate.needs).toContain('windows-native-tests')
     expect(aggregate.needs).not.toContain('windows-observational')
     expect(aggregate.needs).not.toContain('serial-windows')
 
-    // Linux failover is a separate switch: the three required Linux workers
+    // Linux failover is a separate switch: the four required Linux workers
     // and the verdict job resolve their pool through DSH_CI_FAILOVER_LINUX,
     // never the Windows switch.
-    for (const [jobName, job] of [['node-24', node24], ['node-24-coverage', node24Coverage], ['node-24-consumers', node24Consumers]] as const) {
+    for (const [jobName, job] of [['node-24', node24], ['node-24-coverage', node24Coverage], ['node-24-bench', node24Bench], ['node-24-consumers', node24Consumers]] as const) {
       expect(typeof job['runs-on']).toBe('string')
       expect(job['runs-on'], `${jobName} runs-on must use the Linux failover switch`).toContain('DSH_CI_FAILOVER_LINUX')
       expect(job['runs-on'], `${jobName} runs-on must not use the Windows failover switch`).not.toContain('DSH_CI_FAILOVER_WINDOWS')

+ 1 - 0
scripts/run-gates.spec.ts

@@ -141,6 +141,7 @@ describe('gate graph validation', () => {
     'ci-static',
     'ci-lint-contracts-ready',
     'ci-coverage',
+    'ci-bench',
     'ci-snapshot',
     'ci-artifacts',
     'ci-consumers',

+ 5 - 1
scripts/run-gates.ts

@@ -27,6 +27,7 @@ export type Mode =
   | 'ci-static'
   | 'ci-lint-contracts-ready'
   | 'ci-coverage'
+  | 'ci-bench'
   | 'ci-snapshot'
   | 'ci-artifacts'
   | 'ci-consumers'
@@ -137,6 +138,7 @@ function parseMode(raw: string | undefined): Mode {
     case 'ci-static':
     case 'ci-lint-contracts-ready':
     case 'ci-coverage':
+    case 'ci-bench':
     case 'ci-snapshot':
     case 'ci-artifacts':
     case 'ci-consumers':
@@ -151,7 +153,7 @@ function parseMode(raw: string | undefined): Mode {
       return raw
     default:
       throw new Error(
-        `run-gates: expected mode ci-primary | ci-linux-primary | ci-static | ci-lint-contracts-ready | ci-coverage | ci-snapshot | ci-artifacts | ci-consumers | ci-windows-blocking | ci-windows-complete | ci-windows-observational | node-compat | check-all | hygiene | doc-sync | doc-quick, got ${JSON.stringify(raw)}.`,
+        `run-gates: expected mode ci-primary | ci-linux-primary | ci-static | ci-lint-contracts-ready | ci-coverage | ci-bench | ci-snapshot | ci-artifacts | ci-consumers | ci-windows-blocking | ci-windows-complete | ci-windows-observational | node-compat | check-all | hygiene | doc-sync | doc-quick, got ${JSON.stringify(raw)}.`,
       )
   }
 }
@@ -242,6 +244,8 @@ export function gatesForMode(selected: Mode): Gate[] {
       ]
     case 'ci-coverage':
       return coverageGates()
+    case 'ci-bench':
+      return [pnpmScript('bench', 'test:bench', { label: 'performance benchmarks' })]
     case 'ci-snapshot':
       return [ciBuildGate(), snapshotGate()]
     case 'ci-artifacts':

+ 26 - 0
vitest.bench.config.ts

@@ -0,0 +1,26 @@
+import tsconfigPaths from 'vite-tsconfig-paths'
+import { defineConfig } from 'vitest/config'
+import { standardDecoratorPlugin, vitestExecArgv } from './vitest.shared.ts'
+
+/**
+ * CI performance gate. Every `*.bench.ts` file synthesizes its own input from
+ * fixed parameters, measures one owner-visible path, and fails when a
+ * documented time or heap budget is exceeded. Files run one at a time so a
+ * measurement never shares the CPU with another benchmark.
+ */
+export default defineConfig({
+  plugins: [tsconfigPaths({ projects: ['./tsconfig.base.json'] }), standardDecoratorPlugin()],
+  test: {
+    execArgv: vitestExecArgv,
+    setupFiles: ['./scripts/test-proxy-environment.ts'],
+    include: [
+      'packages/*/*/tests/**/*.bench.ts',
+      'packages/*/*/tests/**/*.bench.client.ts',
+    ],
+    fileParallelism: false,
+    maxWorkers: 1,
+    testTimeout: 600_000,
+    hookTimeout: 120_000,
+    disableConsoleIntercept: true,
+  },
+})