Просмотр исходного кода

test(perf): run benchmarks against built artifacts

imccyu 1 месяц назад
Родитель
Сommit
356a332195

+ 2 - 2
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
-2026-09-04-session-open-performance-gate.md: 30c4345a7c1d62c8af0d3d66505e8614935e9618
-2026-09-04-session-open-performance-gate.zh.md: 61e7ed1e87ffcd5c083c6bd4ab6516dbd7031837
+2026-09-04-session-open-performance-gate.md: 8fc2598b7739ef26767a31300e694726a37febb3
+2026-09-04-session-open-performance-gate.zh.md: 57b5179ee9efccdcd7a07d0083e4b55377a666dd

+ 11 - 9
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md

@@ -12,7 +12,7 @@ Measuring only `SessionPersistence.open()` does not stably describe the result f
 
 ## Decision
 
-Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench` → `vitest.bench.config.ts`. The job runs the benchmark lane alone; Vitest runs one file at a time and only prepares input, starts measurement children, aggregates results, and enforces budgets.
+Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench`. The command first builds workspace libraries and dedicated benchmark workers, then invokes `vitest.bench.config.ts`. The job runs the benchmark lane alone; Vitest runs one file at a time and only prepares input, starts measurement children, aggregates results, and enforces budgets. Every timed CPU path executes compiled JavaScript under plain Node with `NODE_OPTIONS` removed and no TypeScript loader; bare workspace imports therefore resolve through package exports to built `lib/` entries.
 
 Required performance gates live under top-level `benchmarks/`, grouped by measured user path rather than package ownership. Host files use `*.bench.ts`, Client-face files use `*.bench.client.ts`, and scenario-specific workers and fixtures stay beside their benchmark without a benchmark suffix. Package-local `.perf.ts` files remain non-gating diagnostics; `scripts/` owns orchestration rather than benchmark cases.
 
@@ -20,7 +20,7 @@ The Session benchmarks synthesize a released-v0 input from fixed parameters: 200
 
 Every Session endpoint runs at two user-lifecycle points. `first-open` starts with only the released V0 generation and therefore includes migration and successor publication. Setup produces `post-upgrade-reopen` once through that same production migration outside measurement, then copies both the unchanged V0 predecessor and published V2 successor into each sample root. Reopen samples use a fresh process, so they measure an upgraded user's later disk open without migration or process-local caches.
 
-Each access-kind and endpoint sample runs in a fresh Node child process. Module imports, Host service initialization, and fixture preparation finish before measurement; the measured process performs no extra parse warm-up. Normal-heap mode runs five independent samples, reports every sample plus minimum, median, and maximum, and enforces access-specific fixed budgets against the median. Another child runs the same path under a fixed 128 MB old-space limit and checks only that it completes; extra GC caused by the constrained heap does not enter the normal timing baseline.
+Each access-kind and endpoint sample runs in a fresh compiled Node child process. Module imports, Host service initialization, and fixture preparation finish before measurement; the measured process performs no extra parse warm-up. Normal-heap mode runs five independent samples, reports every sample plus minimum, median, and maximum, and enforces access-specific fixed budgets against the median. Another child runs the same path under a fixed 128 MB old-space limit and checks only that it completes; extra GC caused by the constrained heap does not enter the normal timing baseline.
 
 The lane contains three independent Session-opening benchmarks and retains the Client-fold benchmark:
 
@@ -37,7 +37,7 @@ Normal-heap mode performs a fixed pair of explicit garbage collections after Hos
 
 The performance gate does not duplicate semantic assertions owned by functional tests; it requires only that the target call completes and reaches its measured endpoint. The Client-fold benchmark continues to use the real `ConversationNodeAssembler` and every Chat Definition, and requires both the large window's absolute time and its scaling relative to the small window to remain below fixed budgets.
 
-Budgets use repeated measurements of the final implementation on the target CI runner, with enough margin for runner noise while remaining below the known regression. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
+Budgets are layered. The first-open `open` limit, constrained-heap completion checks, and Client-fold scaling bound reject the known regressions; first-history and Agent-resume limits are broader orchestration ceilings that catch additional overhead without duplicating those component verdicts. Repeated measurements on the target CI runner leave margin for runner noise. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
 
 ## Calibration evidence
 
@@ -45,12 +45,12 @@ The comparison is orthogonal by user lifecycle, not by artifact representation.
 
 Five-sample medians on the same Node 24 reference machine establish the positive and negative controls:
 
-| Access kind | Implementation | Four-phase total | First history | Agent resume | 128 MB old space |
-|---|---|---:|---:|---:|---|
-| First open | Pre-stack reference | 249.0 ms | 253.8 ms | 100.7 ms | Completes |
-| First open | Repeated-snapshot regression | 4,197.5 ms | 4,284.8 ms | 4,197.9 ms | Exhausts heap |
-| Post-upgrade reopen | Pre-stack reference | 251.1 ms | 253.8 ms | 100.7 ms | Completes |
-| Post-upgrade reopen | Repeated-snapshot regression | 49.2 ms | 50.4 ms | 43.8 ms | Completes |
+| Access kind | Implementation | Four-phase total | First history | Agent resume | Agent retained heap | 128 MB old space |
+|---|---|---:|---:|---:|---:|---|
+| First open | Pre-stack reference | 249.0 ms | 253.8 ms | 100.7 ms | 26.1 MB | Completes |
+| First open | Repeated-snapshot regression | 4,197.5 ms | 4,284.8 ms | 4,197.9 ms | 4.4 MB | Exhausts heap |
+| Post-upgrade reopen | Pre-stack reference | 251.1 ms | 253.8 ms | 100.7 ms | 26.1 MB | Completes |
+| Post-upgrade reopen | Repeated-snapshot regression | 49.2 ms | 50.4 ms | 43.8 ms | 4.5 MB | Completes |
 
 The pre-stack implementation keeps V0 as its current format, so first open does not change its on-disk representation; its native V0 first-history and Agent-resume measurements therefore apply to both lifecycle rows.
 
@@ -66,6 +66,8 @@ The pre-stack implementation keeps V0 as its current format, so first open does
 
 **Add fine-grained timing instrumentation inside production implementations.** Rejected because those probes would expand production APIs and couple the benchmark to implementation details. Tests use existing service and object boundaries; costs that those boundaries cannot attribute remain part of the end-to-end result.
 
+**Run measured workers from TypeScript source.** Rejected because a source loader changes module resolution and startup behavior, and causes nested workers to select source-only bootstrap paths. Vitest remains an unmeasured orchestrator; every timed worker executes the build output exactly as plain Node consumers do.
+
 **Use only time budgets or only post-GC memory.** Rejected because time does not reveal memory regressions, while endpoint live memory cannot expose transient migration spikes. Normal-heap post-GC deltas and constrained-heap completion cover the two risks separately.
 
 **Benchmark the real recorded corpus.** Rejected because corpus fixtures stay small by policy, recorded material must not become benchmark input, and re-recording would silently move the workload.

+ 10 - 8
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md

@@ -12,7 +12,7 @@ Session format v2 的推出改变了两条成本随模型输出增长的路径
 
 ## 决定
 
-Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench` → `vitest.bench.config.ts`。该 job 单独运行 benchmark lane;Vitest 逐文件运行,测试进程只负责准备输入、启动测量子进程、汇总结果和执行预算断言。
+Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench`。该命令先构建 workspace library 和专用 benchmark worker,再调用 `vitest.bench.config.ts`。该 job 单独运行 benchmark lane;Vitest 逐文件运行,只负责准备输入、启动测量子进程、汇总结果和执行预算断言。每条被计时的 CPU 路径都以纯 Node 执行编译后的 JavaScript,并移除 `NODE_OPTIONS` 且不加载 TypeScript runtime;workspace 裸导入因此通过 package exports 解析到构建后的 `lib/` 入口。
 
 必需性能 gate 位于顶层 `benchmarks/`,按被测用户路径而非 package 归属组织。Host 文件使用 `*.bench.ts`,Client 面文件使用 `*.bench.client.ts`,场景专属 worker 与 fixture 留在对应 benchmark 旁且不带 benchmark 后缀。包内 `.perf.ts` 文件仍是非门禁诊断;`scripts/` 负责编排而不承载 benchmark case。
 
@@ -20,7 +20,7 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 每个 Session endpoint 都针对用户生命周期中的两个时点运行。`first-open` 最初只有 released V0 generation,因此包含 migration 与后继 generation 发布。测试准备阶段在计时外通过同一套生产 migration 生成一次 `post-upgrade-reopen`,再把未改动的 V0 前代和已发布的 V2 后继一起复制到每个样本目录。Reopen 样本使用全新进程,因此测量用户升级完成后的磁盘再次打开,不包含 migration 或进程内 cache。
 
-每个 access kind 与 endpoint 的样本都在全新 Node 子进程中运行。模块加载、Host 服务初始化和 fixture 准备在测量开始前完成;测量进程不执行额外的预热解析。正常堆模式运行五个独立样本,报告全部样本及最小值、中位数和最大值,并以中位数执行各访问状态独立的固定预算。另一个子进程使用固定 128 MB old-space 上限运行同一路径,只判断能否完成;低堆限制引起的额外 GC 不进入正常时间基线。
+每个 access kind 与 endpoint 的样本都在全新、已编译的 Node 子进程中运行。模块加载、Host 服务初始化和 fixture 准备在测量开始前完成;测量进程不执行额外的预热解析。正常堆模式运行五个独立样本,报告全部样本及最小值、中位数和最大值,并以中位数执行各访问状态独立的固定预算。另一个子进程使用固定 128 MB old-space 上限运行同一路径,只判断能否完成;低堆限制引起的额外 GC 不进入正常时间基线。
 
 该 lane 包含三个独立的 Session 打开 benchmark,并保留 Client fold benchmark:
 
@@ -37,7 +37,7 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 性能 gate 不重复功能测试的内容断言,只要求目标调用完成并到达对应的可观察终点。Client fold benchmark 继续使用真实 `ConversationNodeAssembler` 与全部 Chat Definition,要求大窗口的绝对时间和相对小窗口的缩放比均低于固定预算。
 
-预算以最终实现于目标 CI runner 上的多次样本为基线,并保留足以吸收 runner 波动、但仍能区分已知退化的余量。栈前参考提交固定为 `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5`,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
+预算分层发挥作用。First-open `open` 上限、受限堆完成性检查和 Client fold 缩放上限会拒绝已知退化;首屏历史与 Agent resume 上限是更宽松的编排总量限制,用于发现额外开销而不重复组件判定。目标 CI runner 上的重复测量为机器波动保留余量。栈前参考提交固定为 `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5`,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
 
 ## 校准证据
 
@@ -45,11 +45,11 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 同一台 Node 24 参考机器上的五次样本中位数构成正反例:
 
-| Access kind | 实现 | 四阶段总时间 | 首屏历史 | Agent resume | 128 MB old space |
-|---|---|---:|---:|---:|---|
-| First open | 栈前参考版本 | 249.0 ms | 253.8 ms | 100.7 ms | 完成 |
-| First open | 重复 snapshot 退化实现 | 4,197.5 ms | 4,284.8 ms | 4,197.9 ms | 堆耗尽 |
-| Post-upgrade reopen | 栈前参考版本 | 251.1 ms | 253.8 ms | 100.7 ms | 完成 |
+| Access kind | 实现 | 四阶段总时间 | 首屏历史 | Agent resume | Agent GC 后增量堆 | 128 MB old space |
+|---|---|---:|---:|---:|---:|---|
+| First open | 栈前参考版本 | 249.0 ms | 253.8 ms | 100.7 ms | 26.1 MB | 完成 |
+| First open | 重复 snapshot 退化实现 | 4,197.5 ms | 4,284.8 ms | 4,197.9 ms | 4.4 MB | 堆耗尽 |
+| Post-upgrade reopen | 栈前参考版本 | 251.1 ms | 253.8 ms | 100.7 ms | 26.1 MB | 完成 |
 | Post-upgrade reopen | 重复 snapshot 退化实现 | 49.2 ms | 50.4 ms | 43.8 ms | 完成 |
 
 栈前实现以 V0 作为当前格式,因此 first open 不改变磁盘表示;它的原生 V0 首屏历史与 Agent resume 测量同时适用于两个生命周期行。
@@ -66,6 +66,8 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 **在生产实现内部添加细粒度计时桩。** 拒绝:这些桩会扩大生产接口并让 benchmark 与实现细节耦合。测试只使用既有服务和对象边界;无法由这些边界解释的成本保留在端到端结果中。
 
+**从 TypeScript 源码运行被测 worker。** 拒绝:源码 loader 会改变模块解析与启动行为,并使嵌套 worker 选择仅适用于源码的启动路径。Vitest 仍可作为不计时的编排层;每个被计时的 worker 都像纯 Node 消费方一样执行构建产物。
+
 **只用时间预算或只看 GC 后内存。** 拒绝:时间无法发现内存退化,终点存活内存也看不到迁移期间的瞬时爆发。正常堆的 GC 后增量与受限堆的完成性分别覆盖两类风险。
 
 **用真实录制语料做 benchmark。** 拒绝:语料 fixture 按策略保持小体量,录制材料不得成为 benchmark 输入,且其重新录制会静默移动 workload。

+ 1 - 1
AGENTS.md

@@ -50,7 +50,7 @@ packages/    @deepseek-ai/dsh-<pkg> workspaces at packages/<group>/<pkg>/
   util/        zero-dependency utilities
 python/      Python SDK/runtime (see python/README.md)
 native/      @deepseek-ai/node-addon-landlock-run source of record (see native/README.md)
-benchmarks/  cross-package performance gates
+benchmarks/  performance gates
 .agents/     Agent workflows and Agent Notes (`notes/`)
 docs/        architecture, generated catalogs, postmortems, cookbook (see docs/AGENTS.md)
 scripts/     gates and generators

+ 2 - 0
benchmarks/AGENTS.md

@@ -4,9 +4,11 @@ This tree owns required, repository-level performance gates whose measured user
 
 - Organize benchmarks by measured user path, one directory per path. Do not mirror the package tree.
 - Host cases use `*.bench.ts`; Client-face cases use `*.bench.client.ts`. Worker, fixture, and support modules do not carry a benchmark suffix.
+- `test:bench` builds workspace libraries and `.dsh-build/benchmarks/` workers before Vitest orchestration. Timed CPU work runs in those workers under plain Node, without a TypeScript loader; runtime package imports must resolve to built `lib/` entries.
 - Synthesize fixed inputs from reviewed constants. Never use recorded Sessions, user material, ambient repositories, or network services.
 - Run process-level wall-clock and retained-memory samples in fresh children with private `mkdtemp` roots. Pure synchronous folds create a fresh object graph per sample and must not mutate process-global state. Bound every child, await exit, and remove owned roots after failure as well as success.
 - Report enough raw and aggregate measurements to explain each verdict, including whether a budget uses a median, minimum, absolute value, or ratio. Enforce reviewed source constants; environment variables must not override performance budgets.
 - Keep scenario-specific support beside its benchmark. Move a helper into `benchmarks/support/` only after at least two benchmark directories require the same behavior.
 - Exercise production entry points. Do not copy product algorithms, add production exports solely for measurement, or turn benchmark completion into duplicate semantic assertions.
+- A compiled worker may bundle a private integration adapter when no public Node export exposes the measured user path. Keep package imports external so product services resolve through their built package exports.
 - Record the workload, timing boundary, memory endpoint, calibration reference, alternatives, and known exclusions in the owning Agent Note.

+ 46 - 190
benchmarks/conversation-fold/conversation-fold.bench.client.ts

@@ -1,218 +1,74 @@
-/**
- * Performance gate for the cold Client fold of a large Session format v2
- * history window: every registered Chat Definition runs over a synthesized
- * window in which each assistant reply embeds its compact stream. The gate
- * bounds the wall time and requires the fold to scale with the number of
- * compact stream records rather than with the number of streamed deltas.
- */
+/** Required performance budget for the compiled cold Client conversation fold. */
 
+import { join } from 'node:path'
 import { describe, expect, it } from 'vitest'
-import { AssistantStreamAccumulator } from '@deepseek-ai/dsh-llm/assistant-stream'
-import type { StreamChunk } from '@deepseek-ai/dsh-llm'
-import type { SessionEvent } from '@deepseek-ai/dsh-session/types'
-import type { ChatSnapshot } from '@deepseek-ai/dsh-client-ui-chat/client'
-import type { SessionEventLikeEntry } from '@deepseek-ai/dsh-api-session-controller/client'
 import {
-  ConversationNodeAssembler,
-  inspectRequestPrompt,
-  type ConversationNodeDefinition,
-  type ConversationViewDefinition,
-} from '@deepseek-ai/dsh-client-ui-conversation/client'
-import { assistantDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/assistant.ts'
-import { chatViewDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/chat-snapshot-builder.ts'
-import { commandDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/command.ts'
-import { compactionDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/compaction.ts'
-import { unknownFallbackDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/fallback.ts'
-import { nextStepInboxDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/inbox.ts'
-import { messageDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/message.ts'
-import { requestPromptDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/request-prompt.ts'
-import { retryDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/retry.ts'
-import { toolDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/tool.ts'
-import { turnErrorDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-error.ts'
-import { turnMaxTokensDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-max-tokens.ts'
-import { turnProcessDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-process.ts'
-import { turnTailDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-tail.ts'
+  runBuiltBenchmarkWorker,
+  type BuiltBenchmarkWorkerRun,
+} from '../support/built-worker.ts'
+import type { ConversationFoldWorkerReport } from './conversation-fold.worker.client.ts'
 
 /** Replies in the folded window; each carries one reasoning block and one text block. */
 const TURNS = 200
-
-/** Text deltas per reply in the large workload; the reply also streams `deltas / 4` reasoning deltas. */
+/** Text deltas per reply in the large workload; each reply adds one quarter as many reasoning deltas. */
 const LARGE_DELTAS = 2_000
-
 /** Text deltas per reply in the small workload used as the scaling reference. */
 const SMALL_DELTAS = 100
+/** Fresh object graphs measured in one compiled worker; the fastest sample removes scheduler delay. */
+const ATTEMPTS = 3
+/** A stuck fold worker is reaped before the outer benchmark deadline. */
+const WORKER_TIMEOUT_MS = 60_000
 
 /**
- * Wall-clock budget for folding the large window (200 replies, 500,000 streamed
- * deltas compacted into 800 stream records). The pre-stack fold processed the
- * equivalent packed chunk rows in a few milliseconds; the budget leaves room
- * for the complete Definition set and slower CI hosts while staying below the
- * per-delta replay that needed hundreds of milliseconds for this window.
+ * The large window contains 500,000 streamed deltas compacted into 1,600
+ * stream records. The budget separates the record-proportional fold from the
+ * per-delta replay that needs hundreds of milliseconds for the same window.
  */
 const LARGE_FOLD_BUDGET_MS = 150
 
 /**
- * Maximum ratio between folding the large and the small window. Both windows
- * hold the same number of events and compact records, so a fold that scales
- * with records plus the joined text stays a few times the small fold
- * (about 2.5× measured); a fold that replays every delta grows with the 20×
- * delta count (about 11× measured).
+ * Both windows contain equal event and compact-record counts. A fold over
+ * records plus joined text measures about 2.5×; replaying every delta measures
+ * about 11× as the delta count grows 20×.
  */
 const MAX_DELTA_SCALING = 5
 
-/** Attempts per workload; the gate compares minima so scheduler noise only adds. */
-const ATTEMPTS = 3
-
-const TIME_ZERO = 1_700_000_000_000
-
-class BenchEventDefinitions {
-  readonly definitions: readonly ConversationNodeDefinition[] = [
-    nextStepInboxDefinition,
-    messageDefinition,
-    requestPromptDefinition(inspectRequestPrompt),
-    assistantDefinition,
-    turnProcessDefinition,
-    toolDefinition,
-    commandDefinition,
-    compactionDefinition,
-    retryDefinition,
-    turnErrorDefinition,
-    turnMaxTokensDefinition,
-    turnTailDefinition,
-  ]
-
-  entries(): readonly ConversationNodeDefinition[] {
-    return this.definitions
-  }
-
-  fallbackEntry(): ConversationNodeDefinition {
-    return unknownFallbackDefinition
-  }
-}
-
-class BenchViewDefinitions {
-  entries(): readonly ConversationViewDefinition[] {
-    return [chatViewDefinition]
-  }
-}
-
-function entry(seq: number, type: string, data: unknown, extra: Record<string, unknown> = {}): SessionEventLikeEntry {
-  return {
-    type: 'event',
-    event: { seq, time: TIME_ZERO + seq, type, data, ...extra } as unknown as SessionEvent,
-  }
-}
-
-/**
- * Synthesize one v2 history window: `turns` completed replies whose compact
- * streams are accumulated from `deltas` text deltas and `deltas / 4`
- * reasoning deltas each.
- */
-function synthesizeWindow(turns: number, deltas: number): { readonly entries: readonly SessionEventLikeEntry[]; readonly records: number } {
-  const entries: SessionEventLikeEntry[] = []
-  let seq = 0
-  let records = 0
-  const push = (type: string, data: unknown, extra: Record<string, unknown> = {}): void => {
-    entries.push(entry(seq, type, data, extra))
-    seq += 1
-  }
-  const reasoningDeltas = Math.floor(deltas / 4)
-  for (let turn = 1; turn <= turns; turn += 1) {
-    push('turn/start', { turn })
-    push('user/message', {
-      id: `user-${String(turn)}`,
-      role: 'user',
-      content: [{ type: 'text', text: `prompt ${String(turn)}` }],
-      source: { kind: 'user' },
-    }, { surfaceOp: 'append' })
-    push('step/start', { turn, step: 1 })
-    const accumulator = new AssistantStreamAccumulator()
-    let time = TIME_ZERO + seq * 1_000
-    const stream = (chunk: StreamChunk): void => {
-      accumulator.push({ time, chunk })
-      time += 1
-    }
-    stream({ type: 'block-start', index: 0, blockType: 'reasoning' })
-    let reasoning = ''
-    for (let index = 0; index < reasoningDeltas; index += 1) {
-      const delta = `r${String(index)} `
-      reasoning += delta
-      stream({ type: 'reasoning-delta', index: 0, text: delta })
-    }
-    stream({ type: 'block-end', index: 0, block: { type: 'reasoning', text: reasoning } })
-    stream({ type: 'block-start', index: 1, blockType: 'text' })
-    let text = ''
-    for (let index = 0; index < deltas; index += 1) {
-      const delta = `w${String(index)} `
-      text += delta
-      stream({ type: 'text-delta', index: 1, text: delta })
-    }
-    stream({ type: 'block-end', index: 1, block: { type: 'text', text } })
-    const usage = { inputTokens: 100, outputTokens: deltas }
-    stream({ type: 'usage', usage })
-    stream({ type: 'finish', reason: { kind: 'stop' } })
-    const snapshot = accumulator.snapshot()
-    records += snapshot.length
-    push('assistant/message', {
-      turn,
-      step: 1,
-      message: {
-        id: `assistant-${String(turn)}`,
-        role: 'assistant',
-        content: [{ type: 'reasoning', text: reasoning }, { type: 'text', text }],
-        source: { kind: 'model', provider: 'bench', model: 'bench' },
-      },
-      usage,
-      stream: snapshot,
-    }, { surfaceOp: 'append' })
-    push('step/end', { turn, step: 1 })
-    push('turn/end', { turn, reason: { kind: 'completed' } })
-  }
-  return { entries, records }
-}
-
-function foldOnce(entries: readonly SessionEventLikeEntry[]): { readonly ms: number; readonly nodes: number } {
-  const started = performance.now()
-  const assembler = new ConversationNodeAssembler(new BenchEventDefinitions(), new BenchViewDefinitions())
-  assembler.replaceWindow(entries, false)
-  assembler.activateTarget('chat')
-  const snapshot = assembler.snapshot('chat') as ChatSnapshot | undefined
-  return { ms: performance.now() - started, nodes: snapshot?.order.length ?? 0 }
-}
-
-function bestOf(entries: readonly SessionEventLikeEntry[]): { readonly ms: number; readonly nodes: number } {
-  let best = foldOnce(entries)
-  for (let attempt = 1; attempt < ATTEMPTS; attempt += 1) {
-    const next = foldOnce(entries)
-    if (next.ms < best.ms) best = next
-  }
-  return best
+const WORKER = join(
+  import.meta.dirname,
+  '..',
+  '..',
+  '.dsh-build',
+  'benchmarks',
+  'conversation-fold',
+  'conversation-fold.worker.js',
+)
+
+function requireReport(
+  run: BuiltBenchmarkWorkerRun<ConversationFoldWorkerReport>,
+): ConversationFoldWorkerReport {
+  if (run.report !== undefined) return run.report
+  const stderrLines = run.stderr.trim().split('\n')
+  throw new Error(
+    `conversation-fold worker failed: exit=${String(run.exitCode)}, signal=${String(run.signal)}, `
+    + `timedOut=${String(run.timedOut)}\n${stderrLines.slice(-10).join('\n')}`,
+  )
 }
 
 describe('cold Chat fold of a large v2 history window', () => {
-  it(`folds ${String(TURNS)} replies with ${String(LARGE_DELTAS)} deltas each within ${String(LARGE_FOLD_BUDGET_MS)} ms and scales with compact records`, () => {
-    const small = synthesizeWindow(TURNS, SMALL_DELTAS)
-    const large = synthesizeWindow(TURNS, LARGE_DELTAS)
-    expect(large.entries.length).toBe(small.entries.length)
-    expect(large.records).toBe(small.records)
-
-    const smallFold = bestOf(small.entries)
-    const largeFold = bestOf(large.entries)
-    const scaling = largeFold.ms / Math.max(smallFold.ms, 1)
+  it(`folds ${String(TURNS)} replies with ${String(LARGE_DELTAS)} deltas each within ${String(LARGE_FOLD_BUDGET_MS)} ms and scales with compact records`, async () => {
+    const report = requireReport(await runBuiltBenchmarkWorker<ConversationFoldWorkerReport>({
+      worker: WORKER,
+      args: [String(TURNS), String(SMALL_DELTAS), String(LARGE_DELTAS), String(ATTEMPTS)],
+      timeoutMs: WORKER_TIMEOUT_MS,
+    }))
     console.log(JSON.stringify({
       benchmark: 'conversation-fold/large-window',
-      events: large.entries.length,
-      compactRecords: large.records,
-      streamedDeltas: TURNS * (LARGE_DELTAS + Math.floor(LARGE_DELTAS / 4)),
-      chatNodes: largeFold.nodes,
-      smallFoldMs: Math.round(smallFold.ms * 10) / 10,
-      largeFoldMs: Math.round(largeFold.ms * 10) / 10,
-      scaling: Math.round(scaling * 100) / 100,
+      ...report,
       budgetMs: LARGE_FOLD_BUDGET_MS,
       maxScaling: MAX_DELTA_SCALING,
     }))
-    expect(largeFold.nodes).toBeGreaterThan(0)
-    expect(largeFold.ms).toBeLessThanOrEqual(LARGE_FOLD_BUDGET_MS)
-    expect(scaling).toBeLessThanOrEqual(MAX_DELTA_SCALING)
+    expect(report.chatNodes).toBeGreaterThan(0)
+    expect(report.largeFoldMs).toBeLessThanOrEqual(LARGE_FOLD_BUDGET_MS)
+    expect(report.scaling).toBeLessThanOrEqual(MAX_DELTA_SCALING)
   })
 })

+ 203 - 0
benchmarks/conversation-fold/conversation-fold.worker.client.ts

@@ -0,0 +1,203 @@
+/** Compiled worker for the cold Client conversation-fold benchmark. */
+
+import { performance } from 'node:perf_hooks'
+import { AssistantStreamAccumulator } from '@deepseek-ai/dsh-llm/assistant-stream'
+import type { StreamChunk } from '@deepseek-ai/dsh-llm'
+import type { SessionEvent } from '@deepseek-ai/dsh-session/types'
+import type { ChatSnapshot } from '@deepseek-ai/dsh-client-ui-chat/client'
+import type { SessionEventLikeEntry } from '@deepseek-ai/dsh-api-session-controller/client'
+// These Client-only fold modules have no plain-Node package export and are compiled into this worker.
+import { ConversationNodeAssembler } from '../../packages/client/ui-conversation/src/client/conversation/assembler.ts'
+import { inspectRequestPrompt } from '../../packages/client/ui-conversation/src/client/contract/request-inspection.ts'
+import type {
+  ConversationNodeDefinition,
+  ConversationViewDefinition,
+} from '../../packages/client/ui-conversation/src/client/contract/conversation.ts'
+import { assistantDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/assistant.ts'
+import { chatViewDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/chat-snapshot-builder.ts'
+import { commandDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/command.ts'
+import { compactionDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/compaction.ts'
+import { unknownFallbackDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/fallback.ts'
+import { nextStepInboxDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/inbox.ts'
+import { messageDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/message.ts'
+import { requestPromptDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/request-prompt.ts'
+import { retryDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/retry.ts'
+import { toolDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/tool.ts'
+import { turnErrorDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-error.ts'
+import { turnMaxTokensDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-max-tokens.ts'
+import { turnProcessDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-process.ts'
+import { turnTailDefinition } from '../../packages/client/ui-chat/src/client/conversation-nodes/turn-tail.ts'
+import { assertBuiltBenchmarkRuntime } from '../support/built-worker.ts'
+
+const TIME_ZERO = 1_700_000_000_000
+
+/** Result emitted by the compiled conversation-fold worker. */
+export interface ConversationFoldWorkerReport {
+  readonly events: number
+  readonly compactRecords: number
+  readonly streamedDeltas: number
+  readonly chatNodes: number
+  readonly smallFoldMs: number
+  readonly largeFoldMs: number
+  readonly scaling: number
+}
+
+class BenchEventDefinitions {
+  readonly definitions: readonly ConversationNodeDefinition[] = [
+    nextStepInboxDefinition,
+    messageDefinition,
+    requestPromptDefinition(inspectRequestPrompt),
+    assistantDefinition,
+    turnProcessDefinition,
+    toolDefinition,
+    commandDefinition,
+    compactionDefinition,
+    retryDefinition,
+    turnErrorDefinition,
+    turnMaxTokensDefinition,
+    turnTailDefinition,
+  ]
+
+  entries(): readonly ConversationNodeDefinition[] {
+    return this.definitions
+  }
+
+  fallbackEntry(): ConversationNodeDefinition {
+    return unknownFallbackDefinition
+  }
+}
+
+class BenchViewDefinitions {
+  entries(): readonly ConversationViewDefinition[] {
+    return [chatViewDefinition]
+  }
+}
+
+function entry(seq: number, type: string, data: unknown, extra: Record<string, unknown> = {}): SessionEventLikeEntry {
+  return {
+    type: 'event',
+    event: { seq, time: TIME_ZERO + seq, type, data, ...extra } as unknown as SessionEvent,
+  }
+}
+
+function synthesizeWindow(
+  turns: number,
+  deltas: number,
+): { readonly entries: readonly SessionEventLikeEntry[]; readonly records: number } {
+  const entries: SessionEventLikeEntry[] = []
+  let seq = 0
+  let records = 0
+  const push = (type: string, data: unknown, extra: Record<string, unknown> = {}): void => {
+    entries.push(entry(seq, type, data, extra))
+    seq += 1
+  }
+  const reasoningDeltas = Math.floor(deltas / 4)
+  for (let turn = 1; turn <= turns; turn += 1) {
+    push('turn/start', { turn })
+    push('user/message', {
+      id: `user-${String(turn)}`,
+      role: 'user',
+      content: [{ type: 'text', text: `prompt ${String(turn)}` }],
+      source: { kind: 'user' },
+    }, { surfaceOp: 'append' })
+    push('step/start', { turn, step: 1 })
+    const accumulator = new AssistantStreamAccumulator()
+    let time = TIME_ZERO + seq * 1_000
+    const stream = (chunk: StreamChunk): void => {
+      accumulator.push({ time, chunk })
+      time += 1
+    }
+    stream({ type: 'block-start', index: 0, blockType: 'reasoning' })
+    let reasoning = ''
+    for (let index = 0; index < reasoningDeltas; index += 1) {
+      const delta = `r${String(index)} `
+      reasoning += delta
+      stream({ type: 'reasoning-delta', index: 0, text: delta })
+    }
+    stream({ type: 'block-end', index: 0, block: { type: 'reasoning', text: reasoning } })
+    stream({ type: 'block-start', index: 1, blockType: 'text' })
+    let text = ''
+    for (let index = 0; index < deltas; index += 1) {
+      const delta = `w${String(index)} `
+      text += delta
+      stream({ type: 'text-delta', index: 1, text: delta })
+    }
+    stream({ type: 'block-end', index: 1, block: { type: 'text', text } })
+    const usage = { inputTokens: 100, outputTokens: deltas }
+    stream({ type: 'usage', usage })
+    stream({ type: 'finish', reason: { kind: 'stop' } })
+    const snapshot = accumulator.snapshot()
+    records += snapshot.length
+    push('assistant/message', {
+      turn,
+      step: 1,
+      message: {
+        id: `assistant-${String(turn)}`,
+        role: 'assistant',
+        content: [{ type: 'reasoning', text: reasoning }, { type: 'text', text }],
+        source: { kind: 'model', provider: 'bench', model: 'bench' },
+      },
+      usage,
+      stream: snapshot,
+    }, { surfaceOp: 'append' })
+    push('step/end', { turn, step: 1 })
+    push('turn/end', { turn, reason: { kind: 'completed' } })
+  }
+  return { entries, records }
+}
+
+function foldOnce(entries: readonly SessionEventLikeEntry[]): { readonly ms: number; readonly nodes: number } {
+  const started = performance.now()
+  const assembler = new ConversationNodeAssembler(new BenchEventDefinitions(), new BenchViewDefinitions())
+  assembler.replaceWindow(entries, false)
+  assembler.activateTarget('chat')
+  const snapshot = assembler.snapshot('chat') as ChatSnapshot | undefined
+  return { ms: performance.now() - started, nodes: snapshot?.order.length ?? 0 }
+}
+
+function bestOf(
+  entries: readonly SessionEventLikeEntry[],
+  attempts: number,
+): { readonly ms: number; readonly nodes: number } {
+  let best = foldOnce(entries)
+  for (let attempt = 1; attempt < attempts; attempt += 1) {
+    const next = foldOnce(entries)
+    if (next.ms < best.ms) best = next
+  }
+  return best
+}
+
+function positiveInteger(value: string | undefined, label: string): number {
+  const parsed = Number(value)
+  if (!Number.isSafeInteger(parsed) || parsed <= 0) throw new Error(`${label} must be a positive integer`)
+  return parsed
+}
+
+assertBuiltBenchmarkRuntime(import.meta.url, {
+  '@deepseek-ai/dsh-client-store': import.meta.resolve('@deepseek-ai/dsh-client-store'),
+  '@deepseek-ai/dsh-llm/assistant-stream': import.meta.resolve('@deepseek-ai/dsh-llm/assistant-stream'),
+  '@deepseek-ai/dsh-session/surface': import.meta.resolve('@deepseek-ai/dsh-session/surface'),
+  '@deepseek-ai/dsh-token-meter/client': import.meta.resolve('@deepseek-ai/dsh-token-meter/client'),
+})
+const [turnsValue, smallDeltasValue, largeDeltasValue, attemptsValue] = process.argv.slice(2)
+const turns = positiveInteger(turnsValue, 'turns')
+const smallDeltas = positiveInteger(smallDeltasValue, 'small deltas')
+const largeDeltas = positiveInteger(largeDeltasValue, 'large deltas')
+const attempts = positiveInteger(attemptsValue, 'attempts')
+const small = synthesizeWindow(turns, smallDeltas)
+const large = synthesizeWindow(turns, largeDeltas)
+if (large.entries.length !== small.entries.length || large.records !== small.records) {
+  throw new Error('conversation-fold workloads must have matching event and compact-record counts')
+}
+const smallFold = bestOf(small.entries, attempts)
+const largeFold = bestOf(large.entries, attempts)
+const report: ConversationFoldWorkerReport = {
+  events: large.entries.length,
+  compactRecords: large.records,
+  streamedDeltas: turns * (largeDeltas + Math.floor(largeDeltas / 4)),
+  chatNodes: largeFold.nodes,
+  smallFoldMs: Math.round(smallFold.ms * 10) / 10,
+  largeFoldMs: Math.round(largeFold.ms * 10) / 10,
+  scaling: Math.round((largeFold.ms / Math.max(smallFold.ms, 1)) * 100) / 100,
+}
+process.stdout.write(`${JSON.stringify(report)}\n`)

+ 24 - 50
benchmarks/session-open/session-open.bench.ts

@@ -1,15 +1,19 @@
 /** Required performance budgets for cold Session preparation, first history, and Agent resume. */
 
-import { spawn } from 'node:child_process'
 import { copyFile, mkdir, mkdtemp, rm } from 'node:fs/promises'
 import { tmpdir } from 'node:os'
 import { join } from 'node:path'
 import { afterAll, beforeAll, describe, expect, it } from 'vitest'
+import {
+  runBuiltBenchmarkWorker,
+  type BuiltBenchmarkWorkerRun,
+} from '../support/built-worker.ts'
 import type {
   SessionOpenBenchmarkScenario,
   SessionOpenWorkerReport,
-} from './session-open.bench.worker.ts'
+} from './session-open.worker.ts'
 import {
+  SYNTHETIC_CURRENT_GENERATION,
   SYNTHETIC_SESSION_DIRECTORY,
   SYNTHETIC_CURRENT_FILENAME,
   SYNTHETIC_V0_FILENAME,
@@ -21,8 +25,8 @@ import {
 const SHAPE = { turns: 200, textDeltas: 500 } as const
 /** Fresh processes per normal-heap scenario; the median enforces each timing budget. */
 const ATTEMPTS = 5
-/** A stuck child is a benchmark failure and must be reaped before another sample starts. */
-const WORKER_TIMEOUT_MS = 120_000
+/** A stuck child is reaped well before the outer test and hook deadlines. */
+const WORKER_TIMEOUT_MS = 60_000
 /** Old-space pressure check, kept independent from normal-heap timing samples. */
 const CONSTRAINED_HEAP_MB = 128
 
@@ -31,7 +35,7 @@ type SessionBenchmarkEndpoint = 'phases' | 'first-history' | 'agent-resume'
 
 const SOURCE_GENERATION_BY_ACCESS = {
   'first-open': 'released-v0',
-  'post-upgrade-reopen': 'current-v2',
+  'post-upgrade-reopen': SYNTHETIC_CURRENT_GENERATION,
 } as const satisfies Record<SessionAccessKind, string>
 
 /** Existing CI calibration: optimized migration is about 2 s and the repeated-snapshot path exceeds 4 s. */
@@ -40,30 +44,24 @@ const MIGRATION_OPEN_BUDGET_MS = 4_000
 const REOPEN_OPEN_BUDGET_MS = 500
 /** Complete event reads remain bounded after either opening path. */
 const READ_BUDGET_MS = 500
-/** Restoring the detached in-memory Session must remain below the migration budget's spare second. */
+/** In-memory Session restore normally takes tens of milliseconds; one second rejects large cloning regressions. */
 const SESSION_RESTORE_BUDGET_MS = 1_000
 /** The fixed production projection set must fold the complete Session within one second. */
 const PROJECTION_BUDGET_MS = 1_000
-/** Host first-history includes migration, restore, projection, and bounded page construction. */
+/** Coarse Host orchestration ceiling; component and constrained-heap gates reject the known migration regression. */
 const FIRST_OPEN_FIRST_HISTORY_BUDGET_MS = 6_000
-/** An already-published V2 Session should produce first history without migration-scale work. */
+/** An already-published current Session should produce first history without migration-scale work. */
 const REOPEN_FIRST_HISTORY_BUDGET_MS = 500
-/** Cold Agent resume includes migration, Session restore, Agent setup, publication, and loop startup. */
+/** Coarse Agent orchestration ceiling; component and constrained-heap gates reject the known migration regression. */
 const FIRST_OPEN_AGENT_RESUME_BUDGET_MS = 7_000
-/** An already-published V2 Session should resume without migration-scale work. */
+/** An already-published current Session should resume without migration-scale work. */
 const REOPEN_AGENT_RESUME_BUDGET_MS = 500
-/** Live Agent, Session, events, and normal caches retained after full GC. */
-const AGENT_RETAINED_HEAP_BUDGET_MB = 192
+/** The 64 MB cap exceeds the 26.1 MB historical-reference median while still detecting retained-graph growth. */
+const AGENT_RETAINED_HEAP_BUDGET_MB = 64
 
-const WORKER = join(import.meta.dirname, 'session-open.bench.worker.ts')
+const WORKER = join(import.meta.dirname, '..', '..', '.dsh-build', 'benchmarks', 'session-open', 'session-open.worker.js')
 
-interface WorkerRun {
-  readonly report: SessionOpenWorkerReport | undefined
-  readonly exitCode: number | null
-  readonly signal: NodeJS.Signals | null
-  readonly timedOut: boolean
-  readonly stderr: string
-}
+type WorkerRun = BuiltBenchmarkWorkerRun<SessionOpenWorkerReport>
 
 function rounded(value: number): number {
   return Math.round(value * 10) / 10
@@ -125,36 +123,12 @@ function runWorker(
   scenario: SessionOpenBenchmarkScenario,
   heapLimitMb?: number,
 ): Promise<WorkerRun> {
-  return new Promise((resolve, reject) => {
-    const child = spawn(process.execPath, [
-      '--expose-gc',
-      ...heapLimitMb === undefined ? [] : [`--max-old-space-size=${String(heapLimitMb)}`],
-      '--import',
-      'tsx/esm',
-      WORKER,
-      root,
-      scenario,
-    ], { cwd: process.cwd(), stdio: ['ignore', 'pipe', 'pipe'] })
-    let stdout = ''
-    let stderr = ''
-    let timedOut = false
-    const timeout = setTimeout(() => {
-      timedOut = true
-      child.kill('SIGKILL')
-    }, WORKER_TIMEOUT_MS)
-    child.stdout.setEncoding('utf8').on('data', (chunk: string) => { stdout += chunk })
-    child.stderr.setEncoding('utf8').on('data', (chunk: string) => { stderr += chunk })
-    child.once('error', (error) => {
-      clearTimeout(timeout)
-      reject(error)
-    })
-    child.once('close', (exitCode, signal) => {
-      clearTimeout(timeout)
-      const line = stdout.trim().split('\n').findLast(candidate => candidate.startsWith('{'))
-      let report: SessionOpenWorkerReport | undefined
-      if (exitCode === 0 && line !== undefined) report = JSON.parse(line) as SessionOpenWorkerReport
-      resolve({ report, exitCode, signal, timedOut, stderr })
-    })
+  return runBuiltBenchmarkWorker({
+    worker: WORKER,
+    args: [root, scenario],
+    timeoutMs: WORKER_TIMEOUT_MS,
+    exposeGc: true,
+    ...(heapLimitMb === undefined ? {} : { heapLimitMb }),
   })
 }
 

+ 12 - 0
benchmarks/session-open/session-open.constants.ts

@@ -0,0 +1,12 @@
+/** Stable identity and released storage location for the Session-opening workload. */
+
+import { join } from 'node:path'
+
+/** Session id encoded in the fixture and used for every measured open. */
+export const SYNTHETIC_SESSION_ID = 'bench-session'
+/** Stable logical working directory encoded in the fixture header. */
+export const SYNTHETIC_SESSION_CWD = '/bench'
+/** Storage directory derived from the fixture's cwd and Session id. */
+export const SYNTHETIC_SESSION_DIRECTORY = join('--bench--', SYNTHETIC_SESSION_ID)
+/** Canonical released-v0 Zstandard generation filename. */
+export const SYNTHETIC_V0_FILENAME = 'session.jsonl.zstd'

+ 21 - 7
benchmarks/session-open/session-open.bench.worker.ts → benchmarks/session-open/session-open.worker.ts

@@ -28,9 +28,11 @@ import * as SessionStatsPlugin from '@deepseek-ai/dsh-session-stats'
 import SessionTitleService from '@deepseek-ai/dsh-session-title'
 import * as SessionTurnOutlinePlugin from '@deepseek-ai/dsh-session-turn-outline'
 import TokenMeter from '@deepseek-ai/dsh-token-meter'
+// These Host-only adapters have no public Node export and are compiled into the benchmark worker.
 import { SessionHistoryController } from '../../packages/api/session-controller/src/history.ts'
 import { installModelSelectionProjection } from '../../packages/api/session-controller/src/model-selection-projection.ts'
-import { SYNTHETIC_SESSION_ID } from './synthetic-released-v0-session.ts'
+import { assertBuiltBenchmarkRuntime } from '../support/built-worker.ts'
+import { SYNTHETIC_SESSION_ID } from './session-open.constants.ts'
 
 /** Worker scenario selected by the parent benchmark. */
 export type SessionOpenBenchmarkScenario =
@@ -127,6 +129,7 @@ function memoryDelta(
 }
 
 async function installProjectionSet(ctx: Context, agentLoopOwnsBoundary: boolean): Promise<void> {
+  // Mirrors projection owners mounted by the base and web-app bundles without timing profile boot.
   if (!agentLoopOwnsBoundary) ctx.sessionProjections.register(turnBoundaryProjectionDefinition)
   ctx.sessionProjections.register(agentPresetProjectionDefinition)
   installModelSelectionProjection(ctx)
@@ -151,6 +154,7 @@ class SessionBenchmarkHost {
   private constructor(
     private readonly ctx: Context,
     private readonly scenario: SessionOpenBenchmarkScenario,
+    private readonly history: SessionHistoryController | undefined,
   ) {}
 
   static async create(root: string, scenario: SessionOpenBenchmarkScenario): Promise<SessionBenchmarkHost> {
@@ -161,9 +165,16 @@ class SessionBenchmarkHost {
     else await ctx.plugin(SessionStore)
     await installProjectionSet(ctx, agentScenario)
     await ctx.plugin(JsonlSessionPersistence, { root, compression: 'zstd' })
-    if (scenario === 'first-history') new BenchmarkSessionQuery(ctx)
+    let history: SessionHistoryController | undefined
+    if (scenario === 'first-history') {
+      new BenchmarkSessionQuery(ctx)
+      history = new SessionHistoryController(ctx, (observation) => {
+        // First-history ends at snapshot delivery; Agent-resume owns live activation and retention.
+        observation[Symbol.dispose]()
+      })
+    }
     if (agentScenario) await ctx.plugin(AgentLoop, { agents: [] })
-    return new SessionBenchmarkHost(ctx, scenario)
+    return new SessionBenchmarkHost(ctx, scenario, history)
   }
 
   async measure(): Promise<SessionOpenWorkerReport> {
@@ -252,9 +263,8 @@ class SessionBenchmarkHost {
   private async measureFirstHistory(): Promise<{ readonly events: number }> {
     const abort = new AbortController()
     this.historyAbort = abort
-    const history = new SessionHistoryController(this.ctx, (observation) => {
-      observation[Symbol.dispose]()
-    })
+    const history = this.history
+    if (history === undefined) throw new Error('first-history benchmark did not initialize its controller')
     const iterator = history.follow({
       address: { kind: 'session', sessionId: SessionId(SYNTHETIC_SESSION_ID) },
     }, abort.signal)[Symbol.asyncIterator]()
@@ -278,6 +288,10 @@ class SessionBenchmarkHost {
   }
 }
 
+assertBuiltBenchmarkRuntime(import.meta.url, {
+  '@deepseek-ai/dsh-session-persistence-jsonl': import.meta.resolve('@deepseek-ai/dsh-session-persistence-jsonl'),
+})
+
 const [root, scenarioValue] = process.argv.slice(2)
 const scenarios: readonly SessionOpenBenchmarkScenario[] = [
   'phase-migrate',
@@ -286,7 +300,7 @@ const scenarios: readonly SessionOpenBenchmarkScenario[] = [
   'agent-resume',
 ]
 if (root === undefined || !scenarios.includes(scenarioValue as SessionOpenBenchmarkScenario)) {
-  throw new Error('usage: session-open.bench.worker.ts <root> <phase-migrate|phase-steady|first-history|agent-resume>')
+  throw new Error('usage: session-open.worker.js <root> <phase-migrate|phase-steady|first-history|agent-resume>')
 }
 const scenario = scenarioValue as SessionOpenBenchmarkScenario
 const host = await SessionBenchmarkHost.create(root, scenario)

+ 19 - 6
benchmarks/session-open/synthetic-released-v0-session.ts

@@ -2,7 +2,22 @@
 
 import { mkdir, writeFile } from 'node:fs/promises'
 import { join } from 'node:path'
+import { SESSION_FORMAT_VERSION } from '@deepseek-ai/dsh-session'
+import { generationLogFilename } from '../../packages/session/session-persistence-jsonl/src/format.ts'
 import { compressZstdFrame } from '../../packages/session/session-persistence-jsonl/src/zstd.ts'
+import {
+  SYNTHETIC_SESSION_CWD,
+  SYNTHETIC_SESSION_DIRECTORY,
+  SYNTHETIC_SESSION_ID,
+  SYNTHETIC_V0_FILENAME,
+} from './session-open.constants.ts'
+
+export {
+  SYNTHETIC_SESSION_CWD,
+  SYNTHETIC_SESSION_DIRECTORY,
+  SYNTHETIC_SESSION_ID,
+  SYNTHETIC_V0_FILENAME,
+} from './session-open.constants.ts'
 
 /** Fixed workload parameters used by every Session-opening scenario. */
 export interface SyntheticV0SessionShape {
@@ -12,12 +27,10 @@ export interface SyntheticV0SessionShape {
   readonly textDeltas: number
 }
 
-/** Stable identity and storage location of the synthesized Session. */
-export const SYNTHETIC_SESSION_ID = 'bench-session'
-export const SYNTHETIC_SESSION_CWD = '/bench'
-export const SYNTHETIC_SESSION_DIRECTORY = join('--bench--', SYNTHETIC_SESSION_ID)
-export const SYNTHETIC_V0_FILENAME = 'session.jsonl.zstd'
-export const SYNTHETIC_CURRENT_FILENAME = 'session.v2.jsonl.zstd'
+/** Canonical current-generation filename produced by the runtime under test. */
+export const SYNTHETIC_CURRENT_FILENAME = generationLogFilename(SESSION_FORMAT_VERSION, 'zstd')
+/** Report label for a fresh-process reopen of the runtime's current generation. */
+export const SYNTHETIC_CURRENT_GENERATION = `current-v${String(SESSION_FORMAT_VERSION)}`
 
 const TIME_ZERO = 1_700_000_000_000
 /** One body frame per row preserves the historical many-frame workload deterministically. */

+ 97 - 0
benchmarks/support/built-worker.ts

@@ -0,0 +1,97 @@
+/** Plain-Node launcher for compiled benchmark workers. */
+
+import { spawn } from 'node:child_process'
+
+/** Process outcome and optional JSON report from one compiled benchmark worker. */
+export interface BuiltBenchmarkWorkerRun<Report> {
+  readonly report: Report | undefined
+  readonly exitCode: number | null
+  readonly signal: NodeJS.Signals | null
+  readonly timedOut: boolean
+  readonly stderr: string
+}
+
+/** Options for one isolated compiled benchmark process. */
+export interface BuiltBenchmarkWorkerOptions {
+  readonly worker: string
+  readonly args?: readonly string[]
+  readonly timeoutMs: number
+  readonly exposeGc?: boolean
+  readonly heapLimitMb?: number
+}
+
+/**
+ * Run one built JavaScript worker without a TypeScript runtime loader.
+ * @param options - worker path, arguments, deadline, and optional V8 limits.
+ * @returns child exit details and its final JSON-line report when successful.
+ */
+export function runBuiltBenchmarkWorker<Report>(
+  options: BuiltBenchmarkWorkerOptions,
+): Promise<BuiltBenchmarkWorkerRun<Report>> {
+  if (!options.worker.endsWith('.js') && !options.worker.endsWith('.cjs')) {
+    throw new Error(`benchmark worker must be compiled JavaScript: ${options.worker}`)
+  }
+  const env = { ...process.env }
+  delete env['NODE_OPTIONS']
+  delete env['TSX_TSCONFIG_PATH']
+  return new Promise((resolve, reject) => {
+    const child = spawn(process.execPath, [
+      ...options.exposeGc === true ? ['--expose-gc'] : [],
+      ...options.heapLimitMb === undefined
+        ? []
+        : [`--max-old-space-size=${String(options.heapLimitMb)}`],
+      options.worker,
+      ...options.args ?? [],
+    ], {
+      cwd: process.cwd(),
+      env,
+      stdio: ['ignore', 'pipe', 'pipe'],
+    })
+    let stdout = ''
+    let stderr = ''
+    let timedOut = false
+    const timeout = setTimeout(() => {
+      timedOut = true
+      child.kill('SIGKILL')
+    }, options.timeoutMs)
+    child.stdout.setEncoding('utf8').on('data', (chunk: string) => { stdout += chunk })
+    child.stderr.setEncoding('utf8').on('data', (chunk: string) => { stderr += chunk })
+    child.once('error', (error) => {
+      clearTimeout(timeout)
+      reject(error)
+    })
+    child.once('close', (exitCode, signal) => {
+      clearTimeout(timeout)
+      const line = stdout.trim().split('\n').findLast(candidate => candidate.startsWith('{'))
+      try {
+        const report = exitCode === 0 && line !== undefined
+          ? JSON.parse(line) as Report
+          : undefined
+        resolve({ report, exitCode, signal, timedOut, stderr })
+      } catch (error: unknown) {
+        reject(error)
+      }
+    })
+  })
+}
+
+/**
+ * Reject a benchmark worker reached through source execution or a TypeScript loader.
+ * @param moduleUrl - `import.meta.url` from the worker entry.
+ * @param packageEntries - resolved production package entries used by the measured path.
+ */
+export function assertBuiltBenchmarkRuntime(
+  moduleUrl: string,
+  packageEntries: Readonly<Record<string, string>>,
+): void {
+  if (!moduleUrl.endsWith('.js') || !moduleUrl.includes('/.dsh-build/benchmarks/')) {
+    throw new Error(`benchmark worker is not running from .dsh-build/benchmarks: ${moduleUrl}`)
+  }
+  const tsRuntime = process.execArgv.find(argument => /(?:^|[/\\])tsx(?:[/\\]|$)|tsx\/esm|tsx\/cjs/.test(argument))
+  if (tsRuntime !== undefined) throw new Error(`benchmark worker received a TypeScript loader: ${tsRuntime}`)
+  for (const [specifier, entry] of Object.entries(packageEntries)) {
+    if (!/\/lib\/(?:[^/]+\/)*[^/]+\.js$/.test(entry)) {
+      throw new Error(`benchmark package ${specifier} did not resolve to lib JavaScript: ${entry}`)
+    }
+  }
+}

+ 33 - 0
benchmarks/tsdown.config.ts

@@ -0,0 +1,33 @@
+import { defineConfig } from 'tsdown'
+
+const shared = {
+  format: 'esm' as const,
+  platform: 'node' as const,
+  target: 'es2024',
+  fixedExtension: false,
+  dts: false,
+  deps: {
+    neverBundle: [/^@deepseek-ai\//],
+    onlyBundle: false as const,
+  },
+}
+
+/** Compile measured benchmark workers while keeping workspace packages on their built `lib` entries. */
+export default defineConfig([
+  {
+    ...shared,
+    entry: { 'session-open.worker': 'session-open/session-open.worker.ts' },
+    outDir: '../.dsh-build/benchmarks/session-open',
+    clean: true,
+    tsconfig: 'tsconfig.host.json',
+  },
+  {
+    ...shared,
+    entry: {
+      'conversation-fold.worker': 'conversation-fold/conversation-fold.worker.client.ts',
+    },
+    outDir: '../.dsh-build/benchmarks/conversation-fold',
+    clean: true,
+    tsconfig: 'tsconfig.client.json',
+  },
+])

+ 2 - 2
docs/testing.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/testing.md
-testing.md: 06abd29971bcf6918373e8d09681490cf743d7cd
-testing.zh.md: 1aa1270f29562c08c73cee141dd6b98c953f752b
+testing.md: e061ba5e801da0cc68097338ef2b40d76fda6077
+testing.zh.md: c1923235ec78a792f03b5429e6d54c7d8cfbee56

+ 1 - 1
docs/testing.md

@@ -10,7 +10,7 @@ How this repo tests, tier by tier, and the rules that keep a green suite meaning
 - **Coverage gate** (`pnpm run test:coverage`): the gating run, per-file 100% on `packages/*/*/src`. An uncovered line is often dead code the gate flags for deletion, not a missing test to bolt on. Line coverage is necessary, never sufficient — it proves lines ran, not that the feature works as shipped. Per-file 100% on `packages/shell/pwsh-local/src` needs a real `pwsh`: without one its executor suites self-skip and `vitest.config.ts` exempts the file so pwsh-less hosts stay green, while CI runners ship pwsh and enforce the full bar.
 - **Real-API e2e** (`pnpm run test:e2e`): with-key tests against live provider APIs — the DeepSeek model plus provider-specific smokes that gate on their own keys (`EXA_API_KEY`, `PERPLEXITY_API_KEY`, …); each suite self-skips without its key so keyless CI stays green ([real-API e2e Agent Note](../.agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.md)).
 - **Owner-local expected output** (`pnpm run test:expected`): keyless assembled CLI/process expectations without a recorded-session round trip. Drivers use `*.expected.e2e.ts` beside `tests/expected/`; CI runs built exports. Package/script expectations use `test`, while browser expectations use `test:web`.
-- **Performance benchmarks** (`pnpm run test:bench`; required Linux PR gate `node 24 / benchmarks`): top-level `benchmarks/` holds `*.bench.ts` and Client-face `*.bench.client.ts` gates grouped by user path. They synthesize fixed inputs, never recordings, and enforce documented wall-clock, heap, or scaling budgets; package-local `.perf.ts` remains diagnostic ([rules](../.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md)).
+- **Performance benchmarks** (`pnpm run test:bench`; required Linux PR gate `node 24 / benchmarks`): `benchmarks/` groups user-path gates. It builds libraries and workers; timed code runs under plain Node, never TSX. Synthetic inputs enforce time, heap, and scaling budgets; package-local `.perf.ts` stays diagnostic ([rules](../.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md)).
 - **Snapshot** (`pnpm run test:snapshot`): a top-level scenario's highest recorded parent generation supplies user input and model replay, then serves as the expected persisted result. Parent filenames are `session[.vN].jsonl`; child roles are `session.<ordinal>[.vN].jsonl`; v0 omits `.v0`, positive versions require lowercase `.vN`, and each filename must agree with its header. Process scenarios start through `dsh`: headless owns one-shot behavior, the SDK owns persistent control, ACP owns automation-protocol behavior, and Web retains browser/ARIA evidence beside the same Session. `snapshot.yml` declares the profile, composition/header class, recording policy, exceptional replay or input metadata, and workspace facts. Typed tokens preserve parent/child identity relationships; only header pins own prompt/schema sidecars. A mutating scenario independently compares the complete `workspace.expected/` tree, which record and refresh never rewrite. Use `test:snapshot:record` when a model transcript changes and `test:snapshot:refresh` when replay input remains valid; review every resulting diff.
 - **Web browser snapshot** (`pnpm run test:web`; required Linux PR gate): Chromium compares session-driven output under `snapshots/web/` and UI-only output under `apps/web/tests/expected/`. CI forces read-only `DSH_SNAPSHOT=replay`, never writing expected outputs; record/refresh stay local and every diff is reviewed ([web e2e lane](../.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.md), [CI gate decision](../.agents/notes/implemented/testing/2026-07-30-web-browser-snapshot-ci-gate.md)). `test:web` builds first for plugin CSS.
 

+ 1 - 1
docs/testing.zh.md

@@ -10,7 +10,7 @@
 - **覆盖率门禁**(`pnpm run test:coverage`):门禁级运行,对 `packages/*/*/src` 按文件 100% 覆盖。未覆盖的行往往是门禁正确标记出的死代码(应删除),而非需要补写的测试。行覆盖率是必要条件,但永远不是充分条件:它证明行被执行过,不证明功能按交付预期工作。`packages/shell/pwsh-local/src` 的按文件 100% 覆盖需要真实的 `pwsh`:缺少它时其执行器套件会自动跳过,`vitest.config.ts` 会豁免该文件以使无 pwsh 的主机保持绿色,而 CI runner 自带 pwsh,仍按完整标准执行门禁。
 - **真实 API e2e**(`pnpm run test:e2e`):带密钥测试调用真实提供方 API,包括 DeepSeek 模型以及各提供方特有的冒烟测试;这些测试各自由自己的密钥控制(`EXA_API_KEY`、`PERPLEXITY_API_KEY` 等),缺少密钥时套件会自动跳过,使 keyless CI 保持绿色([真实 API e2e Agent Note](../.agents/notes/implemented/testing/2026-06-19-real-api-e2e-ci.zh.md))。
 - **所属位置的预期输出**(`pnpm run test:expected`):无录制会话往返的无密钥组装 CLI/进程预期。驱动使用 `*.expected.e2e.ts`,并与 `tests/expected/` 同属一处;CI 针对构建产物运行。包/脚本预期使用 `test`,浏览器预期使用 `test:web`。
-- **性能基准**(`pnpm run test:bench`;必需的 Linux PR gate `node 24 / benchmarks`):顶层 `benchmarks/` 按用户路径组织 `*.bench.ts` 与 Client 面 `*.bench.client.ts` gate。它们使用固定合成输入而非录制材料,并执行有记录的壁钟、堆或缩放预算;包内 `.perf.ts` 仍是诊断([规则](../.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md))。
+- **性能基准**(`pnpm run test:bench`;必需的 Linux PR gate `node 24 / benchmarks`):`benchmarks/` 按用户路径组织门禁。它先构建 library 和 worker;被计时代码在纯 Node 下运行,不使用 TSX。合成输入执行耗时、堆和缩放预算;包内 `.perf.ts` 保留为诊断([规则](../.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md))。
 - **快照**(`pnpm run test:snapshot`):顶层场景数值最高的已录制 parent generation 同时提供用户输入和模型回放,并作为持久化结果的预期值。parent 文件名是 `session[.vN].jsonl`;child 角色使用 `session.<ordinal>[.vN].jsonl`;v0 省略 `.v0`,正版本必须使用小写 `.vN`,且每个文件名必须与其 header 一致。进程级场景都通过 `dsh` 启动:headless 负责一次性行为,SDK 负责持久控制,ACP 负责自动化协议行为,Web 在同一 Session 旁保留浏览器与 ARIA 证据。`snapshot.yml` 声明 profile、组合与请求头类别、录制策略、例外回放或输入元数据以及 workspace 事实。带类型的 token 保留父子身份关系;只有请求头 pin 拥有 prompt/schema sidecar。变更 workspace 的场景会独立比较完整的 `workspace.expected/` 目录,record 与 refresh 绝不改写该目录。当模型 transcript(文本记录)变化时使用 `test:snapshot:record`,回放输入仍有效时使用 `test:snapshot:refresh`;请审查所有结果差异。
 - **Web 浏览器快照**(`pnpm run test:web`;必需的 Linux PR(Pull Request)门禁):Chromium 比较 `snapshots/web/` 下由会话驱动的输出,以及 `apps/web/tests/expected/` 下仅含 UI 的输出。CI 强制只读的 `DSH_SNAPSHOT=replay`,绝不写入预期输出;record/refresh 留在本地,每处 diff 都须评审([web e2e 车道](../.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.zh.md)、[CI 门禁决策](../.agents/notes/implemented/testing/2026-07-30-web-browser-snapshot-ci-gate.zh.md))。`test:web` 会先构建以交付插件 CSS。
 

+ 20 - 1
package.json

@@ -18,6 +18,7 @@
   ],
   "scripts": {
     "build": "tsx scripts/build.ts",
+    "build:bench": "npm run build:lib && tsdown --config benchmarks/tsdown.config.ts",
     "build:official": "tsx scripts/build.ts --profile official",
     "build:lib": "npm run build:lib:host && npm run build:lib:client",
     "build:lib:host": "node --max-old-space-size=4096 ./node_modules/typescript/bin/tsc -b tsconfig.host.json && tsdown --env.DSH_BUILD_FACE host",
@@ -36,7 +37,8 @@
     "test:coverage": "vitest run --coverage",
     "test:coverage:partitioned": "tsx scripts/run-coverage-partitions.ts",
     "test:e2e": "vitest run --config vitest.e2e.config.ts",
-    "test:bench": "vitest run --config vitest.bench.config.ts",
+    "test:bench": "npm run build:bench && npm run test:bench:built",
+    "test:bench:built": "vitest run --config vitest.bench.config.ts",
     "test:expected": "vitest run --config vitest.expected.config.ts",
     "test:expected:refresh": "DSH_SNAPSHOT=refresh vitest run --config vitest.expected.config.ts",
     "test:issue-management": "node .github/issue-management/policy.test.mjs",
@@ -164,7 +166,24 @@
   },
   "devDependencies": {
     "@deepseek-ai/dsh-package-manifest": "workspace:^",
+    "@deepseek-ai/cordis": "workspace:^",
+    "@deepseek-ai/dsh-agent-loop": "workspace:^",
+    "@deepseek-ai/dsh-agent-loop-testkit": "workspace:^",
+    "@deepseek-ai/dsh-agent-presets": "workspace:^",
+    "@deepseek-ai/dsh-client-store": "workspace:^",
+    "@deepseek-ai/dsh-deque": "workspace:^",
+    "@deepseek-ai/dsh-llm": "workspace:^",
+    "@deepseek-ai/dsh-session": "workspace:^",
+    "@deepseek-ai/dsh-session-persistence": "workspace:^",
+    "@deepseek-ai/dsh-session-persistence-jsonl": "workspace:^",
+    "@deepseek-ai/dsh-session-projection": "workspace:^",
+    "@deepseek-ai/dsh-session-query": "workspace:^",
+    "@deepseek-ai/dsh-session-stats": "workspace:^",
+    "@deepseek-ai/dsh-session-title": "workspace:^",
+    "@deepseek-ai/dsh-session-turn-outline": "workspace:^",
+    "@deepseek-ai/dsh-token-meter": "workspace:^",
     "@deepseek-ai/dsh-tool-session-query": "workspace:^",
+    "@deepseek-ai/dsh-typert-protocol": "workspace:^",
     "@deepseek-ai/dsh-web-fetch-http": "workspace:^",
     "@stylistic/eslint-plugin": "^5.10.0",
     "@testing-library/dom": "^10.4.1",

+ 51 - 0
pnpm-lock.yaml

@@ -16,12 +16,63 @@ importers:
 
   .:
     devDependencies:
+      '@deepseek-ai/cordis':
+        specifier: workspace:^
+        version: link:vendor/cordis
+      '@deepseek-ai/dsh-agent-loop':
+        specifier: workspace:^
+        version: link:packages/core/agent-loop
+      '@deepseek-ai/dsh-agent-loop-testkit':
+        specifier: workspace:^
+        version: link:packages/test-support/agent-loop-testkit
+      '@deepseek-ai/dsh-agent-presets':
+        specifier: workspace:^
+        version: link:packages/preset/agent-presets
+      '@deepseek-ai/dsh-client-store':
+        specifier: workspace:^
+        version: link:packages/client/store
+      '@deepseek-ai/dsh-deque':
+        specifier: workspace:^
+        version: link:packages/util/deque
+      '@deepseek-ai/dsh-llm':
+        specifier: workspace:^
+        version: link:packages/llm/llm
       '@deepseek-ai/dsh-package-manifest':
         specifier: workspace:^
         version: link:packages/util/package-manifest
+      '@deepseek-ai/dsh-session':
+        specifier: workspace:^
+        version: link:packages/core/session
+      '@deepseek-ai/dsh-session-persistence':
+        specifier: workspace:^
+        version: link:packages/session/session-persistence
+      '@deepseek-ai/dsh-session-persistence-jsonl':
+        specifier: workspace:^
+        version: link:packages/session/session-persistence-jsonl
+      '@deepseek-ai/dsh-session-projection':
+        specifier: workspace:^
+        version: link:packages/session/session-projection
+      '@deepseek-ai/dsh-session-query':
+        specifier: workspace:^
+        version: link:packages/session-query/session-query
+      '@deepseek-ai/dsh-session-stats':
+        specifier: workspace:^
+        version: link:packages/session/session-stats
+      '@deepseek-ai/dsh-session-title':
+        specifier: workspace:^
+        version: link:packages/session/session-title
+      '@deepseek-ai/dsh-session-turn-outline':
+        specifier: workspace:^
+        version: link:packages/session/session-turn-outline
+      '@deepseek-ai/dsh-token-meter':
+        specifier: workspace:^
+        version: link:packages/llm/token-meter
       '@deepseek-ai/dsh-tool-session-query':
         specifier: workspace:^
         version: link:packages/session-query/tool-session-query
+      '@deepseek-ai/dsh-typert-protocol':
+        specifier: workspace:^
+        version: link:packages/typert/protocol
       '@deepseek-ai/dsh-web-fetch-http':
         specifier: workspace:^
         version: link:packages/web/web-fetch-http

+ 1 - 0
tsconfig.client.json

@@ -21,6 +21,7 @@
     "packages/client/*/tests/**/*.tsx",
     "benchmarks/**/*.client.ts",
     "benchmarks/**/*.client.tsx",
+    "benchmarks/support/**/*.ts",
     "packages/*/*/tests/**/*.client.spec.ts",
     "packages/*/*/tests/**/*.client.spec.tsx",
     "packages/*/*/tests/**/*.client.tsx",

+ 4 - 4
vitest.bench.config.ts

@@ -3,10 +3,10 @@ import { defineConfig } from 'vitest/config'
 import { standardDecoratorPlugin, vitestExecArgv } from './vitest.shared.ts'
 
 /**
- * CI performance gate. Every benchmark under `benchmarks/` synthesizes its
- * own input from fixed parameters, measures one user-visible path, and fails when a
- * documented time or heap budget is exceeded. Files run one at a time so a
- * measurement never shares the CPU with another benchmark.
+ * CI performance gate. Vitest orchestrates compiled plain-Node workers under
+ * `.dsh-build/benchmarks/`; timed product work never runs through its source transform.
+ * Files run one at a time so a measurement never shares the CPU with another
+ * benchmark.
  */
 export default defineConfig({
   plugins: [tsconfigPaths({ projects: ['./tsconfig.base.json'] }), standardDecoratorPlugin()],