Ver código fonte

test(perf): isolate benchmark dependencies

imccyu 3 semanas atrás
pai
commit
a548150f86

+ 2 - 2
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
-2026-09-04-session-open-performance-gate.md: 723bdf7825792b3beae46cf6b43c9eb945cf4508
-2026-09-04-session-open-performance-gate.zh.md: 047104be1d50ad53fa6b44804791d14c0cf16c44
+2026-09-04-session-open-performance-gate.md: 30e65eb52dea48426cf87cf91053a9133ab42763
+2026-09-04-session-open-performance-gate.zh.md: 3e202c966ea4764e135c79e1d238af707dca1f8d

+ 1 - 1
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md

@@ -12,7 +12,7 @@ Measuring only `SessionPersistence.open()` does not stably describe the result f
 
 ## Decision
 
-Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench`. The command first builds workspace libraries and dedicated benchmark workers, then invokes `vitest.bench.config.ts`. The job runs the benchmark lane alone; Vitest runs one file at a time and only prepares input, starts measurement children, aggregates results, and enforces budgets. Every timed CPU path executes compiled JavaScript under plain Node with `NODE_OPTIONS` removed and no TypeScript loader; bare workspace imports therefore resolve through package exports to built `lib/` entries.
+Linux pull requests run a required `node 24 / benchmarks` job that executes `pnpm run check:ci:bench` → `pnpm run test:bench`. The private `@deepseek-ai/dsh-benchmarks` workspace owns benchmark-only dependencies. The command first builds workspace libraries and dedicated workers under `benchmarks/.dsh-build/`, then invokes `vitest.bench.config.ts`. The job runs the benchmark lane alone; Vitest runs one file at a time and only prepares input, starts measurement children, aggregates results, and enforces budgets. Every timed CPU path executes compiled JavaScript under plain Node with `NODE_OPTIONS` removed and no TypeScript loader; bare workspace imports therefore resolve from `benchmarks/node_modules` through package exports to built `lib/` entries.
 
 Required performance gates live under top-level `benchmarks/`, grouped by measured user path rather than package ownership. Host files use `*.bench.ts`, Client-face files use `*.bench.client.ts`, and scenario-specific workers and fixtures stay beside their benchmark without a benchmark suffix. Package-local `.perf.ts` files remain non-gating diagnostics; `scripts/` owns orchestration rather than benchmark cases.
 

+ 1 - 1
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md

@@ -12,7 +12,7 @@ Session format v2 的推出改变了两条成本随模型输出增长的路径
 
 ## 决定
 
-Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench`。该命令先构建 workspace library 和专用 benchmark worker,再调用 `vitest.bench.config.ts`。该 job 单独运行 benchmark lane;Vitest 逐文件运行,只负责准备输入、启动测量子进程、汇总结果和执行预算断言。每条被计时的 CPU 路径都以纯 Node 执行编译后的 JavaScript,并移除 `NODE_OPTIONS` 且不加载 TypeScript runtime;workspace 裸导入因此通过 package exports 解析到构建后的 `lib/` 入口。
+Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run check:ci:bench` → `pnpm run test:bench`。私有 `@deepseek-ai/dsh-benchmarks` workspace 拥有 benchmark 专属依赖。该命令先构建 workspace library 和 `benchmarks/.dsh-build/` 下的专用 worker,再调用 `vitest.bench.config.ts`。该 job 单独运行 benchmark lane;Vitest 逐文件运行,只负责准备输入、启动测量子进程、汇总结果和执行预算断言。每条被计时的 CPU 路径都以纯 Node 执行编译后的 JavaScript,并移除 `NODE_OPTIONS` 且不加载 TypeScript runtime;workspace 裸导入因此从 `benchmarks/node_modules` 通过 package exports 解析到构建后的 `lib/` 入口。
 
 必需性能 gate 位于顶层 `benchmarks/`,按被测用户路径而非 package 归属组织。Host 文件使用 `*.bench.ts`,Client 面文件使用 `*.bench.client.ts`,场景专属 worker 与 fixture 留在对应 benchmark 旁且不带 benchmark 后缀。包内 `.perf.ts` 文件仍是非门禁诊断;`scripts/` 负责编排而不承载 benchmark case。
 

+ 1 - 1
benchmarks/AGENTS.md

@@ -4,7 +4,7 @@ This tree owns required, repository-level performance gates whose measured user
 
 - Organize benchmarks by measured user path, one directory per path. Do not mirror the package tree.
 - Host cases use `*.bench.ts`; Client-face cases use `*.bench.client.ts`. Worker, fixture, and support modules do not carry a benchmark suffix.
-- `test:bench` builds workspace libraries and `.dsh-build/benchmarks/` workers before Vitest orchestration. Timed CPU work runs in those workers under plain Node, without a TypeScript loader; runtime package imports must resolve to built `lib/` entries.
+- The private `@deepseek-ai/dsh-benchmarks` workspace owns benchmark-only dependencies. `test:bench` builds workspace libraries and `benchmarks/.dsh-build/` workers before Vitest orchestration. Timed CPU work runs in those workers under plain Node, without a TypeScript loader; runtime package imports must resolve to built `lib/` entries.
 - Synthesize fixed inputs from reviewed constants. Never use recorded Sessions, user material, ambient repositories, or network services.
 - Run process-level wall-clock and retained-memory samples in fresh children with private `mkdtemp` roots. Pure synchronous folds create a fresh object graph per sample and must not mutate process-global state. Bound every child, await exit, and remove owned roots after failure as well as success.
 - Record reference-machine expectations separately from the shared CI time scale and variance headroom. Do not apply the time scale to memory or dimensionless ratios.

+ 0 - 2
benchmarks/conversation-fold/conversation-fold.bench.client.ts

@@ -42,9 +42,7 @@ const MAX_DELTA_SCALING = EXPECTED_DELTA_SCALING * PERFORMANCE_BUDGET_HEADROOM
 const WORKER = join(
   import.meta.dirname,
   '..',
-  '..',
   '.dsh-build',
-  'benchmarks',
   'conversation-fold',
   'conversation-fold.worker.js',
 )

+ 32 - 0
benchmarks/package.json

@@ -0,0 +1,32 @@
+{
+  "name": "@deepseek-ai/dsh-benchmarks",
+  "version": "0.1.3-alpha.1",
+  "license": "MIT",
+  "private": true,
+  "type": "module",
+  "devDependencies": {
+    "@deepseek-ai/cordis": "workspace:^",
+    "@deepseek-ai/dsh-agent": "workspace:^",
+    "@deepseek-ai/dsh-agent-loop": "workspace:^",
+    "@deepseek-ai/dsh-agent-loop-testkit": "workspace:^",
+    "@deepseek-ai/dsh-agent-presets": "workspace:^",
+    "@deepseek-ai/dsh-api-session-controller": "workspace:^",
+    "@deepseek-ai/dsh-client-store": "workspace:^",
+    "@deepseek-ai/dsh-client-ui-chat": "workspace:^",
+    "@deepseek-ai/dsh-deque": "workspace:^",
+    "@deepseek-ai/dsh-llm": "workspace:^",
+    "@deepseek-ai/dsh-session": "workspace:^",
+    "@deepseek-ai/dsh-session-persistence": "workspace:^",
+    "@deepseek-ai/dsh-session-persistence-jsonl": "workspace:^",
+    "@deepseek-ai/dsh-session-projection": "workspace:^",
+    "@deepseek-ai/dsh-session-query": "workspace:^",
+    "@deepseek-ai/dsh-session-stats": "workspace:^",
+    "@deepseek-ai/dsh-session-title": "workspace:^",
+    "@deepseek-ai/dsh-session-turn-outline": "workspace:^",
+    "@deepseek-ai/dsh-token-meter": "workspace:^",
+    "@deepseek-ai/dsh-typert-protocol": "workspace:^"
+  },
+  "peerDependencies": {
+    "@deepseek-ai/cordis": "workspace:^"
+  }
+}

+ 1 - 1
benchmarks/session-open/session-open.bench.ts

@@ -70,7 +70,7 @@ const AGENT_RETAINED_HEAP_BUDGET_MB = Math.ceil(
   EXPECTED_AGENT_RETAINED_HEAP_MB * PERFORMANCE_BUDGET_HEADROOM,
 )
 
-const WORKER = join(import.meta.dirname, '..', '..', '.dsh-build', 'benchmarks', 'session-open', 'session-open.worker.js')
+const WORKER = join(import.meta.dirname, '..', '.dsh-build', 'session-open', 'session-open.worker.js')
 
 type WorkerRun = BuiltBenchmarkWorkerRun<SessionOpenWorkerReport>
 

+ 2 - 2
benchmarks/support/built-worker.ts

@@ -84,8 +84,8 @@ export function assertBuiltBenchmarkRuntime(
   moduleUrl: string,
   packageEntries: Readonly<Record<string, string>>,
 ): void {
-  if (!moduleUrl.endsWith('.js') || !moduleUrl.includes('/.dsh-build/benchmarks/')) {
-    throw new Error(`benchmark worker is not running from .dsh-build/benchmarks: ${moduleUrl}`)
+  if (!moduleUrl.endsWith('.js') || !moduleUrl.includes('/benchmarks/.dsh-build/')) {
+    throw new Error(`benchmark worker is not running from benchmarks/.dsh-build: ${moduleUrl}`)
   }
   const tsRuntime = process.execArgv.find(argument => /(?:^|[/\\])tsx(?:[/\\]|$)|tsx\/esm|tsx\/cjs/.test(argument))
   if (tsRuntime !== undefined) throw new Error(`benchmark worker received a TypeScript loader: ${tsRuntime}`)

+ 2 - 2
benchmarks/tsdown.config.ts

@@ -17,7 +17,7 @@ export default defineConfig([
   {
     ...shared,
     entry: { 'session-open.worker': 'session-open/session-open.worker.ts' },
-    outDir: '../.dsh-build/benchmarks/session-open',
+    outDir: '.dsh-build/session-open',
     clean: true,
     tsconfig: 'tsconfig.host.json',
   },
@@ -26,7 +26,7 @@ export default defineConfig([
     entry: {
       'conversation-fold.worker': 'conversation-fold/conversation-fold.worker.client.ts',
     },
-    outDir: '../.dsh-build/benchmarks/conversation-fold',
+    outDir: '.dsh-build/conversation-fold',
     clean: true,
     tsconfig: 'tsconfig.client.json',
   },

+ 2 - 2
docs/testing.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write docs/testing.md
-testing.md: e061ba5e801da0cc68097338ef2b40d76fda6077
-testing.zh.md: c1923235ec78a792f03b5429e6d54c7d8cfbee56
+testing.md: b227aea937c63e641580486e6481234168231258
+testing.zh.md: f3938fc7773ed2dcd1a37e257310a060d8dafa6e

+ 2 - 2
docs/testing.md

@@ -14,11 +14,11 @@ How this repo tests, tier by tier, and the rules that keep a green suite meaning
 - **Snapshot** (`pnpm run test:snapshot`): a top-level scenario's highest recorded parent generation supplies user input and model replay, then serves as the expected persisted result. Parent filenames are `session[.vN].jsonl`; child roles are `session.<ordinal>[.vN].jsonl`; v0 omits `.v0`, positive versions require lowercase `.vN`, and each filename must agree with its header. Process scenarios start through `dsh`: headless owns one-shot behavior, the SDK owns persistent control, ACP owns automation-protocol behavior, and Web retains browser/ARIA evidence beside the same Session. `snapshot.yml` declares the profile, composition/header class, recording policy, exceptional replay or input metadata, and workspace facts. Typed tokens preserve parent/child identity relationships; only header pins own prompt/schema sidecars. A mutating scenario independently compares the complete `workspace.expected/` tree, which record and refresh never rewrite. Use `test:snapshot:record` when a model transcript changes and `test:snapshot:refresh` when replay input remains valid; review every resulting diff.
 - **Web browser snapshot** (`pnpm run test:web`; required Linux PR gate): Chromium compares session-driven output under `snapshots/web/` and UI-only output under `apps/web/tests/expected/`. CI forces read-only `DSH_SNAPSHOT=replay`, never writing expected outputs; record/refresh stay local and every diff is reviewed ([web e2e lane](../.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.md), [CI gate decision](../.agents/notes/implemented/testing/2026-07-30-web-browser-snapshot-ci-gate.md)). `test:web` builds first for plugin CSS.
 
-Session fixtures retain headers and payloads but omit body sequence/time envelopes; replay synthesizes them and, like record and refresh, selects each role's highest generation. Retained v0 and v1 generations may keep packed rows for migration coverage; [the migrator](../scripts/migrate-packed-session-fixtures.ts) rewrites older layouts.
+Session fixtures retain headers and payloads but omit body sequence/time envelopes; replay synthesizes them. Replay, record, and refresh select each parent/child role's highest generation. Current v2 uses `.v2`, one row per event, and embedded compact Assistant streams; retained v0 (suffixless) and v1 (`.v1`) may keep canonical packed rows for migration coverage. [The migrator](../scripts/migrate-packed-session-fixtures.ts) rewrites older historical layouts.
 
 ## How specs execute
 
-Forked workers run several spec files at once, the coverage gate splits into concurrent partitions beside the other gates in its job, and the self-hosted runners share one host and one volume. Only the process is isolated: ports, predictable paths, external namespaces, and inherited children are not. Own each acquired resource through its teardown; a spec that passes only when run alone is defective, not the runner. [dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) owns the allocation, restoration, synchronization, timeout-budget, platform, and teardown rules; its [flake diagnosis workflow](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md) classifies an existing probabilistic failure.
+Forked workers run several spec files at once, the coverage gate splits into concurrent partitions beside the other gates in its job, and the self-hosted runners share one host and one volume. Only the process is isolated: ports, predictable paths, external namespaces, and inherited children are not. Own each acquired resource through its teardown, and read a spec that passes only when it runs alone as a defect in the spec rather than an unstable runner. [dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) owns the allocation, restoration, synchronization, timeout-budget, platform, and teardown rules; its [flake diagnosis workflow](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md) classifies an existing probabilistic failure.
 
 ## The with-key policy: inference is cheap here
 

+ 2 - 2
docs/testing.zh.md

@@ -14,11 +14,11 @@
 - **快照**(`pnpm run test:snapshot`):顶层场景数值最高的已录制 parent generation 同时提供用户输入和模型回放,并作为持久化结果的预期值。parent 文件名是 `session[.vN].jsonl`;child 角色使用 `session.<ordinal>[.vN].jsonl`;v0 省略 `.v0`,正版本必须使用小写 `.vN`,且每个文件名必须与其 header 一致。进程级场景都通过 `dsh` 启动:headless 负责一次性行为,SDK 负责持久控制,ACP 负责自动化协议行为,Web 在同一 Session 旁保留浏览器与 ARIA 证据。`snapshot.yml` 声明 profile、组合与请求头类别、录制策略、例外回放或输入元数据以及 workspace 事实。带类型的 token 保留父子身份关系;只有请求头 pin 拥有 prompt/schema sidecar。变更 workspace 的场景会独立比较完整的 `workspace.expected/` 目录,record 与 refresh 绝不改写该目录。当模型 transcript(文本记录)变化时使用 `test:snapshot:record`,回放输入仍有效时使用 `test:snapshot:refresh`;请审查所有结果差异。
 - **Web 浏览器快照**(`pnpm run test:web`;必需的 Linux PR(Pull Request)门禁):Chromium 比较 `snapshots/web/` 下由会话驱动的输出,以及 `apps/web/tests/expected/` 下仅含 UI 的输出。CI 强制只读的 `DSH_SNAPSHOT=replay`,绝不写入预期输出;record/refresh 留在本地,每处 diff 都须评审([web e2e 车道](../.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.zh.md)、[CI 门禁决策](../.agents/notes/implemented/testing/2026-07-30-web-browser-snapshot-ci-gate.zh.md))。`test:web` 会先构建以交付插件 CSS。
 
-Session fixture 保留 header 与 payload,但省略正文 seq/time envelope;replay 会合成这些 envelope,并与 record、refresh 一样选择每个角色的最高 generation。保留的 v0 与 v1 generation 可以为迁移覆盖保留 packed row;[迁移器](../scripts/migrate-packed-session-fixtures.ts)会改写更旧的布局。
+Session fixture 保留 header 与 payload,但省略正文 seq/time envelope;replay 会合成这些 envelope。Replay、record 与 refresh 会选择每个 parent/child 角色的最高 generation。当前 v2 使用 `.v2`、每个事件一行,并嵌入紧凑 Assistant stream;保留的 v0(无后缀)与 v1(`.v1`)可以为迁移覆盖保留规范 packed row。[迁移器](../scripts/migrate-packed-session-fixtures.ts)会改写更旧的历史布局。
 
 ## spec 如何被执行
 
-fork 出的 worker 会同时运行多个 spec 文件,coverage gate 会拆成并发的 partition,与同一个 job 中的其它 gate 并排运行,而自托管 runner 共用同一台宿主机和同一个卷。被隔离的只有进程:端口、可预测路径、外部命名空间和继承而来的子进程都不隔离。为每个占用的资源负责到它的 teardown;只有单独运行时才通过的 spec 是缺陷,而不是 runner 不稳定。[dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) 负责资源分配、状态恢复、同步、超时预算、平台差异与 teardown 规则;它的 [flake 诊断流程](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md)用于归类已经存在的概率性失败。
+fork 出的 worker 会同时运行多个 spec 文件,coverage gate 会拆成并发的 partition,与同一个 job 中的其它 gate 并排运行,而自托管 runner 共用同一台宿主机和同一个卷。被隔离的只有进程:端口、可预测路径、外部命名空间和继承而来的子进程都不隔离。为每个占用的资源负责到它的 teardown,并把「只有单独运行时才通过」的 spec 读作该 spec 的缺陷,而不是 runner 不稳定。[dsh-ci-test-reliability](../.agents/skills/dsh-ci-test-reliability/SKILL.md) 负责资源分配、状态恢复、同步、超时预算、平台差异与 teardown 规则;它的 [flake 诊断流程](../.agents/skills/dsh-ci-test-reliability/references/ci-flake-diagnosis.md)用于归类已经存在的概率性失败。
 
 ## 带密钥策略:推理(inference)在这里很便宜
 

+ 0 - 17
package.json

@@ -166,24 +166,7 @@
   },
   "devDependencies": {
     "@deepseek-ai/dsh-package-manifest": "workspace:^",
-    "@deepseek-ai/cordis": "workspace:^",
-    "@deepseek-ai/dsh-agent-loop": "workspace:^",
-    "@deepseek-ai/dsh-agent-loop-testkit": "workspace:^",
-    "@deepseek-ai/dsh-agent-presets": "workspace:^",
-    "@deepseek-ai/dsh-client-store": "workspace:^",
-    "@deepseek-ai/dsh-deque": "workspace:^",
-    "@deepseek-ai/dsh-llm": "workspace:^",
-    "@deepseek-ai/dsh-session": "workspace:^",
-    "@deepseek-ai/dsh-session-persistence": "workspace:^",
-    "@deepseek-ai/dsh-session-persistence-jsonl": "workspace:^",
-    "@deepseek-ai/dsh-session-projection": "workspace:^",
-    "@deepseek-ai/dsh-session-query": "workspace:^",
-    "@deepseek-ai/dsh-session-stats": "workspace:^",
-    "@deepseek-ai/dsh-session-title": "workspace:^",
-    "@deepseek-ai/dsh-session-turn-outline": "workspace:^",
-    "@deepseek-ai/dsh-token-meter": "workspace:^",
     "@deepseek-ai/dsh-tool-session-query": "workspace:^",
-    "@deepseek-ai/dsh-typert-protocol": "workspace:^",
     "@deepseek-ai/dsh-web-fetch-http": "workspace:^",
     "@stylistic/eslint-plugin": "^5.10.0",
     "@testing-library/dom": "^10.4.1",

+ 63 - 51
pnpm-lock.yaml

@@ -16,63 +16,12 @@ importers:
 
   .:
     devDependencies:
-      '@deepseek-ai/cordis':
-        specifier: workspace:^
-        version: link:vendor/cordis
-      '@deepseek-ai/dsh-agent-loop':
-        specifier: workspace:^
-        version: link:packages/core/agent-loop
-      '@deepseek-ai/dsh-agent-loop-testkit':
-        specifier: workspace:^
-        version: link:packages/test-support/agent-loop-testkit
-      '@deepseek-ai/dsh-agent-presets':
-        specifier: workspace:^
-        version: link:packages/preset/agent-presets
-      '@deepseek-ai/dsh-client-store':
-        specifier: workspace:^
-        version: link:packages/client/store
-      '@deepseek-ai/dsh-deque':
-        specifier: workspace:^
-        version: link:packages/util/deque
-      '@deepseek-ai/dsh-llm':
-        specifier: workspace:^
-        version: link:packages/llm/llm
       '@deepseek-ai/dsh-package-manifest':
         specifier: workspace:^
         version: link:packages/util/package-manifest
-      '@deepseek-ai/dsh-session':
-        specifier: workspace:^
-        version: link:packages/core/session
-      '@deepseek-ai/dsh-session-persistence':
-        specifier: workspace:^
-        version: link:packages/session/session-persistence
-      '@deepseek-ai/dsh-session-persistence-jsonl':
-        specifier: workspace:^
-        version: link:packages/session/session-persistence-jsonl
-      '@deepseek-ai/dsh-session-projection':
-        specifier: workspace:^
-        version: link:packages/session/session-projection
-      '@deepseek-ai/dsh-session-query':
-        specifier: workspace:^
-        version: link:packages/session-query/session-query
-      '@deepseek-ai/dsh-session-stats':
-        specifier: workspace:^
-        version: link:packages/session/session-stats
-      '@deepseek-ai/dsh-session-title':
-        specifier: workspace:^
-        version: link:packages/session/session-title
-      '@deepseek-ai/dsh-session-turn-outline':
-        specifier: workspace:^
-        version: link:packages/session/session-turn-outline
-      '@deepseek-ai/dsh-token-meter':
-        specifier: workspace:^
-        version: link:packages/llm/token-meter
       '@deepseek-ai/dsh-tool-session-query':
         specifier: workspace:^
         version: link:packages/session-query/tool-session-query
-      '@deepseek-ai/dsh-typert-protocol':
-        specifier: workspace:^
-        version: link:packages/typert/protocol
       '@deepseek-ai/dsh-web-fetch-http':
         specifier: workspace:^
         version: link:packages/web/web-fetch-http
@@ -600,6 +549,69 @@ importers:
         specifier: 8.21.0
         version: 8.21.0
 
+  benchmarks:
+    devDependencies:
+      '@deepseek-ai/cordis':
+        specifier: workspace:^
+        version: link:../vendor/cordis
+      '@deepseek-ai/dsh-agent':
+        specifier: workspace:^
+        version: link:../packages/core/agent
+      '@deepseek-ai/dsh-agent-loop':
+        specifier: workspace:^
+        version: link:../packages/core/agent-loop
+      '@deepseek-ai/dsh-agent-loop-testkit':
+        specifier: workspace:^
+        version: link:../packages/test-support/agent-loop-testkit
+      '@deepseek-ai/dsh-agent-presets':
+        specifier: workspace:^
+        version: link:../packages/preset/agent-presets
+      '@deepseek-ai/dsh-api-session-controller':
+        specifier: workspace:^
+        version: link:../packages/api/session-controller
+      '@deepseek-ai/dsh-client-store':
+        specifier: workspace:^
+        version: link:../packages/client/store
+      '@deepseek-ai/dsh-client-ui-chat':
+        specifier: workspace:^
+        version: link:../packages/client/ui-chat
+      '@deepseek-ai/dsh-deque':
+        specifier: workspace:^
+        version: link:../packages/util/deque
+      '@deepseek-ai/dsh-llm':
+        specifier: workspace:^
+        version: link:../packages/llm/llm
+      '@deepseek-ai/dsh-session':
+        specifier: workspace:^
+        version: link:../packages/core/session
+      '@deepseek-ai/dsh-session-persistence':
+        specifier: workspace:^
+        version: link:../packages/session/session-persistence
+      '@deepseek-ai/dsh-session-persistence-jsonl':
+        specifier: workspace:^
+        version: link:../packages/session/session-persistence-jsonl
+      '@deepseek-ai/dsh-session-projection':
+        specifier: workspace:^
+        version: link:../packages/session/session-projection
+      '@deepseek-ai/dsh-session-query':
+        specifier: workspace:^
+        version: link:../packages/session-query/session-query
+      '@deepseek-ai/dsh-session-stats':
+        specifier: workspace:^
+        version: link:../packages/session/session-stats
+      '@deepseek-ai/dsh-session-title':
+        specifier: workspace:^
+        version: link:../packages/session/session-title
+      '@deepseek-ai/dsh-session-turn-outline':
+        specifier: workspace:^
+        version: link:../packages/session/session-turn-outline
+      '@deepseek-ai/dsh-token-meter':
+        specifier: workspace:^
+        version: link:../packages/llm/token-meter
+      '@deepseek-ai/dsh-typert-protocol':
+        specifier: workspace:^
+        version: link:../packages/typert/protocol
+
   native/landlock-run:
     devDependencies:
       '@deepseek-ai/node-addon-landlock-run':

+ 2 - 0
pnpm-workspace.yaml

@@ -7,6 +7,8 @@ packages:
   - native/landlock-run/packages/*
   # Product assemblies over the package tier; apps/cli owns the `dsh` bin.
   - apps/*
+  # Private package owning repository-level benchmark dependencies.
+  - benchmarks
   - website
   # Deploy root of the single-exe build: a pure dependency manifest whose
   # closure is what the exe bundles and what the Python runtime distributes.

+ 1 - 1
scripts/doc-budgets.manifest.json

@@ -4,7 +4,7 @@
   "docs/architecture.md": 2400,
   "docs/cordis-primer.md": 600,
   "docs/defensive-patterns.md": 550,
-  "docs/testing.md": 1300,
+  "docs/testing.md": 1350,
   "packages/AGENTS.md": 750,
   "packages/README.md": 994
 }