ソースを参照

fix(benchmarks): keep Session opening fixture migratable to V3

Begin each synthetic step before its user surface message so the structural V2-to-V3 migration can reserve its protected system head without reordering historical events. Keep the workload size, frame partition, original chunk provenance, and every timing/heap budget unchanged.

Add a small real persistence migration/publication/reopen prerequisite that preserves all user and assistant content and original V0 bytes. Both this regression and bounded stderr-headline regression fail before the fix and pass afterward. Worker errors retain a bounded head and tail instead of hiding migration refusal behind a generic stack tail.

Validation: full Session-opening benchmark16/16 passed, including all six128MB completion cases; focused full type-aware lint passed. First Agent resume median178.7ms and reopen27.4ms are below unchanged450/100ms budgets.
Tianyi Cui 1 ヶ月 前
親
コミット
fcbadb07db

+ 2 - 2
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md
-2026-09-04-session-open-performance-gate.md: 2820c9d7d0e5b7d9382c7f8d6540154440175f26
-2026-09-04-session-open-performance-gate.zh.md: 965b9035074504870bcb2f1ca8166962c264d75a
+2026-09-04-session-open-performance-gate.md: 4ce62e18427525a3f66a96487bab50db5445fd2c
+2026-09-04-session-open-performance-gate.zh.md: c607b9090f7100be47d3629a2cd87c316a781a5f

+ 3 - 3
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.md

@@ -16,9 +16,9 @@ Linux pull requests run a required `node 24 / benchmarks` job that executes `pnp
 
 Required performance gates live under top-level `benchmarks/`, grouped by measured user path rather than package ownership. Host files use `*.bench.ts`, Client-face files use `*.bench.client.ts`, and scenario-specific workers and fixtures stay beside their benchmark without a benchmark suffix. Package-local `.perf.ts` files remain non-gating diagnostics; `scripts/` owns orchestration rather than benchmark cases.
 
-The Session benchmarks synthesize a released-v0 input from fixed parameters: 200 turns with 500 text deltas and 125 reasoning deltas per turn, for 127,400 logical events. The input uses Zstandard with fixed logical-row grouping and frame partitioning, so every run processes the same events, bytes, and frame distribution. The fixture constructs the immutable released-v0 physical rows directly instead of depending on a current-runtime historical encoder; compression and every measured read or migration entry point still use production code. Setup writes the input into a private temporary directory for each sample before timing starts; benchmarks never use recorded Sessions.
+The Session benchmarks synthesize a released-v0 input from fixed parameters: 200 turns with 500 text deltas and 125 reasoning deltas per turn, for 127,400 logical events. The input uses Zstandard with fixed logical-row grouping and frame partitioning, so every run processes the same events, bytes, and frame distribution. Each synthetic turn starts its step before appending user surface input, allowing the V2-to-V3 migration to reserve a protected system head without reordering history. The fixture constructs released-v0 physical rows directly instead of depending on a current-runtime historical encoder; compression and every measured read or migration entry point still use production code. Setup writes the input into a private temporary directory for each sample before timing starts; benchmarks never use recorded Sessions.
 
-Every Session endpoint runs at two user-lifecycle points. `first-open` starts with only the released V0 generation and therefore includes migration and successor publication. Setup produces `post-upgrade-reopen` once through that same production migration outside measurement, then copies both the unchanged V0 predecessor and published V2 successor into each sample root. Reopen samples use a fresh process, so they measure an upgraded user's later disk open without migration or process-local caches.
+Every Session endpoint runs at two user-lifecycle points. `first-open` starts with only the released V0 generation and therefore includes migration and successor publication. Setup produces `post-upgrade-reopen` once through that same production migration outside measurement, then copies both the unchanged V0 predecessor and published current-generation successor into each sample root. Reopen samples use a fresh process, so they measure an upgraded user's later disk open without migration or process-local caches.
 
 Each access-kind and endpoint sample runs in a fresh compiled Node child process. Module imports, Host service initialization, and fixture preparation finish before measurement; the measured process performs no extra parse warm-up. Normal-heap mode runs five independent samples, reports every sample plus minimum, median, and maximum, and enforces access-specific fixed budgets against the median. Another child runs the same path under a fixed 128 MB old-space limit and checks only that it completes; extra GC caused by the constrained heap does not enter the normal timing baseline.
 
@@ -35,7 +35,7 @@ The phase profile invokes each layer's production entry point explicitly and doe
 
 Normal-heap mode performs a fixed pair of explicit garbage collections after Host initialization and before the cold Session is touched, then records starting memory. It stops operation timing before performing the same garbage-collection sequence while the scenario's intended long-lived objects remain explicitly reachable, then records ending memory. The Agent-resume endpoint retains the Agent, Session, complete events, and normal service caches; its `heapUsed` delta is the primary resident-Session memory budget. Every scenario also reports `external`, `arrayBuffers`, post-GC RSS, and `process.resourceUsage().maxRSS`; the 128 MB mode prevents transient allocation peaks from being hidden by endpoint collection. Explicit garbage-collection time is excluded from operation timing.
 
-The performance gate does not duplicate semantic assertions owned by functional tests; it requires only that the target call completes and reaches its measured endpoint. The Client-fold benchmark continues to use the real `ConversationNodeAssembler` and every Chat Definition, and requires both the large window's absolute time and its scaling relative to the small window to remain below fixed budgets.
+A small untimed fixture prerequisite verifies current migration, message preservation, immutable V0 bytes, and successor reopen. Worker failures retain the first and last ten stderr lines, or fatal heap diagnostics, so setup rejection remains distinguishable from a budget breach. The timed performance cases do not duplicate semantic assertions owned by functional tests; it requires only that the target call completes and reaches its measured endpoint. The Client-fold benchmark continues to use the real `ConversationNodeAssembler` and every Chat Definition, and requires both the large window's absolute time and its scaling relative to the small window to remain below fixed budgets.
 
 Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Except for current-generation `open`, source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. Current-generation `open` uses a directly measured standard-runner expectation of 50 ms with only the 1.25× headroom, rounded up to a 63 ms budget. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
 

+ 3 - 3
.agents/notes/implemented/testing/2026-09-04-session-open-performance-gate.zh.md

@@ -16,9 +16,9 @@ Linux pull request 运行必需的 `node 24 / benchmarks` job,执行 `pnpm run
 
 必需性能 gate 位于顶层 `benchmarks/`,按被测用户路径而非 package 归属组织。Host 文件使用 `*.bench.ts`,Client 面文件使用 `*.bench.client.ts`,场景专属 worker 与 fixture 留在对应 benchmark 旁且不带 benchmark 后缀。包内 `.perf.ts` 文件仍是非门禁诊断;`scripts/` 负责编排而不承载 benchmark case。
 
-Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮 500 个 text delta 与 125 个 reasoning delta,共 127,400 个逻辑事件。输入使用 Zstandard,并固定 logical rows 的分组与 frame 拆分,使每次运行处理相同的事件、字节与 frame 分布。fixture 直接构造不可变的 released-v0 physical rows,不依赖当前 runtime 的历史 encoder;压缩以及所有被测读取和 migration 入口仍使用生产代码。输入在计时前写入每个样本独占的临时目录;benchmark 不使用录制的 Session。
+Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮 500 个 text delta 与 125 个 reasoning delta,共 127,400 个逻辑事件。输入使用 Zstandard,并固定 logical rows 的分组与 frame 拆分,使每次运行处理相同的事件、字节与 frame 分布。每个合成 turn 先开始 step,再追加 user surface 输入,使 V2-to-V3 migration 能预留受保护的 system head 而不重排历史。fixture 直接构造 released-v0 physical rows,不依赖当前 runtime 的历史 encoder;压缩以及所有被测读取和 migration 入口仍使用生产代码。输入在计时前写入每个样本独占的临时目录;benchmark 不使用录制的 Session。
 
-每个 Session endpoint 都针对用户生命周期中的两个时点运行。`first-open` 最初只有 released V0 generation,因此包含 migration 与后继 generation 发布。测试准备阶段在计时外通过同一套生产 migration 生成一次 `post-upgrade-reopen`,再把未改动的 V0 前代和已发布的 V2 后继一起复制到每个样本目录。Reopen 样本使用全新进程,因此测量用户升级完成后的磁盘再次打开,不包含 migration 或进程内 cache。
+每个 Session endpoint 都针对用户生命周期中的两个时点运行。`first-open` 最初只有 released V0 generation,因此包含 migration 与后继 generation 发布。测试准备阶段在计时外通过同一套生产 migration 生成一次 `post-upgrade-reopen`,再把未改动的 V0 前代和已发布的当前 generation 后继一起复制到每个样本目录。Reopen 样本使用全新进程,因此测量用户升级完成后的磁盘再次打开,不包含 migration 或进程内 cache。
 
 每个 access kind 与 endpoint 的样本都在全新、已编译的 Node 子进程中运行。模块加载、Host 服务初始化和 fixture 准备在测量开始前完成;测量进程不执行额外的预热解析。正常堆模式运行五个独立样本,报告全部样本及最小值、中位数和最大值,并以中位数执行各访问状态独立的固定预算。另一个子进程使用固定 128 MB old-space 上限运行同一路径,只判断能否完成;低堆限制引起的额外 GC 不进入正常时间基线。
 
@@ -35,7 +35,7 @@ Session benchmark 使用固定参数合成 released-v0 输入:200 轮,每轮
 
 正常堆模式在 Host 初始化完成且 Session 尚未访问时执行固定的两轮显式 GC,记录起点内存;操作计时结束后,在该场景要求的长期对象仍明确可达时再次执行同样的 GC,再记录终点内存。Agent resume 场景在终点保留 Agent、Session、完整 events 与正常服务 cache,它的 `heapUsed` 增量是常驻 Session 内存预算的主指标。每个场景同时报告 `external`、`arrayBuffers`、GC 后 RSS 和 `process.resourceUsage().maxRSS`;128 MB 模式继续防止瞬时分配峰值被终点 GC 隐藏。显式 GC 时间不计入操作时间。
 
-性能 gate 不重复功能测试的内容断言,只要求目标调用完成并到达对应的可观察终点。Client fold benchmark 继续使用真实 `ConversationNodeAssembler` 与全部 Chat Definition,要求大窗口的绝对时间和相对小窗口的缩放比均低于固定预算。
+一个不计时的小型 fixture 前置用例验证当前 migration、消息保留、V0 字节不变及后继再次打开。Worker 失败时保留 stderr 首尾各十行或致命堆错误,使准备阶段拒绝与预算超限可区分。计时性能用例不重复功能测试的内容断言,只要求目标调用完成并到达对应的可观察终点。Client fold benchmark 继续使用真实 `ConversationNodeAssembler` 与全部 Chat Definition,要求大窗口的绝对时间和相对小窗口的缩放比均低于固定预算。
 
 预算按各测量终点分别校准。两次 Node 24.19 x64 CI 运行的中位数最大相差 5.2%;其 CPU 密集型壁钟时间是 Node 24.18 arm64 参考运行的 1.95–2.06 倍。除当前 generation `open` 外,源码常量记录参考机器上的预期耗时;`ciTimeBudget()` 将其乘以实测的 2 倍 CI 时间系数和 1.25 倍波动余量。当前 generation `open` 使用标准运行器直接测得的 50 ms 预期值,仅乘以 1.25 倍余量,向上取整得到 63 ms 预算。GC 后增量堆与 Client fold 缩放预算不属于壁钟时间,因此只使用 1.25 倍余量。128 MB 完成性检查仍是独立的瞬时分配限制。由此得到的 first-open 时间上限、受限堆检查与 Client fold 上限都会拒绝已知退化。栈前参考提交固定为 `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5`,只用于校准和评审预算;CI 不 checkout 或执行历史仓库。预算是源码中的受评审常量,不由环境变量覆盖。
 

+ 70 - 2
benchmarks/session-open/session-open.bench.ts

@@ -1,9 +1,12 @@
 /** Required performance budgets for cold Session preparation, first history, and Agent resume. */
 
-import { copyFile, mkdir, mkdtemp, rm } from 'node:fs/promises'
+import { copyFile, mkdir, mkdtemp, readFile, readdir, rm } from 'node:fs/promises'
 import { tmpdir } from 'node:os'
 import { join } from 'node:path'
 import { afterAll, beforeAll, describe, expect, it } from 'vitest'
+import { Context } from '@deepseek-ai/cordis'
+import { Session, SessionId, SESSION_FORMAT_VERSION } from '@deepseek-ai/dsh-session'
+import JsonlSessionPersistence from '@deepseek-ai/dsh-session-persistence-jsonl'
 import {
   runBuiltBenchmarkWorker,
   type BuiltBenchmarkWorkerRun,
@@ -18,6 +21,7 @@ import type {
 } from './session-open.worker.ts'
 import {
   SYNTHETIC_CURRENT_GENERATION,
+  SYNTHETIC_SESSION_ID,
   SYNTHETIC_SESSION_DIRECTORY,
   SYNTHETIC_CURRENT_FILENAME,
   SYNTHETIC_V0_FILENAME,
@@ -152,7 +156,10 @@ function requireReport(
   if (run.report !== undefined) return run.report
   const stderrLines = run.stderr.trim().split('\n')
   const fatal = stderrLines.filter(line => /FATAL ERROR|heap limit|out of memory/i.test(line))
-  const detail = (fatal.length > 0 ? fatal : stderrLines.slice(-10)).join('\n')
+  const context = stderrLines.length <= 20
+    ? stderrLines
+    : [...stderrLines.slice(0, 10), '... stderr middle omitted ...', ...stderrLines.slice(-10)]
+  const detail = (fatal.length > 0 ? fatal : context).join('\n')
   const limit = heapLimitMb === undefined ? 'normal heap' : `${String(heapLimitMb)} MB old space`
   throw new Error(
     `${scenario} failed under ${limit}: exit=${String(run.exitCode)}, signal=${String(run.signal)}, `
@@ -293,6 +300,67 @@ describe('standard hosted reopen calibration', () => {
   })
 })
 
+describe('Session opening benchmark prerequisites', () => {
+  it('retains the exception headline and bounded stderr tail when a worker fails', () => {
+    const headline = 'SessionFormatUnsupportedError: source chronology cannot be migrated'
+    const stderr = [headline, ...Array.from({ length: 30 }, (_, index) => 'stack frame ' + String(index)), 'Node.js test'].join('\n')
+    const run: WorkerRun = { report: undefined, exitCode: 1, signal: null, timedOut: false, stderr }
+    expect(() => requireReport(run, 'agent-resume')).toThrow(headline)
+    expect(() => requireReport(run, 'agent-resume')).toThrow('Node.js test')
+    expect(() => requireReport(run, 'agent-resume')).not.toThrow('stack frame 15')
+    expect(() => requireReport({ ...run, stderr: 'FATAL ERROR: heap limit' }, 'agent-resume', 128))
+      .toThrow('128 MB old space')
+  })
+
+  it('migrates the generated workload and reopens its successor without changing V0', async () => {
+    const root = await mkdtemp(join(tmpdir(), 'dsh-session-bench-fixture-'))
+    const contexts: Context[] = []
+    const mount = async () => {
+      const ctx = new Context()
+      contexts.push(ctx)
+      await ctx.plugin(JsonlSessionPersistence, { root, compression: 'zstd' })
+      return ctx.sessionPersistence
+    }
+    try {
+      const facts = await writeSyntheticReleasedV0Session(root, { turns: 2, textDeltas: 12 })
+      expect({ events: facts.events, rows: facts.rows, frames: facts.frames }).toEqual({ events: 54, rows: 28, frames: 29 })
+      const original = await readFile(facts.path)
+      const directory = join(root, SYNTHETIC_SESSION_DIRECTORY)
+      const persistence = await mount()
+      const read = await persistence.open(SessionId(SYNTHETIC_SESSION_ID), 'read')
+      const initial = await read.read()
+      const session = Session.fromRestore(read.header.id, initial.events, read.header, read.inheritedEventCount, initial.eventState)
+      expect(read.header.version).toBe(SESSION_FORMAT_VERSION)
+      expect(initial.events.filter(event => event.type === 'system/message')).toHaveLength(1)
+      expect(session.deriveMessages().map(({ id, role, content }) => ({ id, role, content }))).toEqual(
+        [1, 2].flatMap(turn => [
+          { id: 'user-' + String(turn), role: 'user', content: [{ type: 'text', text: 'prompt ' + String(turn) }] },
+          { id: 'assistant-' + String(turn), role: 'assistant', content: [
+            { type: 'reasoning', text: 'r0 r1 r2 ' },
+            { type: 'text', text: Array.from({ length: 12 }, (_, index) => 'w' + String(index) + ' ').join('') },
+          ] },
+        ]),
+      )
+      await read.close()
+      expect(await readdir(directory)).toEqual([SYNTHETIC_V0_FILENAME])
+      const writer = await persistence.open(SessionId(SYNTHETIC_SESSION_ID), 'write')
+      expect((await writer.read()).events).toEqual(initial.events)
+      await writer.close()
+      expect(await readdir(directory)).toContain(SYNTHETIC_CURRENT_FILENAME)
+      const reopened = await (await mount()).open(SessionId(SYNTHETIC_SESSION_ID), 'read')
+      expect((await reopened.read()).events).toEqual(initial.events)
+      await reopened.close()
+      expect(await readFile(facts.path)).toEqual(original)
+    } finally {
+      try {
+        for (const ctx of contexts.reverse()) await ctx.fiber.dispose()
+      } finally {
+        await rm(root, { recursive: true, force: true })
+      }
+    }
+  })
+})
+
 describe('opening a large Session for first open and post-upgrade reopen', () => {
   const suite = new SessionOpenBenchmarkSuite()
 

+ 1 - 1
benchmarks/session-open/synthetic-released-v0-session.ts

@@ -67,13 +67,13 @@ class ReleasedV0FixtureBuilder {
 
   private appendTurn(turn: number, reasoningDeltaCount: number, textDeltaCount: number): void {
     this.appendEvent('turn/start', { turn })
+    this.appendEvent('step/start', { turn, step: 1 })
     this.appendEvent('user/message', {
       id: `user-${String(turn)}`,
       role: 'user',
       content: [{ type: 'text', text: `prompt ${String(turn)}` }],
       source: { kind: 'user' },
     }, { surfaceOp: 'append' })
-    this.appendEvent('step/start', { turn, step: 1 })
     const firstChunkSeq = this.appendChunk(turn, { type: 'block-start', index: 0, blockType: 'reasoning' })
     const reasoningDeltas = Array.from(
       { length: reasoningDeltaCount },