Просмотр исходного кода

test(perf): measure real history streams and live input overlap

Tianyi Cui 4 недель назад
Родитель
Сommit
7ac5e08246

+ 2 - 2
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md
-2026-09-06-frontend-performance-budgets.md: 0ebe97db9532c4922d2e0e8f2bd41613b9e80b6e
-2026-09-06-frontend-performance-budgets.zh.md: f58afd5aae763b145887204a63dd5b90d16566b4
+2026-09-06-frontend-performance-budgets.md: 92cdb27fa64efe304f501916f8f5bb205612e454
+2026-09-06-frontend-performance-budgets.zh.md: 38efa82251f3dcc3886051b434df4018c1665888

+ 11 - 11
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md

@@ -14,29 +14,29 @@ The existing serial benchmark inventory includes two frontend owners: [active re
 
 `build:bench` keeps the Node-only library and worker build. `test:bench` additionally builds the Web shell before running all cases; the required benchmark CI job provisions Chromium. Browser cases reuse the shipped-composition Web scaffold with private temporary roots and an atomically assigned loopback port. Only the nondeterministic model is replaced by synthetic replay. The scaffold Host runs under the existing Vitest source resolver; measured Client rendering runs built bundles in fresh Chromium processes. Browser wall times therefore include this test Host, transport, Playwright actionability, and rendering, and are not claims about a published Host process.
 
-The browser input contains 240 closed turns, 40 tool results, and 20 code fences, plus mixed-language prose and reasoning. Nine older-page actions exhaust this input from its observed 25-turn initial window; the readiness probe follows mounted turn growth rather than duplicating the pagination algorithm. Each sample uses a fresh scaffold and browser. Setup, seeding, browser launch, initial shell load, and sidebar expansion are excluded from open timing. Open ends at transcript availability and an editable composer; page and navigation timings end at their target DOM state. Two animation frames include a rendering opportunity, not hardware presentation or a guarantee that every offscreen node painted.
+The browser input contains 240 closed turns, 40 tool results, and 20 code fences, plus mixed-language prose and reasoning. Historical Assistant records carry matching compact streams built through the production accumulator with 12-character reasoning/text deltas and 8-character tool-argument deltas; empty streams would omit stored and transferred payload costs. Nine older-page actions exhaust this input from its observed 25-turn initial window; the readiness probe follows mounted turn growth rather than duplicating the pagination algorithm. Each sample uses a fresh scaffold and browser. Setup, seeding, browser launch, initial shell load, and sidebar expansion are excluded from open timing. Open ends at transcript availability and an editable composer; page and navigation timings end at their target DOM state. Two animation frames include a rendering opportunity, not hardware presentation or a guarantee that every offscreen node painted.
 
-The continuation sends 120 text deltas at 8 ms replay pacing. It records click-to-first-visible-reply, trusted draft typing while the completion marker is absent, complete reply wall time through settled persistence, and Chromium main-thread task duration. The complete wall budget adds the fixed 992 ms scripted pacing to a scaled overhead allowance; input and completion have their own enforced budgets. Post-GC browser heap and DOM counts remain diagnostics because one endpoint does not prove a leak.
+The continuation sends 120 text deltas at 8 ms replay pacing. It records click-to-first-visible-reply, trusted draft typing whose first actual input event observes the first reply but no completion marker, complete reply wall time through settled persistence and the new rendered turn-tail, and Chromium main-thread task duration. The complete wall budget adds the fixed 992 ms scripted pacing to a scaled overhead allowance; input and completion have their own enforced budgets. Post-GC browser heap and DOM counts remain diagnostics because one endpoint does not prove a leak.
 
 Reconnect uses three fresh compiled plain-Node children. Each creates a 100,000-delta reasoning prefix with distinct timestamps and two compact records before timing `ClientAssistantStream.replace()`. GC precedes the baseline and follows replacement while the result remains reachable; replacement time excludes both collections. The report consumes the result after collection and checks that the next dense live frame remains accepted. This measures reconstruction, not transport, rendering, or an entire reconnect workflow.
 
 ## Calibration
 
-Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at product revision `925e012340`, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. Source reference constants round above observed values; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The existing scale comes from Node CI calibration, not a measured x64 browser comparison; browser-specific runner calibration remains an explicit gap.
+Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at product revision `925e012340`, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. Source reference constants retain the original allowances rather than increasing them after the compact-payload correction; the corrected 302.25 ms paging median exceeds its 260 ms reference allowance but remains below its 650 ms CI limit; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The existing scale comes from Node CI calibration, not a measured x64 browser comparison; browser-specific runner calibration remains an explicit gap.
 
 | Endpoint | Measured median | Reference allowance | CI limit |
 |---|---:|---:|---:|
-| Browser open | 166.77 ms | 200 ms | 500 ms |
-| Slowest older page | 245.52 ms | 260 ms | 650 ms |
-| First Trajectory | 133.45 ms | 160 ms | 400 ms |
-| First reply | 1033.07 ms | 1100 ms | 2750 ms |
-| Stream main-thread task | 1668.46 ms | 1800 ms | 4500 ms |
-| Draft typing | 228.08 ms | 500 ms | 1250 ms |
-| Complete response | 1699.26 ms | 1000 ms overhead + 992 ms pacing | 3492 ms |
+| Browser open | 194.04 ms | 200 ms | 500 ms |
+| Slowest older page | 302.25 ms | 260 ms | 650 ms |
+| First Trajectory | 140.44 ms | 160 ms | 400 ms |
+| First reply | 1093.12 ms | 1100 ms | 2750 ms |
+| Stream main-thread task | 1712.99 ms | 1800 ms | 4500 ms |
+| Draft typing | 126.71 ms | 500 ms | 1250 ms |
+| Complete response | 1751.04 ms | 1000 ms overhead + 992 ms pacing | 3492 ms |
 | Reconnect replacement | 13.83 ms | 16 ms | 40 ms |
 | Reconnect retained heap | 23.03 MiB | 24 MiB | 30 MiB |
 
-Draft typing spans 138.80–420.28 ms across the three isolated samples; its reference covers that observed spread instead of treating the median as a per-keystroke bound. No budget is an environment override. Temporary zero allowances exercise every rejection path; these negative controls prove enforcement, not an optimization or a historical regression.
+Draft typing spans 101.58–415.26 ms across the three isolated samples; its reference covers that observed spread instead of treating the median as a per-keystroke bound. No budget is an environment override. Temporary zero allowances exercise every rejection path; these negative controls prove enforcement, not an optimization or a historical regression. A separate control waits for the final reply marker before typing and fails the actual-input overlap assertion. The compact synthetic JSONL is 3,262,577 bytes; all three corrected samples report an overlapping trusted input event and end after the 241st rendered turn-tail.
 
 ## Alternatives considered
 

+ 11 - 11
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.zh.md

@@ -14,29 +14,29 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 `build:bench` 保留仅 Node 的 library 与 worker 构建。`test:bench` 额外构建 Web shell 后再运行所有用例;必需的基准 CI job 安装 Chromium。浏览器用例复用产品组合的 Web scaffold,使用私有临时目录和原子分配的回环端口。只有不确定的模型被合成重放替代。scaffold Host 通过现有 Vitest 源码解析器运行;被测 Client 渲染在全新 Chromium 进程中执行构建后的 bundle。因此浏览器壁钟时间包含测试 Host、传输、Playwright 可交互性等待及渲染,不代表发布版 Host 进程。
 
-浏览器输入包含 240 个已关闭轮次、40 个工具结果和 20 个代码块,以及混合语言正文和推理。从观察到的初始 25 轮窗口开始,九次更早分页操作读完该输入;就绪探针跟踪已挂载轮次增长,不复制分页算法。每个样本使用全新 scaffold 和浏览器。环境准备、数据播种、浏览器启动、初始 shell 加载及侧栏展开不计入打开时间。打开测量在对话可用且输入框可编辑时结束;分页与导航测量在目标 DOM 状态出现时结束。两次动画帧包含一次渲染机会,不代表硬件显示或保证每个屏幕外节点都已绘制。
+浏览器输入包含 240 个已关闭轮次、40 个工具结果和 20 个代码块,以及混合语言正文和推理。历史 Assistant 记录携带匹配的紧凑 stream,通过生产 accumulator 按 12 字符推理/文本 delta 和 8 字符工具参数 delta 构建;空 stream 会遗漏存储与传输负载成本。从观察到的初始 25 轮窗口开始,九次更早分页操作读完该输入;就绪探针跟踪已挂载轮次增长,不复制分页算法。每个样本使用全新 scaffold 和浏览器。环境准备、数据播种、浏览器启动、初始 shell 加载及侧栏展开不计入打开时间。打开测量在对话可用且输入框可编辑时结束;分页与导航测量在目标 DOM 状态出现时结束。两次动画帧包含一次渲染机会,不代表硬件显示或保证每个屏幕外节点都已绘制。
 
-续接以 8 ms 重放间隔发送 120 个文本 delta。它记录点击到首段可见回复的时间、完成标记尚未出现时的真实草稿键入、直到持久化结算的完整回复壁钟时间,以及 Chromium 主线程任务时间。完整壁钟预算在缩放后的额外开销额度上加固定的 992 ms 脚本节奏;输入和完成均有独立执行的预算。强制 GC 后的浏览器 heap 和 DOM 数量仍仅供诊断,因为单个终点不能证明泄漏。
+续接以 8 ms 重放间隔发送 120 个文本 delta。它记录点击到首段可见回复的时间、首个实际输入事件观察到首段回复且完成标记尚未出现时的真实草稿键入、直到持久化结算并渲染新 turn-tail 的完整回复壁钟时间,以及 Chromium 主线程任务时间。完整壁钟预算在缩放后的额外开销额度上加固定的 992 ms 脚本节奏;输入和完成均有独立执行的预算。强制 GC 后的浏览器 heap 和 DOM 数量仍仅供诊断,因为单个终点不能证明泄漏。
 
 重连使用三个全新编译后的纯 Node 子进程。各进程在计时 `ClientAssistantStream.replace()` 前创建包含不同时间戳、两条紧凑记录和 100,000 个 delta 的推理前缀。在基线前执行 GC,并在结果仍可达时于替换后再次 GC;替换时间不含两次回收。报告在回收后消费结果,并检查下一个稠密序号的实时 frame 仍被接受。这测量重建,不测量传输、渲染或完整重连工作流。
 
 ## 校准
 
-在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和产品版本 `925e012340` 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。源码参考常量向上取整覆盖观察值;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。现有倍率来自 Node CI 校准,并非实测 x64 浏览器对比;浏览器专用 runner 校准仍是明确缺口。
+在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和产品版本 `925e012340` 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。源码参考常量保留原额度,不因紧凑负载修正而提高;修正后的分页中位数 302.25 ms 超过 260 ms 参考额度,但仍低于 650 ms CI 限制;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。现有倍率来自 Node CI 校准,并非实测 x64 浏览器对比;浏览器专用 runner 校准仍是明确缺口。
 
 | 终点 | 实测中位数 | 参考额度 | CI 限制 |
 |---|---:|---:|---:|
-| 浏览器打开 | 166.77 ms | 200 ms | 500 ms |
-| 最慢更早分页 | 245.52 ms | 260 ms | 650 ms |
-| 首次 Trajectory | 133.45 ms | 160 ms | 400 ms |
-| 首段回复 | 1033.07 ms | 1100 ms | 2750 ms |
-| 流式主线程任务 | 1668.46 ms | 1800 ms | 4500 ms |
-| 草稿键入 | 228.08 ms | 500 ms | 1250 ms |
-| 完整回复 | 1699.26 ms | 1000 ms 额外开销 + 992 ms 节奏 | 3492 ms |
+| 浏览器打开 | 194.04 ms | 200 ms | 500 ms |
+| 最慢更早分页 | 302.25 ms | 260 ms | 650 ms |
+| 首次 Trajectory | 140.44 ms | 160 ms | 400 ms |
+| 首段回复 | 1093.12 ms | 1100 ms | 2750 ms |
+| 流式主线程任务 | 1712.99 ms | 1800 ms | 4500 ms |
+| 草稿键入 | 126.71 ms | 500 ms | 1250 ms |
+| 完整回复 | 1751.04 ms | 1000 ms 额外开销 + 992 ms 节奏 | 3492 ms |
 | 重连替换 | 13.83 ms | 16 ms | 40 ms |
 | 重连保留 heap | 23.03 MiB | 24 MiB | 30 MiB |
 
-三个隔离样本中的草稿键入时间为 138.80–420.28 ms;参考额度覆盖观察到的波动,而不把中位数作为单次按键上限。预算不能通过环境变量覆盖。临时零额度覆盖每条拒绝路径;这些负向对照证明预算执行,而非优化或历史回归。
+三个隔离样本中的草稿键入时间为 101.58–415.26 ms;参考额度覆盖观察到的波动,而不把中位数作为单次按键上限。预算不能通过环境变量覆盖。临时零额度覆盖每条拒绝路径;这些负向对照证明预算执行,而非优化或历史回归。另一项对照在键入前等待最终回复标记,实际输入重叠断言因此失败。紧凑合成 JSONL 为 3,262,577 字节;三个修正样本均报告重叠的真实输入事件,并在第 241 个 turn-tail 渲染后结束。
 
 ## 考虑过的替代方案
 

+ 2 - 2
benchmarks/long-session-browser/README.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write benchmarks/long-session-browser/README.md
-README.md: 53385706703736a76568b0b55141d40731068ed2
-README.zh.md: 5e709ec6bbd53a420fcf8ad4f4f414b4552549b6
+README.md: 421f2a904e0b40b55b4b5cdeb3a15c66a3e68df9
+README.zh.md: ee26bb28d5c080d202b0285403daff7ab28563b3

+ 2 - 2
benchmarks/long-session-browser/README.md

@@ -10,8 +10,8 @@ This reference describes the required Chromium workflow in [long-session.bench.t
 
 ## Measurements
 
-Three fresh browser processes and scaffold worlds produce raw samples and median verdicts. Open and paging end after the expected transcript state and two animation frames; this includes a rendering opportunity, not a hardware presentation timestamp. Paging reports every page and gates the median of each sample’s slowest page. Stream reports first visible reply, trusted draft typing, complete reply wall time, and Chromium main-thread task duration. Heap after forced GC and DOM counts are diagnostics, not leak budgets.
+Three fresh browser processes and scaffold worlds produce raw samples and median verdicts. Open and paging end after the expected transcript state and two animation frames; this includes a rendering opportunity, not a hardware presentation timestamp. Paging reports every page and gates the median of each sample’s slowest page. Stream reports first visible reply, trusted draft typing, complete reply wall time, and Chromium main-thread task duration. The actual first input event must observe an unfinished reply; completion waits for the new rendered turn-tail after Host settlement. Heap after forced GC and DOM counts are diagnostics, not leak budgets.
 
-The fixture contains mixed-language prompts, prose, reasoning, 20 code fences, and 40 synthetic tool results. No model, tool, external network, recorded Session, or private Harness home supplies its content. Streaming uses 120 text deltas at 8 ms replay pacing through the real composer, agent loop, transport, and persistence.
+The fixture contains mixed-language prompts, prose, reasoning, 20 code fences, and 40 synthetic tool results. Every historical Assistant includes a compact stream built by the production accumulator from matching reasoning, text, tool arguments, usage, and finish chunks. No model, tool, external network, recorded Session, or private Harness home supplies its content. Streaming uses 120 text deltas at 8 ms replay pacing through the real composer, agent loop, transport, and persistence.
 
 The [decision record](../../.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md) owns calibration, exclusions, and alternatives. The larger [manual diagnostic](../../apps/web/tests/complex-history.perf.ts) remains separate.

+ 2 - 2
benchmarks/long-session-browser/README.zh.md

@@ -10,8 +10,8 @@
 
 ## 测量
 
-三个全新浏览器进程与 scaffold 环境产生原始样本及中位数判定。打开和分页在预期对话状态出现且经过两次动画帧后结束;这包含一次渲染机会,而非硬件显示时间戳。分页报告每一页,并对各样本最慢分页时间的中位数执行预算检查。流式报告首段可见回复、真实草稿键入、完整回复壁钟时间和 Chromium 主线程任务时间。强制 GC 后的 heap 与 DOM 数量仅供诊断,不作为泄漏预算。
+三个全新浏览器进程与 scaffold 环境产生原始样本及中位数判定。打开和分页在预期对话状态出现且经过两次动画帧后结束;这包含一次渲染机会,而非硬件显示时间戳。分页报告每一页,并对各样本最慢分页时间的中位数执行预算检查。流式报告首段可见回复、真实草稿键入、完整回复壁钟时间和 Chromium 主线程任务时间。实际首个输入事件必须观察到未完成的回复;完成测量在 Host 结算后等待新 turn-tail 渲染。强制 GC 后的 heap 与 DOM 数量仅供诊断,不作为泄漏预算。
 
-fixture(测试前置数据)包含混合语言提示、正文、推理、20 个代码块和 40 个合成工具结果。其内容不来自模型、工具、外部网络、录制 Session 或私有 Harness 主目录。流式回复以 8 ms 重放间隔发送 120 个文本 delta,经过真实输入框、agent loop(智能体循环)、传输与持久化。
+fixture(测试前置数据)包含混合语言提示、正文、推理、20 个代码块和 40 个合成工具结果。每条历史 Assistant 都含紧凑 stream,由生产 accumulator 从匹配的推理、文本、工具参数、usage 和 finish chunk 构建。其内容不来自模型、工具、外部网络、录制 Session 或私有 Harness 主目录。流式回复以 8 ms 重放间隔发送 120 个文本 delta,经过真实输入框、agent loop(智能体循环)、传输与持久化。
 
 [决策记录](../../.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.zh.md)拥有校准、排除项与替代方案。更大规模的[手动诊断](../../apps/web/tests/complex-history.perf.ts)保持独立。

+ 15 - 5
benchmarks/long-session-browser/long-session.bench.ts

@@ -40,7 +40,7 @@ function median(values: number[]): number {
 
 it('opens, pages, navigates and streams into a 240-turn browser history', async () => {
   if (webSnapshotMode() !== 'replay') throw new Error('browser benchmarks require keyless replay mode')
-  const samples: { open: number; page: number; trajectory: number; first: number; streamTask: number; streamWall: number; input: number; heapMb: number; nodes: number }[] = []
+  const samples: { open: number; page: number; trajectory: number; first: number; streamTask: number; streamWall: number; input: number; inputOverlapped: boolean; heapMb: number; nodes: number }[] = []
   for (let sample = 0; sample < SAMPLES; sample++) {
     const failures: unknown[] = []
     const root = await mkdtemp(join(tmpdir(), 'dsh-browser-benchmark-'))
@@ -49,7 +49,9 @@ it('opens, pages, navigates and streams into a 240-turn browser history', async
       await writeFile(replayOverride, JSON.stringify([{ kind: 'chunks', chunks: syntheticReply() }]))
       const scaffold = await launchWebScaffold({ replayFixture: join(root, 'override-only.jsonl'), replayOverride, paceMs: PACE_MS, replayContextWindow: 10000000 })
       try {
-        await seedSession(scaffold, syntheticHistory(), SESSION_ID)
+        const history = syntheticHistory()
+        await seedSession(scaffold, history, SESSION_ID)
+        console.log(JSON.stringify({ benchmark: 'long-session-browser/fixture', bytes: Buffer.byteLength(history) }))
         const browser = await chromium.launch({ headless: true })
         try {
           const page = await newEnglishPage(browser)
@@ -100,16 +102,24 @@ it('opens, pages, navigates and streams into a 240-turn browser history', async
           await page.getByText(FIRST, { exact: false }).last().waitFor()
           await painted(page)
           const first = performance.now() - started
-          expect(await page.getByText(DONE, { exact: false }).count()).toBe(0)
-          // Trusted keyboard input while the response is live, rather than a synthetic heartbeat.
+          await composer.evaluate((element, markers) => {
+            element.addEventListener('input', (event) => {
+              const transcript = document.querySelector('[data-conversation-scroll]')?.textContent ?? ''
+              element.setAttribute('data-benchmark-input-overlap', String(event.isTrusted && transcript.includes(markers.first) && !transcript.includes(markers.done)))
+            }, { once: true })
+          }, { first: FIRST, done: DONE })
+          // Observe the actual trusted input event, not state before asynchronous click/typing.
           const input = await measure(page, async () => {
             await composer.click()
             await page.keyboard.type('next synthetic question')
             await expect.poll(() => composer.textContent()).toBe('next synthetic question')
           })
+          const inputOverlapped = await composer.getAttribute('data-benchmark-input-overlap') === 'true'
+          expect(inputOverlapped).toBe(true)
           await page.getByText(DONE, { exact: false }).last().waitFor()
           const settlement = await settled
           if (!settlement.ok) throw settlement.error
+          await page.waitForFunction(({ selector, expected }) => document.querySelectorAll(selector).length === expected, { selector: TAIL, expected: HISTORY_TURNS + 1 })
           await painted(page)
           const streamWall = performance.now() - started
           const streamTask = await taskMs(cdp) - beforeTask
@@ -117,7 +127,7 @@ it('opens, pages, navigates and streams into a 240-turn browser history', async
           const metrics = (await cdp.send('Performance.getMetrics')).metrics
           const heap = metrics.find(metric => metric.name === 'JSHeapUsedSize')
           if (heap === undefined) throw new Error('Chromium heap metric missing')
-          samples.push({ open, page: Math.max(...pages), trajectory, first, streamTask, streamWall, input, heapMb: heap.value / 1048576, nodes: await page.locator('*').count() })
+          samples.push({ open, page: Math.max(...pages), trajectory, first, streamTask, streamWall, input, inputOverlapped, heapMb: heap.value / 1048576, nodes: await page.locator('*').count() })
           console.log(JSON.stringify({ benchmark: 'long-session-browser/sample', sample, initialTurns, pages, ...samples.at(-1) }))
           expect(consoleWatch.pageErrors).toEqual([])
           expect(consoleWatch.warnings).toEqual([])

+ 28 - 5
benchmarks/long-session-browser/synthetic-history.ts

@@ -1,6 +1,7 @@
 /** Synthetic current-generation history and paced reply for browser measurements. */
 import { createAssistantMessage, createUserMessage, createToolResultMessage, ToolCallId } from '@deepseek-ai/dsh-llm'
 import type { StreamChunk } from '@deepseek-ai/dsh-llm'
+import { AssistantStreamAccumulator } from '@deepseek-ai/dsh-llm/assistant-stream'
 import { Session, SessionId, SESSION_FORMAT_VERSION } from '@deepseek-ai/dsh-session'
 import type {} from '@deepseek-ai/dsh-session-title'
 
@@ -36,20 +37,42 @@ export function syntheticHistory(): string {
     const code = turn % 12 === 0
       ? '\n\n```ts\n' + Array.from({ length: 60 }, (_, i) => 'const value' + String(i) + ' = ' + String(i)).join('\n') + '\n```'
       : ''
+    const reasoning = 'Compare the synthetic module and test. '.repeat(40)
+    const text = 'Synthetic answer ' + String(turn) + '. ' + 'Preserve ordering and validate the output. '.repeat(30) + code
+    const args = '{"path":"src/example.ts"}'
+    const stream = new AssistantStreamAccumulator()
+    let time = 1700000000000 + turn * 10000
+    const push = (chunk: StreamChunk): void => { stream.push({ time: time++, chunk }) }
+    for (const [index, block] of [{ type: 'reasoning' as const, text: reasoning }, { type: 'text' as const, text }].entries()) {
+      push({ type: 'block-start', index, blockType: block.type })
+      for (let offset = 0; offset < block.text.length; offset += 12) {
+        push({ type: block.type === 'reasoning' ? 'reasoning-delta' : 'text-delta', index, text: block.text.slice(offset, offset + 12) })
+      }
+      push({ type: 'block-end', index, block })
+    }
+    if (tool) {
+      push({ type: 'block-start', index: 2, blockType: 'tool-call' })
+      for (let offset = 0; offset < args.length; offset += 8) {
+        push({ type: 'tool-call-delta', index: 2, id: callId, ...offset === 0 ? { name: 'synthetic_tool' } : {}, argumentsDelta: args.slice(offset, offset + 8) })
+      }
+      push({ type: 'block-end', index: 2, block: { type: 'tool-call', id: callId, name: 'synthetic_tool', arguments: args } })
+    }
+    push({ type: 'usage', usage: { inputTokens: 4000, outputTokens: 800 } })
+    push({ type: 'finish', reason: { kind: tool ? 'tool-calls' : 'stop' } })
     session.append('assistant/message', {
-      turn, step: 1, stream: [],
+      turn, step: 1, stream: [...stream.snapshot()],
       message: createAssistantMessage({
         source: { provider: 'deepseek-official', model: 'deepseek-v4-flash' },
         content: [
-          { type: 'reasoning', text: 'Compare the synthetic module and test. '.repeat(40) },
-          { type: 'text', text: 'Synthetic answer ' + String(turn) + '. ' + 'Preserve ordering and validate the output. '.repeat(30) + code },
-          ...tool ? [{ type: 'tool-call' as const, id: callId, name: 'synthetic_tool', arguments: '{"path":"src/example.ts"}' }] : [],
+          { type: 'reasoning', text: reasoning },
+          { type: 'text', text },
+          ...tool ? [{ type: 'tool-call' as const, id: callId, name: 'synthetic_tool', arguments: args }] : [],
         ],
       }),
       usage: { inputTokens: 4000, outputTokens: 800 },
     }, { surfaceOp: 'append' })
     if (tool) {
-      const call = session.append('tool/call', { turn, step: 1, callId, name: 'synthetic_tool', arguments: '{"path":"src/example.ts"}' })
+      const call = session.append('tool/call', { turn, step: 1, callId, name: 'synthetic_tool', arguments: args })
       session.append('tool/result', { turn, step: 1, message: createToolResultMessage({
         callId, isError: false, content: [{ type: 'text', text: 'Synthetic tool output line.\n'.repeat(160) }],
       }) }, { surfaceOp: 'append', sourceEventSeqs: [call.seq] })