فهرست منبع

test(perf): baseline tool-heavy backend workflows

Tianyi Cui 1 ماه پیش
والد
کامیت
c599ef87c4

+ 6 - 0
.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.md
+2026-09-06-backend-continuation-performance.md: 4c18e440c98135f140c4957a7c5e7c73188c694c
+2026-09-06-backend-continuation-performance.zh.md: 2cf1793faf39ab0c3bc10b1932db15e64961e8b8

+ 60 - 0
.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.md

@@ -0,0 +1,60 @@
+# Agent Note: Performance baselines for tool-heavy backend continuation
+
+Status: implemented
+
+English | [中文](2026-09-06-backend-continuation-performance.zh.md)
+
+## Problem
+
+Opening one Session does not measure the repeated cost of preparing model requests after a long tool conversation, executing another tool-heavy turn, or discovering multiple inactive fork children. The [Session-opening gate](2026-09-04-session-open-performance-gate.md) covers first history and activation but deliberately stops before new model work. Its text/reasoning workload also lacks historical tool-call arguments and large tool results.
+
+## Decision
+
+The [agent-continuation benchmark](../../../../benchmarks/agent-continuation/agent-continuation.bench.ts) adds three scenario groups, including a shipped-profile variant, without changing product implementations. They use current-generation Zstandard Sessions authored through production append, stream accumulation, and persistence APIs. A separate seed process creates the deterministic source before measurement; each sample copies that source into its private root and starts a fresh compiled plain-Node worker. No recorded Session, ambient repository, network, private Harness home, or deployed GUI supplies input.
+
+The shared history has 800 completed two-step turns, four tool calls per turn, and 2,048-character tool results: 13,600 events and 5,600 model messages. Each assistant reply carries reasoning, text, and compact streamed records; tool replies additionally carry fragmented arguments. Fixed timestamps and ids describe the seed. Live synthetic replies use the real loop's clocks and ids without overriding process globals.
+
+| Case | Timed operation | Endpoint |
+|---|---|---|
+| Request history | After unmeasured cold resume, deliver 40 sequential text-only turns over the tool-heavy history, then flush | Idle Agent with all 40 model requests completed; reports turn and final-flush time separately |
+| Tool continuation | Cold resume, 20 sequential turns with eight parallel-safe synthetic tool calls and a final reply per turn, then flush | Idle Agent with 40 model requests and 160 completed tool executions; reports resume, turns, and final flush separately |
+| Shipped SDK workflow | Launch built dsh with the sdk-minimal profile, deliver 100 sequential turns with eight real file-view calls per turn, then close the SDK | SDK receives 200 assistant messages and 800 successful file results; includes Loader boot, stdio JSON-RPC, persistence, and shutdown |
+| Child catalog | List 16 inactive seeded fork children twice through the real subagent and Session query services | Two complete healthy catalogs with observations released; each child inherits 80 tool-heavy turns and owns its descriptor after the exact fork cut |
+
+The tool execution pipeline, request preparation, Session projections required by those services, persistence, and catalog observations remain production code. Only the model adapter and bounded tool body are synthetic. The adapter retains a request counter, not request objects, so the fixture cannot manufacture a growing retention cost. Sequential input means each idle interval belongs to the one request delivered by this worker; it does not generalize idle to a per-message completion API under concurrent input.
+
+Five samples report raw wall time, CPU user/system time, peak RSS, endpoint counts, and the minimum, median, and maximum total wall time. Budgets enforce the unrounded median. Continuation additionally measures retained heap against an initialized Host: two explicit GCs separated by an event-loop yield precede and follow the timed operation, while the idle Agent remains reachable. The measured delta therefore includes the resident historical Session and live additions, not just newly appended turns. GC and teardown are outside timing; flush is inside. Request-history retention starts after resume and is diagnostic only. Catalog peak RSS is diagnostic; no retained-heap budget claims to measure already-released child observations.
+
+The parent bounds every child to 60 seconds, checks timeout, signal, exit, and report independently, awaits process close, and removes private roots after failures. Context and Agent teardown run in finally blocks. Seed processes cannot warm the measured process's caches. Filesystem caches are not forcibly evicted: cold means a fresh process, not cold physical storage.
+
+## Calibration evidence
+
+The implementation reference is `925e012340f033f0521e802ba8569ce6dd7ef1ac` on Apple M4 Pro, macOS arm64, Node 24.19.0. Two exclusive five-sample runs use the same seed and no product optimization. Durations below are milliseconds; source expectations round above the observed run medians rather than imposing an unimplemented optimization target.
+
+| Case | Run 1 raw totals | Run 2 raw totals | Medians | Reference expectation | CI budget |
+|---|---|---|---|---:|---:|
+| Request history | 209.134, 210.333, 208.959, 236.355, 238.685 | 222.833, 213.911, 208.089, 211.494, 209.137 | 210.333 / 211.494 | 220 | 550 |
+| Tool continuation | 358.953, 324.790, 318.861, 320.119, 322.896 | 324.280, 321.952, 340.409, 325.470, 324.312 | 322.896 / 324.312 | 340 | 850 |
+| Child catalog | 318.730, 309.006, 311.404, 308.565, 310.105 | 308.670, 310.030, 280.086, 303.084, 284.829 | 310.105 / 303.084 | 320 | 800 |
+
+Continuation retains approximately 22.295 MiB; its source expectation is 23 MiB and its budget is 28.75 MiB. Time expectations use the existing [calibration helper](../../../../benchmarks/support/calibration.ts): 2× shared CI time scale and 1.25× variance headroom. Memory uses only 1.25× headroom. The scale is inherited from the existing lane's calibration, not a new Linux measurement of these cases; CI evidence remains necessary when runner characteristics change. Baseline budgets protect the measured implementation; tighter budgets belong with a measured behavior-preserving fix.
+
+A separate plain-Node request-history CPU profile attributes 132.876 ms of sampled self time to deepFreeze called by buildRequest during a 211.300 ms operation. This identifies repeated traversal of already-frozen history as a focused investigation target, not a proven optimization result. Catalog first/repeat timings remain separate because a second listing still reads body-bearing seeded children after observations are released.
+
+The shipped SDK variant completes 100 turns, 200 requests, and 800 real file reads. Its five-sample smoke totals are 1,521.773, 1,463.465, 1,689.701, 1,365.485, and 1,417.106 ms (median 1,463.465 ms); a full-suite repeat reports 1,596.183, 1,784.536, 2,120.082, 1,405.365, and 1,355.894 ms (median 1,596.183 ms). Its 1,700 ms reference expectation yields a 4,250 ms CI budget. The repeat also slows the unchanged service cases, so it is validation under variable host load rather than evidence to relax their exclusive calibration. The SDK process receives an allowlisted environment and private home/workspace. A 40-second deadline starts SDK shutdown; every path awaits the same memoized close promise before the outer worker’s 60-second deadline. Profile timing includes boot, all turns, and shutdown, reported separately; no parent-process CPU or heap metric is presented as server memory. The adapter does not serialize requests for an external model provider.
+
+## Alternatives considered
+
+**Repeat existing migration and first-open variants.** Rejected: those twelve cases already distinguish read-only preparation from writable publication. These cases use the current generation and begin or continue actual model work, or enumerate a corpus rather than open one Session.
+
+**Measure only deriveMessages.** Rejected: its incremental cache does not include complete request freezing, adapter dispatch, live append, or persistence. Actual sequential requests protect the cost the Agent pays per step.
+
+**Use only unseeded children with warm projection-cache rows.** Rejected: that path bypasses body observations and misses the exact inherited-cut requirement of fork children. The catalog intentionally omits the optional projection cache and reports the seeded fallback path; it does not characterize cache-hit discovery.
+
+**Apply an optimization and its desired budget together with the first measurements.** Rejected: a baseline-only layer remains independently mergeable and records the current workload before attribution or implementation changes. Source constants cannot be overridden by environment variables.
+
+## Consequences
+
+The lane adds four cases in three scenario groups and twenty measured workers, plus two seed processes. The integrated continuation case spans resume through completed model/tool work and durable flush. The shipped SDK workflow additionally includes profile boot, SDK transport, real file tools, and shutdown; only its model adapter is synthetic. It starts a fresh Session because the public SDK prompt API creates rather than resumes stored identities. Neither path includes network model latency, provider-specific request serialization, optional user plugins, compaction, failed tool results, images, cancellation, or browser rendering. Functional tests retain responsibility for event contents, immutable messages, tool semantics, fork lineage, and read-only versus writable side effects; endpoint counts prevent timing a skipped workload without duplicating those assertions.
+
+This note supplements, rather than supersedes, the Session-opening gate's isolation and calibration rationale. No existing active decision is retired.

+ 60 - 0
.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.zh.md

@@ -0,0 +1,60 @@
+# Agent Note: 工具密集后端续聊的性能基线
+
+Status: implemented
+
+[English](2026-09-06-backend-continuation-performance.md) | 中文
+
+## 问题
+
+打开一个 Session 不能衡量长工具对话后重复准备模型请求、执行更多工具密集轮次或发现多个非活动 fork 子会话的成本。[Session 打开门禁](2026-09-04-session-open-performance-gate.zh.md)覆盖首屏历史和激活,但有意停在新的模型工作开始前。它的文本与推理负载也不包含历史工具调用参数和大型工具结果。
+
+## 决定
+
+[agent-continuation 基准](../../../../benchmarks/agent-continuation/agent-continuation.bench.ts)增加三个场景组,包含一个已发布 profile 变体,不修改产品实现。它们通过生产追加、流累积和持久化 API 构造当前代际的 Zstandard Session。独立播种进程在测量前生成确定性源数据;每个样本将其复制到私有根目录,并启动新的已编译纯 Node worker。输入不来自录制 Session、环境仓库、网络、私有 Harness 主目录或已部署 GUI。
+
+共享历史包含 800 个已完成的双步骤轮次,每轮四次工具调用,工具结果为 2,048 字符:共 13,600 个事件和 5,600 条模型消息。每条助手回复携带推理、文本和紧凑流记录;请求工具的回复还携带分片参数。播种数据使用固定时间戳和 id。实时合成回复使用真实循环的时钟和 id,不覆盖进程全局状态。
+
+| 用例 | 计时操作 | 终点 |
+|---|---|---|
+| 请求历史 | 在不计时的冷恢复后,向工具密集历史顺序提交 40 个纯文本轮次,然后 flush | 空闲 Agent,已完成全部 40 次模型请求;分别报告轮次和最终 flush 时间 |
+| 工具续聊 | 冷恢复,顺序执行 20 个轮次,每轮八次可安全并行的合成工具调用和一条最终回复,然后 flush | 空闲 Agent,已完成 40 次模型请求和 160 次工具执行;分别报告恢复、轮次和最终 flush 时间 |
+| 已发布 SDK 工作流 | 使用 sdk-minimal profile 启动已构建 dsh,顺序提交 100 个轮次,每轮八次真实文件查看调用,然后关闭 SDK | SDK 收到 200 条助手消息和 800 个成功文件结果;包含 Loader 启动、stdio JSON-RPC、持久化和关闭 |
+| 子会话目录 | 通过真实 subagent 和 Session 查询服务,两次列出 16 个非活动、带种子的 fork 子会话 | 两份完整健康目录,观察已释放;每个子会话继承 80 个工具密集轮次,并在精确 fork 切点后拥有自己的描述符 |
+
+工具执行管线、请求准备、这些服务所需的 Session 投影、持久化和目录观察均保留生产代码。只有模型适配器和有界工具体是合成的。适配器只保留请求计数,不保留请求对象,因此 fixture(测试前置数据)不会制造不断增长的保留成本。顺序输入使每个空闲区间对应此 worker 提交的唯一请求;这不代表并发输入时可以把空闲状态推广为逐消息完成 API。
+
+五个样本报告原始壁钟时间、CPU 用户态/内核态时间、峰值 RSS、终点计数及总壁钟时间的最小值、中位数和最大值。预算约束未经舍入的中位数。续聊还相对已初始化 Host 测量保留堆内存:计时操作前后各执行两次显式 GC,中间让出一次事件循环,空闲 Agent 始终可达。因此该增量包含常驻历史 Session 和实时追加,而不只是新轮次。GC 与资源释放不计时;flush 计时。请求历史的内存基线从恢复后开始,只作诊断。目录峰值 RSS 仅作诊断;没有保留堆预算声称衡量已经释放的子会话观察。
+
+父进程为每个子进程设置 60 秒上限,独立检查超时、信号、退出状态和报告,等待进程关闭,并在失败后删除私有根目录。Context 和 Agent 在 finally 中释放。播种进程无法预热被测进程的缓存。不强制清除文件系统缓存:冷指新进程,不指冷物理存储。
+
+## 校准证据
+
+实现参考为 Apple M4 Pro、macOS arm64、Node 24.19.0 上的 `925e012340f033f0521e802ba8569ce6dd7ef1ac`。两轮独占的五样本运行使用相同播种数据,没有产品优化。下表时间单位为毫秒;源码期望值向上取整至实测各轮中位数以上,而不是施加尚未实现的优化目标。
+
+| 用例 | 第一轮原始总时间 | 第二轮原始总时间 | 中位数 | 参考期望 | CI 预算 |
+|---|---|---|---|---:|---:|
+| 请求历史 | 209.134, 210.333, 208.959, 236.355, 238.685 | 222.833, 213.911, 208.089, 211.494, 209.137 | 210.333 / 211.494 | 220 | 550 |
+| 工具续聊 | 358.953, 324.790, 318.861, 320.119, 322.896 | 324.280, 321.952, 340.409, 325.470, 324.312 | 322.896 / 324.312 | 340 | 850 |
+| 子会话目录 | 318.730, 309.006, 311.404, 308.565, 310.105 | 308.670, 310.030, 280.086, 303.084, 284.829 | 310.105 / 303.084 | 320 | 800 |
+
+续聊保留约 22.295 MiB;源码期望值为 23 MiB,预算为 28.75 MiB。时间期望值使用现有[校准辅助函数](../../../../benchmarks/support/calibration.ts):2× 共享 CI 时间比例和 1.25× 波动余量。内存只使用 1.25× 余量。比例继承现有通道的校准,并非这些用例的新 Linux 实测值;runner 特征变化时仍需 CI 证据。基线预算保护实测实现;更紧预算属于有测量依据且保持行为的修复。
+
+独立的纯 Node 请求历史 CPU profile 在一次 211.300 ms 操作中,将 132.876 ms 采样自身时间归因于 buildRequest 调用的 deepFreeze。这把重复遍历已冻结历史定位为聚焦调查目标,不是已证实的优化结果。目录首次/重复时间分别保留,因为观察释放后第二次列举仍读取带种子子会话的正文。
+
+已发布 SDK 变体完成 100 个轮次、200 次请求和 800 次真实文件读取。五样本 smoke 总时间为 1,521.773、1,463.465、1,689.701、1,365.485 和 1,417.106 ms(中位数 1,463.465 ms);完整套件重复运行报告 1,596.183、1,784.536、2,120.082、1,405.365 和 1,355.894 ms(中位数 1,596.183 ms)。1,700 ms 参考期望对应 4,250 ms CI 预算。重复运行中未改变的服务用例也变慢,因此这是可变主机负载下的验证,不是放宽其独占校准预算的依据。SDK 进程使用白名单环境和私有主目录/工作区。40 秒截止时间启动 SDK 关闭;所有路径等待同一个记忆化 close Promise,并早于外层 worker 的 60 秒截止时间。Profile 时间包含启动、全部轮次和关闭,分别报告;不把父进程 CPU 或堆指标当作服务端内存。适配器不为外部模型服务商序列化请求。
+
+## 考虑过的替代方案
+
+**重复现有迁移和首次打开变体。** 拒绝:现有十二个用例已经区分只读准备与可写发布。这些用例使用当前代际并开始或继续实际模型工作,或者列举语料集合而不是打开单个 Session。
+
+**只测 deriveMessages。** 拒绝:它的增量缓存不包含完整请求冻结、适配器分发、实时追加或持久化。实际顺序请求保护 Agent 每一步支付的成本。
+
+**只使用投影缓存行已预热的无种子子会话。** 拒绝:该路径绕过正文观察,遗漏 fork 子会话的精确继承切点要求。目录用例有意不挂载可选投影缓存,报告带种子的回退路径;它不代表缓存命中的发现过程。
+
+**将优化及其目标预算与首次测量一起应用。** 拒绝:纯基线层可以独立合并,并在归因或实现改变前记录当前负载。环境变量不能覆盖源码常量。
+
+## 后果
+
+通道增加三个场景组中的四个用例、二十个测量 worker 和两个播种进程。集成续聊用例覆盖恢复、完成模型/工具工作及持久化 flush。已发布 SDK 工作流额外包含 profile 启动、SDK 传输、真实文件工具和关闭;只有模型适配器是合成的。它创建新 Session,因为公共 SDK prompt API 创建而非恢复已存储身份。两条路径均不包含网络模型延迟、服务商专属请求序列化、可选用户插件、压缩、失败工具结果、图像、取消或浏览器渲染。功能测试仍负责事件内容、不可变消息、工具语义、fork 谱系以及只读/可写副作用;终点计数防止把跳过的工作当作测量结果,不重复这些断言。
+
+本记录补充而非取代 Session 打开门禁的隔离和校准依据。不退役任何现有活跃决策。

+ 6 - 0
benchmarks/agent-continuation/README.i18n.yaml

@@ -0,0 +1,6 @@
+# Bilingual-pair consistency record (docs/i18n/README.md): the git blob hash of each
+# side as of the last confirmed-consistent state. Both languages carry equal authority;
+# after editing either side, bring the other along and re-record with:
+#   pnpm run verify-translation-pairing --write benchmarks/agent-continuation/README.md
+README.md: 96f6915ebe2cfa3b334939e87d03828526026281
+README.zh.md: 5889e7b6d04c99f7eff388f2606c8f6bbb201c86

+ 31 - 0
benchmarks/agent-continuation/README.md

@@ -0,0 +1,31 @@
+# Backend continuation benchmarks
+
+English | [中文](README.zh.md)
+
+## Summary
+
+Measure long-history request processing, cold tool-heavy continuation, and repeated discovery of inactive fork children without network services or recorded user data. The SDK variant drives 100 turns and 800 real file reads through the shipped sdk-minimal profile; other cases isolate backend service costs. No case renders a browser.
+
+## Table of Contents
+
+- [Run](#run)
+- [Measurements](#measurements)
+- [Dev Note](#dev-note)
+
+<a id="run"></a>
+
+## Run
+
+From the repository root, build the libraries and workers with `pnpm run build:bench`, then run `pnpm exec vitest run --config vitest.bench.config.ts benchmarks/agent-continuation/agent-continuation.bench.ts`. Do not overlap timing runs with builds or other benchmarks.
+
+The test reports all five fresh-process samples and enforces reviewed median budgets. A failed worker reports its exit, signal, timeout, and stderr; temporary roots are removed even on failure. The required benchmark lane discovers this file automatically.
+
+<a id="measurements"></a>
+
+## Measurements
+
+[workload.ts](workload.ts) owns synthetic dimensions. [The Agent Note](../../.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.md) owns timing endpoints, calibration evidence, memory interpretation, and exclusions. The model adapter does not perform provider serialization or network calls; integrated cases use synthetic tool bodies, while the SDK profile variant performs real file reads.
+
+## Dev Note
+
+None.

+ 31 - 0
benchmarks/agent-continuation/README.zh.md

@@ -0,0 +1,31 @@
+# 后端续聊基准
+
+[English](README.md) | 中文
+
+## Summary
+
+在不使用网络服务或录制用户数据的情况下,测量长历史请求处理、冷工具密集续聊和重复发现非活动 fork 子会话。SDK 变体通过已发布 sdk-minimal profile 执行 100 个轮次和 800 次真实文件读取;其他用例隔离后端服务成本。所有用例均不渲染浏览器。
+
+## Table of Contents
+
+- [运行](#run)
+- [测量](#measurements)
+- [Dev Note](#dev-note)
+
+<a id="run"></a>
+
+## 运行
+
+在仓库根目录使用 `pnpm run build:bench` 构建库和 worker,然后运行 `pnpm exec vitest run --config vitest.bench.config.ts benchmarks/agent-continuation/agent-continuation.bench.ts`。不要让计时运行与构建或其他基准重叠。
+
+测试报告全部五个新进程样本,并约束经审查的中位数预算。worker 失败时报告退出状态、信号、超时和 stderr;失败时也会删除临时根目录。必需基准通道自动发现此文件。
+
+<a id="measurements"></a>
+
+## 测量
+
+[workload.ts](workload.ts)拥有合成维度。[Agent Note](../../.agents/notes/implemented/testing/2026-09-06-backend-continuation-performance.zh.md)拥有计时终点、校准证据、内存解释和排除项。模型适配器不执行服务商序列化或网络调用;集成用例的合成工具经过真实执行管线,SDK profile 变体则执行真实文件读取。
+
+## Dev Note
+
+无。

+ 86 - 0
benchmarks/agent-continuation/agent-continuation.bench.ts

@@ -0,0 +1,86 @@
+/** Baseline budgets for long-history requests, tool continuation, and fork-child discovery. */
+
+import { cp, mkdir, mkdtemp, rm } from 'node:fs/promises'
+import { tmpdir } from 'node:os'
+import { join } from 'node:path'
+import { afterAll, beforeAll, describe, expect, it } from 'vitest'
+import { runBuiltBenchmarkWorker } from '../support/built-worker.ts'
+import { ciTimeBudget, PERFORMANCE_BUDGET_HEADROOM } from '../support/calibration.ts'
+import type { ContinuationReport } from './agent-continuation.worker.ts'
+import type { CatalogReport } from './child-catalog.worker.ts'
+import type { ProfileReport } from './profile-continuation.worker.ts'
+import { WORKLOAD } from './workload.ts'
+
+const ATTEMPTS = 5
+const WORKER_TIMEOUT_MS = 60_000
+/** M4 Pro / Node 24.19 baseline expectations, before shared CI scaling and variance headroom. */
+const EXPECTED_MS = { 'request-history': 220, 'tool-continuation': 340, catalog: 320, 'profile-continuation': 1_700 } as const
+const EXPECTED_RETAINED_HEAP_MB = 23
+const WORKERS = join(import.meta.dirname, '..', '.dsh-build', 'agent-continuation')
+
+type Scenario = keyof typeof EXPECTED_MS
+type Report = ContinuationReport | CatalogReport | ProfileReport
+
+function workerName(scenario: Scenario): string {
+  if (scenario === 'profile-continuation') return 'profile-continuation.worker.js'
+  return scenario === 'catalog' ? 'child-catalog.worker.js' : 'agent-continuation.worker.js'
+}
+
+async function run<Output>(root: string, scenario: Scenario, mode: string): Promise<Output> {
+  const outcome = await runBuiltBenchmarkWorker<Output>({
+    worker: join(WORKERS, workerName(scenario)), args: [root, mode],
+    timeoutMs: WORKER_TIMEOUT_MS, exposeGc: true,
+  })
+  if (outcome.timedOut || outcome.signal !== null || outcome.exitCode !== 0 || outcome.report === undefined) {
+    throw new Error('backend worker failed: ' + JSON.stringify(outcome))
+  }
+  return outcome.report
+}
+
+function median(values: readonly number[]): number {
+  return [...values].sort((a, b) => a - b)[Math.floor(values.length / 2)] as number
+}
+
+describe('continuing tool-heavy Sessions with large histories', () => {
+  let scratch: string | undefined
+  const sources = new Map<Scenario, string>()
+
+  beforeAll(async () => {
+    scratch = await mkdtemp(join(tmpdir(), 'dsh-agent-continuation-bench-'))
+    for (const scenario of ['request-history', 'catalog'] as const) {
+      const root = join(scratch, 'source-' + scenario)
+      await run(root, scenario, 'seed')
+      sources.set(scenario, root)
+    }
+    sources.set('tool-continuation', sources.get('request-history') as string)
+  })
+  afterAll(async () => {
+    if (scratch !== undefined) await rm(scratch, { recursive: true, force: true })
+  })
+
+  for (const scenario of ['request-history', 'tool-continuation', 'catalog', 'profile-continuation'] as const) {
+    it(scenario, async () => {
+      const samples: Report[] = []
+      for (let attempt = 0; attempt < ATTEMPTS; attempt++) {
+        const root = join(scratch as string, scenario + '-' + String(attempt))
+        if (scenario === 'profile-continuation') await mkdir(root)
+        else await cp(sources.get(scenario) as string, root, { recursive: true })
+        try { samples.push(await run<Report>(root, scenario, scenario)) }
+        finally { await rm(root, { recursive: true, force: true }) }
+      }
+      const totalMs = samples.map(sample => sample.totalMs)
+      const budgetMs = ciTimeBudget(EXPECTED_MS[scenario])
+      const retainedHeapBudgetMb = EXPECTED_RETAINED_HEAP_MB * PERFORMANCE_BUDGET_HEADROOM
+      console.log(JSON.stringify({
+        benchmark: 'agent-continuation/' + scenario, workload: WORKLOAD,
+        samples, totalMs: { min: Math.min(...totalMs), median: median(totalMs), max: Math.max(...totalMs) },
+        budgetMs, ...(scenario === 'tool-continuation' ? { retainedHeapBudgetMb } : {}),
+      }))
+      expect(median(totalMs)).toBeLessThanOrEqual(budgetMs)
+      if (scenario === 'tool-continuation') {
+        expect(median((samples as ContinuationReport[]).map(sample => sample.retainedHeapMb)))
+          .toBeLessThanOrEqual(retainedHeapBudgetMb)
+      }
+    })
+  }
+})

+ 144 - 0
benchmarks/agent-continuation/agent-continuation.worker.ts

@@ -0,0 +1,144 @@
+/** Plain-Node measurements of active request history and cold tool-heavy continuation. */
+
+import { performance } from 'node:perf_hooks'
+import { scheduler } from 'node:timers/promises'
+import { Context } from '@deepseek-ai/cordis'
+import AgentLoop from '@deepseek-ai/dsh-agent-loop'
+import type { Agent, AgentHandle } from '@deepseek-ai/dsh-agent'
+import { mountAgentLoopTestDependencies } from '@deepseek-ai/dsh-agent-loop-testkit'
+import { createUserMessage, LlmAdapter } from '@deepseek-ai/dsh-llm'
+import type { GenerateOptions, LlmResolvedModelInfo, StreamChunk } from '@deepseek-ai/dsh-llm'
+import { SESSION_FORMAT_VERSION } from '@deepseek-ai/dsh-session'
+import JsonlSessionPersistence from '@deepseek-ai/dsh-session-persistence-jsonl'
+import SessionProjectionRegistry from '@deepseek-ai/dsh-session-projection'
+import { defineContentToolFixture } from '@deepseek-ai/dsh-tools'
+import { assertBuiltBenchmarkRuntime } from '../support/built-worker.ts'
+import { PARENT_ID, response, resultText, syntheticHistory, TIME_ZERO, WORKLOAD } from './workload.ts'
+
+/** Raw timing and retained-memory report from one isolated backend process. */
+export interface ContinuationReport {
+  readonly totalMs: number
+  readonly resumeMs: number
+  readonly turnsMs: number
+  readonly flushMs: number
+  readonly cpuUserMs: number
+  readonly cpuSystemMs: number
+  readonly retainedHeapMb: number
+  readonly peakRssMb: number
+  readonly requests: number
+  readonly toolCalls: number
+  readonly events: number
+}
+
+class SyntheticAdapter extends LlmAdapter {
+  requests = 0
+  constructor(private readonly toolsPerTurn: number) { super() }
+
+  override resolveModel(provider: string, model: string): Promise<LlmResolvedModelInfo> {
+    return Promise.resolve({ provider, id: model, name: model })
+  }
+
+  async * stream(_options: GenerateOptions): AsyncIterable<StreamChunk> {
+    const tools = this.toolsPerTurn > 0 && this.requests % 2 === 0 ? this.toolsPerTurn : 0
+    const reply = response(100_000 + this.requests++, tools)
+    yield* reply.chunks
+  }
+}
+
+async function collectHeap(): Promise<number> {
+  if (globalThis.gc === undefined) throw new Error('backend benchmark requires --expose-gc')
+  globalThis.gc()
+  await scheduler.yield()
+  globalThis.gc()
+  return process.memoryUsage().heapUsed / 1_048_576
+}
+
+async function seed(root: string): Promise<void> {
+  const ctx = new Context()
+  try {
+    await mountAgentLoopTestDependencies(ctx)
+    await ctx.plugin(JsonlSessionPersistence, { root, compression: 'zstd' })
+    const handle = await ctx.sessionPersistence.create({
+      version: SESSION_FORMAT_VERSION, id: PARENT_ID, createdAt: TIME_ZERO, cwd: '/bench', isSeeded: false,
+    }, {})
+    try {
+      await handle.append(syntheticHistory(WORKLOAD.historyTurns))
+      await handle.flush()
+    } finally { await handle.close() }
+  } finally { await ctx.fiber.dispose() }
+}
+
+async function runTurns(agent: Agent, turns: number): Promise<void> {
+  for (let turn = 0; turn < turns; turn++) {
+    agent.followup(createUserMessage({ content: [{ type: 'text', text: 'Continue synthetic task ' + String(turn) }], source: { kind: 'user' } }))
+    await agent.whenIdle()
+  }
+}
+
+async function measure(root: string, scenario: string): Promise<ContinuationReport> {
+  const ctx = new Context()
+  let handle: AgentHandle | undefined
+  const toolHeavy = scenario === 'tool-continuation'
+  const adapter = new SyntheticAdapter(toolHeavy ? WORKLOAD.toolsPerLiveTurn : 0)
+  let toolCalls = 0
+  try {
+    await ctx.plugin(SessionProjectionRegistry)
+    await mountAgentLoopTestDependencies(ctx)
+    await ctx.plugin(JsonlSessionPersistence, { root, compression: 'zstd' })
+    await ctx.plugin(AgentLoop, { agents: [] })
+    ctx.effect(() => ctx.llm.registerAdapter(['bench'], adapter))
+    ctx.effect(() => ctx.tools.register(defineContentToolFixture({
+      name: 'bench_tool', description: 'Read a bounded synthetic module.',
+      parameters: { ordinal: { type: 'number', required: true } },
+      isConcurrencySafe: () => true,
+      execute(args) {
+        toolCalls++
+        return Promise.resolve([{ type: 'text', text: resultText(args.ordinal) }])
+      },
+    })))
+    if (!toolHeavy) {
+      handle = await ctx.agents.resume({ resumeSessionId: PARENT_ID, agentOptions: { provider: 'bench', model: 'bench' } })
+    }
+    const beforeHeap = await collectHeap()
+    const cpuStart = process.cpuUsage()
+    const start = performance.now()
+    if (handle === undefined) {
+      handle = await ctx.agents.resume({ resumeSessionId: PARENT_ID, agentOptions: { provider: 'bench', model: 'bench' } })
+    }
+    const resumed = performance.now()
+    await runTurns(handle.agent, toolHeavy ? WORKLOAD.continuationTurns : WORKLOAD.requestTurns)
+    const turnsDone = performance.now()
+    await ctx.sessions.flush(handle.agent.session)
+    const end = performance.now()
+    const cpu = process.cpuUsage(cpuStart)
+    const retainedHeapMb = (await collectHeap()) - beforeHeap
+    if (adapter.requests !== (toolHeavy ? WORKLOAD.continuationTurns * 2 : WORKLOAD.requestTurns)
+      || toolCalls !== (toolHeavy ? WORKLOAD.continuationTurns * WORKLOAD.toolsPerLiveTurn : 0)) {
+      throw new Error('backend benchmark did not complete every requested model/tool step')
+    }
+    return {
+      totalMs: end - start, resumeMs: resumed - start, turnsMs: turnsDone - resumed, flushMs: end - turnsDone,
+      cpuUserMs: cpu.user / 1_000, cpuSystemMs: cpu.system / 1_000,
+      retainedHeapMb, peakRssMb: process.resourceUsage().maxRSS / 1_024,
+      requests: adapter.requests, toolCalls, events: handle.agent.session.seq,
+    }
+  } finally {
+    await handle?.dispose()
+    await ctx.fiber.dispose()
+  }
+}
+
+assertBuiltBenchmarkRuntime(import.meta.url, Object.fromEntries([
+  '@deepseek-ai/dsh-agent-loop', '@deepseek-ai/dsh-session', '@deepseek-ai/dsh-llm',
+  '@deepseek-ai/dsh-tools', '@deepseek-ai/dsh-session-persistence-jsonl',
+].map(name => [name, import.meta.resolve(name)])))
+const [root, scenario] = process.argv.slice(2)
+if (root === undefined || scenario === undefined || !['seed', 'request-history', 'tool-continuation'].includes(scenario)) {
+  throw new Error('usage: agent-continuation.worker.js <root> <seed|request-history|tool-continuation>')
+}
+if (scenario === 'seed') {
+  await seed(root)
+  process.stdout.write(JSON.stringify({ seeded: true }) + '\n')
+} else {
+  process.stdout.write(JSON.stringify(await measure(root, scenario)) + '\n')
+}

+ 95 - 0
benchmarks/agent-continuation/child-catalog.worker.ts

@@ -0,0 +1,95 @@
+/** Cold catalog observations of persisted fork children with tool-heavy inherited histories. */
+
+import { performance } from 'node:perf_hooks'
+import { Context } from '@deepseek-ai/cordis'
+import SessionStore, { SESSION_FORMAT_VERSION, SessionId, SessionLogOffset, SessionSeq } from '@deepseek-ai/dsh-session'
+import type { SessionEvent } from '@deepseek-ai/dsh-session'
+import JsonlSessionPersistence from '@deepseek-ai/dsh-session-persistence-jsonl'
+import SessionProjectionRegistry from '@deepseek-ai/dsh-session-projection'
+import SessionQueryEngine from '@deepseek-ai/dsh-session-query'
+import SubagentRuntime, { SUBAGENT_DESCRIPTOR_VERSION } from '@deepseek-ai/dsh-subagent'
+import { assertBuiltBenchmarkRuntime } from '../support/built-worker.ts'
+import { PARENT_ID, syntheticHistory, TIME_ZERO, WORKLOAD } from './workload.ts'
+
+/** Two complete catalog reads in one fresh Host, with every child observation released. */
+export interface CatalogReport {
+  readonly totalMs: number
+  readonly firstMs: number
+  readonly repeatMs: number
+  readonly cpuUserMs: number
+  readonly cpuSystemMs: number
+  readonly children: number
+  readonly peakRssMb: number
+}
+
+class CatalogQuery extends SessionQueryEngine {
+  override searchSessions(): Promise<never> {
+    return Promise.reject(new Error('search is outside the child-catalog benchmark'))
+  }
+  override searchEvents(): Promise<never> {
+    return Promise.reject(new Error('search is outside the child-catalog benchmark'))
+  }
+}
+
+async function seed(ctx: Context): Promise<void> {
+  const inherited = syntheticHistory(WORKLOAD.childHistoryTurns)
+  for (let child = 0; child < WORKLOAD.children; child++) {
+    const id = SessionId('bench-child-' + String(child))
+    const events: SessionEvent[] = [
+      ...inherited,
+      { type: 'session/end-seed', seq: SessionSeq(inherited.length), time: TIME_ZERO + inherited.length, data: { inherited: true } },
+      { type: 'subagent/descriptor', seq: SessionSeq(inherited.length + 1), time: TIME_ZERO + inherited.length + 1, data: {
+        version: SUBAGENT_DESCRIPTOR_VERSION, mode: 'continuable', provider: 'fork', label: 'Synthetic child ' + String(child),
+      } },
+    ]
+    const handle = await ctx.sessionPersistence.create({
+      version: SESSION_FORMAT_VERSION, id, createdAt: TIME_ZERO + child, cwd: '/bench',
+      parentSession: PARENT_ID, isSeeded: true, origin: 'subagent', delegationDepth: 1,
+    }, { inheritedEventCount: SessionLogOffset(inherited.length) })
+    try {
+      await handle.append(events)
+      await handle.flush()
+    } finally { await handle.close() }
+  }
+}
+
+async function run(root: string, mode: string): Promise<CatalogReport | { seeded: true }> {
+  const ctx = new Context()
+  try {
+    await ctx.plugin(SessionStore)
+    await ctx.plugin(SessionProjectionRegistry)
+    await ctx.plugin(JsonlSessionPersistence, { root, compression: 'zstd' })
+    await ctx.plugin(CatalogQuery)
+    await ctx.plugin(SubagentRuntime)
+    if (mode === 'seed') {
+      await seed(ctx)
+      return { seeded: true }
+    }
+    const cpuStart = process.cpuUsage()
+    const start = performance.now()
+    const first = await ctx.subagents.listChildren(PARENT_ID)
+    const firstDone = performance.now()
+    const repeated = await ctx.subagents.listChildren(PARENT_ID)
+    const end = performance.now()
+    const cpu = process.cpuUsage(cpuStart)
+    if (first.length !== WORKLOAD.children || repeated.length !== WORKLOAD.children
+      || [...first, ...repeated].some(row => row.kind !== 'child')) {
+      throw new Error('child-catalog benchmark did not reach the complete healthy catalog')
+    }
+    return {
+      totalMs: end - start, firstMs: firstDone - start, repeatMs: end - firstDone,
+      cpuUserMs: cpu.user / 1_000, cpuSystemMs: cpu.system / 1_000,
+      children: first.length, peakRssMb: process.resourceUsage().maxRSS / 1_024,
+    }
+  } finally { await ctx.fiber.dispose() }
+}
+
+assertBuiltBenchmarkRuntime(import.meta.url, Object.fromEntries([
+  '@deepseek-ai/dsh-subagent', '@deepseek-ai/dsh-session-query',
+  '@deepseek-ai/dsh-session-persistence-jsonl',
+].map(name => [name, import.meta.resolve(name)])))
+const [root, mode] = process.argv.slice(2)
+if (root === undefined || (mode !== 'seed' && mode !== 'catalog')) {
+  throw new Error('usage: child-catalog.worker.js <root> <seed|catalog>')
+}
+process.stdout.write(JSON.stringify(await run(root, mode)) + '\n')

+ 42 - 0
benchmarks/agent-continuation/profile-adapter.ts

@@ -0,0 +1,42 @@
+/** Compiled synthetic model for the shipped sdk-minimal profile; tools remain production plugins. */
+
+import type { Context } from '@deepseek-ai/cordis'
+import { LlmAdapter, ToolCallId } from '@deepseek-ai/dsh-llm'
+import type { GenerateOptions, LlmResolvedModelInfo, StreamChunk } from '@deepseek-ai/dsh-llm'
+import { response, WORKLOAD } from './workload.ts'
+
+class ProfileAdapter extends LlmAdapter {
+  private requests = 0
+  override resolveModel(provider: string, model: string): Promise<LlmResolvedModelInfo> {
+    return Promise.resolve({ provider, id: model, name: model, contextWindow: 1_000_000 })
+  }
+
+  async * stream(_options: GenerateOptions): AsyncIterable<StreamChunk> {
+    const serial = this.requests++
+    if (serial % 2 === 1) {
+      yield* response(200_000 + serial, 0).chunks
+      return
+    }
+    for (let index = 0; index < WORKLOAD.toolsPerLiveTurn; index++) {
+      const id = ToolCallId('profile-call-' + String(serial) + '-' + String(index))
+      const args = JSON.stringify({ command: 'view', path: process.cwd() + '/synthetic.txt' })
+      yield { type: 'block-start', index, blockType: 'tool-call' }
+      yield { type: 'tool-call-delta', index, id, name: 'str_replace_editor', argumentsDelta: args }
+      yield { type: 'block-end', index, block: { type: 'tool-call', id, name: 'str_replace_editor', arguments: args } }
+    }
+    yield { type: 'finish', reason: { kind: 'tool-calls' } }
+  }
+}
+
+/** Loader plugin identity. */
+export const name = 'backend-profile-benchmark-model'
+/** The scripted provider requires the production LLM registry. */
+export const inject = ['llm']
+
+/**
+ * Register the synthetic provider without changing any runtime services or tools.
+ * @param ctx - profile-owned plugin context.
+ */
+export function apply(ctx: Context): void {
+  ctx.effect(() => ctx.llm.registerAdapter(['bench'], new ProfileAdapter()))
+}

+ 85 - 0
benchmarks/agent-continuation/profile-continuation.worker.ts

@@ -0,0 +1,85 @@
+/** End-to-end SDK continuation through the built dsh sdk-minimal profile and real file tools. */
+
+import { mkdir, writeFile } from 'node:fs/promises'
+import { join } from 'node:path'
+import { performance } from 'node:perf_hooks'
+import { DeepSeekHarness } from '@deepseek-ai/dsh-sdk-client'
+import { assertBuiltBenchmarkRuntime } from '../support/built-worker.ts'
+import { PARENT_ID, resultText, WORKLOAD } from './workload.ts'
+
+/** Parent-observed wall time, including profile launch and SDK shutdown. */
+export interface ProfileReport {
+  readonly totalMs: number
+  readonly bootMs: number
+  readonly turnsMs: number
+  readonly closeMs: number
+  readonly requests: number
+  readonly toolCalls: number
+}
+
+async function run(root: string): Promise<ProfileReport> {
+  const home = join(root, 'home')
+  const cwd = join(root, 'workspace')
+  await mkdir(cwd, { recursive: true })
+  await mkdir(home, { recursive: true })
+  await writeFile(join(cwd, 'synthetic.txt'), resultText(0))
+  const patch = join(root, 'profile.patch.yml')
+  await writeFile(patch, [
+    '- id: llm-deepseek', '  disabled: true',
+    '- id: sessions', '  config:', '    root: ' + JSON.stringify(join(root, 'profile-sessions')), '    compression: zstd',
+    '- insert:', '    - id: benchmark-model', '      name: ' + JSON.stringify(join(import.meta.dirname, 'profile-adapter.js')),
+    '',
+  ].join('\n'))
+  const env: NodeJS.ProcessEnv = {
+    PATH: process.env.PATH, HOME: home, USERPROFILE: home,
+    DSH_AGENTS_HOME: join(home, 'agents'),
+  }
+  const harness = new DeepSeekHarness({
+    dshBin: join(import.meta.dirname, '..', '..', '..', 'apps', 'cli', 'lib', 'bin.js'),
+    profile: 'sdk-minimal', dshHome: home, processCwd: cwd, cwd,
+    provider: 'bench', model: 'bench', patches: [patch], env,
+    initializeTimeoutMs: 15_000, requestTimeoutMs: 15_000,
+  })
+  let closing: Promise<void> | undefined
+  const close = (): Promise<void> => closing ??= harness.close()
+  let expired = false
+  const deadline = setTimeout(() => {
+    expired = true
+    // The awaited finally close below reports shutdown failures; this only requests cancellation.
+    void close().catch(() => undefined)
+  }, 40_000)
+  let requests = 0
+  let toolCalls = 0
+  const start = performance.now()
+  try {
+    await harness.start()
+    const booted = performance.now()
+    for (let turn = 0; turn < WORKLOAD.profileTurns; turn++) {
+      const result = await harness.run('Read the synthetic file ' + String(turn), { sessionId: PARENT_ID })
+      requests += result.events.filter(event => event.type === 'assistant/message').length
+      for (const event of result.events) {
+        if (event.type !== 'tool/result') continue
+        const result = event.data.message.content[0]
+        if (result.isError || !result.content.some(block => block.type === 'text' && block.text.includes('export const synthetic = 42;'))) {
+          throw new Error('profile benchmark did not read the synthetic file')
+        }
+        toolCalls++
+      }
+    }
+    const turnsDone = performance.now()
+    await close()
+    const end = performance.now()
+    if (expired || requests !== WORKLOAD.profileTurns * 2 || toolCalls !== WORKLOAD.profileTurns * WORKLOAD.toolsPerLiveTurn) {
+      throw new Error('profile benchmark did not finish every model request and real tool call')
+    }
+    return { totalMs: end - start, bootMs: booted - start, turnsMs: turnsDone - booted, closeMs: end - turnsDone, requests, toolCalls }
+  } finally {
+    clearTimeout(deadline)
+    await close()
+  }
+}
+
+assertBuiltBenchmarkRuntime(import.meta.url, { '@deepseek-ai/dsh-sdk-client': import.meta.resolve('@deepseek-ai/dsh-sdk-client') })
+const [root] = process.argv.slice(2)
+if (root === undefined) throw new Error('usage: profile-continuation.worker.js <root>')
+process.stdout.write(JSON.stringify(await run(root)) + '\n')

+ 110 - 0
benchmarks/agent-continuation/workload.ts

@@ -0,0 +1,110 @@
+/** Reviewed synthetic tool history shared by continuation and child-catalog measurements. */
+
+import { AssistantStreamAccumulator } from '@deepseek-ai/dsh-llm/assistant-stream'
+import { MessageId, ToolCallId } from '@deepseek-ai/dsh-llm'
+import type { ContentBlock, StreamChunk } from '@deepseek-ai/dsh-llm'
+import { Session, SessionId } from '@deepseek-ai/dsh-session'
+import type { SessionEvent } from '@deepseek-ai/dsh-session'
+
+/** Workload dimensions, independent of environment and recorded user material. */
+export const WORKLOAD = {
+  historyTurns: 800,
+  toolsPerHistoricalTurn: 4,
+  toolResultChars: 2_048,
+  requestTurns: 40,
+  continuationTurns: 20,
+  profileTurns: 100,
+  toolsPerLiveTurn: 8,
+  children: 16,
+  childHistoryTurns: 80,
+} as const
+
+/** Fixed clock used only to author persisted synthetic input. */
+export const TIME_ZERO = 1_700_000_000_000
+/** Durable parent identity of the measured continuation. */
+export const PARENT_ID = SessionId('bench-parent')
+
+/**
+ * Construct a deterministic model reply without retaining past requests.
+ * @param serial - unique response ordinal.
+ * @param tools - number of synthetic tool calls, or zero for a final text reply.
+ * @returns streamed chunks and their known final blocks.
+ */
+export function response(serial: number, tools: number): { chunks: StreamChunk[]; content: ContentBlock[] } {
+  const content: ContentBlock[] = [
+    { type: 'reasoning', text: 'Inspect the synthetic result. '.repeat(8) },
+    { type: 'text', text: 'Synthetic response. '.repeat(8) },
+    ...Array.from({ length: tools }, (_, index): ContentBlock => ({
+      type: 'tool-call', id: ToolCallId('call-' + String(serial) + '-' + String(index)), name: 'bench_tool',
+      arguments: JSON.stringify({ ordinal: serial * 100 + index }),
+    })),
+  ]
+  const chunks: StreamChunk[] = []
+  content.forEach((block, index) => {
+    chunks.push({ type: 'block-start', index, blockType: block.type })
+    if (block.type === 'text' || block.type === 'reasoning') {
+      for (let offset = 0; offset < block.text.length; offset += 16) {
+        chunks.push({ type: block.type === 'text' ? 'text-delta' : 'reasoning-delta', index, text: block.text.slice(offset, offset + 16) })
+      }
+    } else if (block.type === 'tool-call') {
+      for (let offset = 0; offset < block.arguments.length; offset += 8) {
+        chunks.push({ type: 'tool-call-delta', index, id: block.id, name: block.name, argumentsDelta: block.arguments.slice(offset, offset + 8) })
+      }
+    }
+    chunks.push({ type: 'block-end', index, block })
+  })
+  chunks.push({ type: 'usage', usage: { inputTokens: 10_000, outputTokens: 100 } })
+  chunks.push({ type: 'finish', reason: { kind: tools === 0 ? 'stop' : 'tool-calls' } })
+  return { chunks, content }
+}
+
+/**
+ * Author completed two-step turns through production Session append and stream compaction.
+ * @param turns - completed historical turns.
+ * @returns detached current-generation events with fixed ids, timestamps and payloads.
+ */
+export function syntheticHistory(turns: number): SessionEvent[] {
+  const session = Session.create(PARENT_ID)
+  for (let turn = 1; turn <= turns; turn++) {
+    session.append('turn/start', { turn })
+    session.append('step/start', { turn, step: 1 })
+    session.append('user/message', {
+      id: MessageId('prompt-' + String(turn)), role: 'user',
+      content: [{ type: 'text', text: 'Inspect synthetic module ' + String(turn) }], source: { kind: 'user' },
+    }, { surfaceOp: 'append' })
+    for (const step of [1, 2]) {
+      if (step === 2) session.append('step/start', { turn, step })
+      const reply = response(turn * 2 + step, step === 1 ? WORKLOAD.toolsPerHistoricalTurn : 0)
+      const stream = new AssistantStreamAccumulator()
+      reply.chunks.forEach((chunk, index) => { stream.push({ time: TIME_ZERO + turn * 1_000 + step * 100 + index, chunk }) })
+      session.append('assistant/message', {
+        turn, step,
+        message: { id: MessageId('reply-' + String(turn) + '-' + String(step)), role: 'assistant', content: reply.content, source: { kind: 'model', provider: 'bench', model: 'bench' } },
+        stream: [...stream.snapshot()],
+      }, { surfaceOp: 'append' })
+      for (const block of reply.content) {
+        if (block.type !== 'tool-call') continue
+        const call = session.append('tool/call', { turn, step, callId: block.id, name: block.name, arguments: block.arguments })
+        session.append('tool/result', {
+          turn, step,
+          message: {
+            id: MessageId('result-' + block.id), role: 'user', source: { kind: 'tool', callId: block.id },
+            content: [{ type: 'tool-result', toolCallId: block.id, content: [{ type: 'text', text: resultText(turn) }], isError: false }],
+          },
+        }, { surfaceOp: 'append', sourceEventSeqs: [call.seq] })
+      }
+      session.append('step/end', { turn, step })
+    }
+    session.append('turn/end', { turn, reason: { kind: 'completed' } })
+  }
+  return session.snapshotEvents().map(event => ({ ...event, time: TIME_ZERO + event.seq }))
+}
+
+/**
+ * Build a bounded synthetic file-read result with a varying prefix.
+ * @param ordinal - deterministic result identifier.
+ * @returns exactly the reviewed number of UTF-16 characters.
+ */
+export function resultText(ordinal: number): string {
+  return ('module ' + String(ordinal) + '\n' + 'export const synthetic = 42;\n'.repeat(100)).slice(0, WORKLOAD.toolResultChars)
+}

+ 3 - 0
benchmarks/package.json

@@ -15,6 +15,7 @@
     "@deepseek-ai/dsh-client-ui-chat": "workspace:^",
     "@deepseek-ai/dsh-deque": "workspace:^",
     "@deepseek-ai/dsh-llm": "workspace:^",
+    "@deepseek-ai/dsh-sdk-client": "workspace:^",
     "@deepseek-ai/dsh-session": "workspace:^",
     "@deepseek-ai/dsh-session-persistence": "workspace:^",
     "@deepseek-ai/dsh-session-persistence-jsonl": "workspace:^",
@@ -23,6 +24,8 @@
     "@deepseek-ai/dsh-session-stats": "workspace:^",
     "@deepseek-ai/dsh-session-title": "workspace:^",
     "@deepseek-ai/dsh-session-turn-outline": "workspace:^",
+    "@deepseek-ai/dsh-subagent": "workspace:^",
+    "@deepseek-ai/dsh-tools": "workspace:^",
     "@deepseek-ai/dsh-token-meter": "workspace:^",
     "@deepseek-ai/dsh-typert-protocol": "workspace:^"
   },

+ 12 - 0
benchmarks/tsdown.config.ts

@@ -14,6 +14,18 @@ const shared = {
 
 /** Compile measured benchmark workers while keeping workspace packages on their built `lib` entries. */
 export default defineConfig([
+  {
+    ...shared,
+    entry: {
+      'agent-continuation.worker': 'agent-continuation/agent-continuation.worker.ts',
+      'child-catalog.worker': 'agent-continuation/child-catalog.worker.ts',
+      'profile-continuation.worker': 'agent-continuation/profile-continuation.worker.ts',
+      'profile-adapter': 'agent-continuation/profile-adapter.ts',
+    },
+    outDir: '.dsh-build/agent-continuation',
+    clean: true,
+    tsconfig: 'tsconfig.host.json',
+  },
   {
     ...shared,
     entry: { 'session-open.worker': 'session-open/session-open.worker.ts' },

+ 9 - 0
pnpm-lock.yaml

@@ -581,6 +581,9 @@ importers:
       '@deepseek-ai/dsh-llm':
         specifier: workspace:^
         version: link:../packages/llm/llm
+      '@deepseek-ai/dsh-sdk-client':
+        specifier: workspace:^
+        version: link:../packages/sdk/client
       '@deepseek-ai/dsh-session':
         specifier: workspace:^
         version: link:../packages/core/session
@@ -605,9 +608,15 @@ importers:
       '@deepseek-ai/dsh-session-turn-outline':
         specifier: workspace:^
         version: link:../packages/session/session-turn-outline
+      '@deepseek-ai/dsh-subagent':
+        specifier: workspace:^
+        version: link:../packages/subagent/subagent
       '@deepseek-ai/dsh-token-meter':
         specifier: workspace:^
         version: link:../packages/llm/token-meter
+      '@deepseek-ai/dsh-tools':
+        specifier: workspace:^
+        version: link:../packages/core/tools
       '@deepseek-ai/dsh-typert-protocol':
         specifier: workspace:^
         version: link:../packages/typert/protocol