Browse Source

test(perf): calibrate request history on standard hosted CI

Tianyi Cui 1 month ago
parent
commit
8e270960ed

+ 2 - 2
.agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-provenance.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-provenance.md
-2026-09-06-agent-request-freeze-provenance.md: bfefe39a0c481250d45318c02199cbd947c4eb4f
-2026-09-06-agent-request-freeze-provenance.zh.md: 4a333c845d11e2e1cbe8ef2325f92f56f48b203d
+2026-09-06-agent-request-freeze-provenance.md: 7a4816df61f6490647aba6f0603719e1b4662a20
+2026-09-06-agent-request-freeze-provenance.zh.md: 239d7e69df1596010ef0f3c8789250f654a75cb1

+ 11 - 1
.agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-provenance.md

@@ -26,7 +26,7 @@ Apple M4 Pro, macOS arm64, Node 24.19.0; independent worktree dependencies and b
 | Original, 07:17:06–07:17:10 | 249.050708, 238.275291, 242.172084, 250.093166, 246.130875 | 246.130875 | Fail |
 | Optimized repeat, 07:18:17–07:18:20 | 66.693500, 67.402083, 68.665000, 66.642083, 66.609125 | 66.693500 | Pass |
 
-The same 800-turn, four-tools-per-historical-turn history and 40 live requests complete in every sample: 13,923 events, no live tool calls. The repeat median is 72.9% below the isolated original. A 70 ms source expectation rounds above both optimized medians; the existing 2× CI scale and 1.25× headroom produce 175 ms. This is local calibration, not proof that the shared scale fits every CI runner; the required CI lane owns runner validation. No other case or memory budget changes here.
+The same 800-turn, four-tools-per-historical-turn history and 40 live requests complete in every sample: 13,923 events, no live tool calls. The repeat median is 72.9% below the isolated original. The historical 70 ms M4 expectation rounds above both optimized medians; applying the shared 2× CI scale and 1.25× headroom produced the 175 ms budget used in the table. These remain local reference measurements, not hosted-runner expectations. The explicit hosted calibration below owns the enforced request-history budget; no other case or memory budget changes here.
 
 The first optimized slot also measures cold tool continuation: totals 185.839958, 185.235583, 185.865917, 189.213459, 185.279417; median 185.839958 ms. Every sample completes 40 requests and 160 tool calls with 14,143 events. Retained heap samples are 22.591591, 22.590355, 22.594795, 22.591743, 22.594681 MiB, below the unchanged 28.75 MiB budget. The earlier baseline's approximately 22.295 MiB highlights the small provenance-table cost; weak keys prevent the table itself retaining replaced messages.
 
@@ -34,6 +34,16 @@ The same slot's shipped SDK profile completes 100 turns, 200 requests, and 800 r
 
 An earlier original-code run at 06:58:28 UTC overlaps a sibling build because of scheduling-message latency: totals 264.269792, 282.442000, 365.836334, 293.172791, 288.719500 ms; median 288.719500 ms. It also fails 175 ms but is not calibration evidence. The isolated original row replaces that comparison, without removing or averaging away the contaminated samples.
 
+### Standard hosted CI calibration
+
+The standard two-CPU `ubuntu-24.04` lane runs Node 24.20.0. [Run 34033336380, job 101487280801](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336380/job/101487280801) measures the optimized request path at merge commit `8fba64d9ae06d1a9a778a95487bb915d24cb0644` in Azure eastus: 183.355397, 184.468253, 185.042397, 182.160790, 182.924728 ms; median 183.355397 ms. Every sample completes the same 40 requests and 13,923 events. All five exceed the historical 175 ms budget without changing the WeakSet implementation or workload.
+
+A second hosted run of the same request implementation, [run 34033336246, job 101487216170](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336246/job/101487216170), records 145.644577, 144.204300, 143.072572, 145.985903, 146.834474 ms; median 145.644577 ms. It uses the same Ubuntu image and Node version but a different worker in Azure westus3 at merge commit `c366e49`. This faster run does not replace the eastus evidence or establish why the workers differ. The older self-hosted `VM-7-113-ubuntu-ci-9` run with Node 24.18.1 ([run 34021903421, job 101456015028](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34021903421/job/101456015028)) records 110.025154, 119.958978, 108.266860, 107.557950, 108.538902 ms; median 108.538902 ms. Its runner and Node version do not calibrate the standard hosted lane.
+
+The request-history CI expectation is 190 ms, rounded above this observed range. The enforced median budget is `ceil(190 × 1.25) = 238 ms`; the shared 2× reference-machine scale does not apply again to a CI measurement. This matches the direct-CI calibration method of the [63 ms Session-reopen budget](../../../../benchmarks/session-open/session-open.bench.ts), rather than relabeling the M4 reference as hosted evidence. The 238 ms budget remains below the isolated original implementation’s 246.130875 ms M4 median.
+
+Deterministic controls call the same `assertRequestHistoryBudget` assertion as the timed case. They accept the recorded hosted median and maximum (185.042397 ms), reject the recorded original M4 median, and reject a synthetic 250 ms median from 248, 250, 252, 251, 249 ms inputs. The synthetic inputs model a material regression; they are not runtime measurements. Replaying recorded values verifies the assertion, not a new hosted run. The acceptance control fails at 175 ms before calibration; all three controls and the five request-freeze behavior tests pass at 238 ms.
+
 ## Alternatives considered
 
 **Return immediately for `Object.isFrozen`.** A frozen root does not prove its descendants frozen. Applying this shortcut to the shared helper would weaken every caller, including restore and projection paths.

+ 11 - 1
.agents/notes/implemented/simplification/2026-09-06-agent-request-freeze-provenance.zh.md

@@ -26,7 +26,7 @@ Apple M4 Pro、macOS arm64、Node 24.19.0;worktree 使用独立依赖和构建
 | 原版,07:17:06–07:17:10 | 249.050708, 238.275291, 242.172084, 250.093166, 246.130875 | 246.130875 | 失败 |
 | 优化版复测,07:18:17–07:18:20 | 66.693500, 67.402083, 68.665000, 66.642083, 66.609125 | 66.693500 | 通过 |
 
-每个样本都完成相同的 800 轮历史(每个历史轮次四个工具)和 40 个实时请求:13,923 个事件,无实时工具调用。复测中位数比独占原版低 72.9%。70 ms 的源码期望值向上取整并高于两次优化版中位数;现有 2× CI 系数和 1.25× 余量得到 175 ms。这是本地校准,不能证明共享系数适合所有 CI 运行器;必跑 CI 测试负责验证运行器。本文不改变其他场景或内存预算。
+每个样本都完成相同的 800 轮历史(每个历史轮次四个工具)和 40 个实时请求:13,923 个事件,无实时工具调用。复测中位数比独占原版低 72.9%。历史 M4 期望值 70 ms 向上取整并高于两次优化版中位数;应用共享 2× CI 系数和 1.25× 余量,得到表中使用的 175 ms 预算。这些仍是本地参考测量,而非托管运行器期望值。下文的显式托管校准拥有实际执行的请求历史预算;本文不改变其他场景或内存预算。
 
 首个优化版时段还测量冷启动工具续跑:总耗时 185.839958, 185.235583, 185.865917, 189.213459, 185.279417;中位数 185.839958 ms。每个样本都完成 40 个请求、160 个工具调用和 14,143 个事件。保留堆样本为 22.591591, 22.590355, 22.594795, 22.591743, 22.594681 MiB,低于不变的 28.75 MiB 预算。先前基线约 22.295 MiB,显示了证明表的小额成本;弱键防止表本身保留已替换消息。
 
@@ -34,6 +34,16 @@ Apple M4 Pro、macOS arm64、Node 24.19.0;worktree 使用独立依赖和构建
 
 较早的原版运行始于 06:58:28 UTC,因调度消息延迟而与其他构建重叠:总耗时 264.269792, 282.442000, 365.836334, 293.172791, 288.719500 ms;中位数 288.719500 ms。它也超过 175 ms,但不属于校准证据。独占原版行替代该比较,没有删除受污染样本或通过取平均掩盖它们。
 
+### 标准托管 CI 校准
+
+标准双 CPU `ubuntu-24.04` 测试通道运行 Node 24.20.0。[运行 34033336380、任务 101487280801](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336380/job/101487280801)在 Azure eastus 上测量合并提交 `8fba64d9ae06d1a9a778a95487bb915d24cb0644` 的优化请求路径:183.355397, 184.468253, 185.042397, 182.160790, 182.924728 ms;中位数 183.355397 ms。每个样本都完成相同的 40 个请求和 13,923 个事件。在 WeakSet 实现与工作负载未变的情况下,全部五个样本均超过历史 175 ms 预算。
+
+相同请求实现的另一次托管运行,[运行 34033336246、任务 101487216170](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34033336246/job/101487216170),记录了 145.644577, 144.204300, 143.072572, 145.985903, 146.834474 ms;中位数 145.644577 ms。它在合并提交 `c366e49` 上使用相同的 Ubuntu 镜像和 Node 版本,但运行于 Azure westus3 的另一台工作机。较快的运行不能替代 eastus 证据,也不能证明工作机差异的原因。较早的自托管 `VM-7-113-ubuntu-ci-9` 运行使用 Node 24.18.1([运行 34021903421、任务 101456015028](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34021903421/job/101456015028)),记录了 110.025154, 119.958978, 108.266860, 107.557950, 108.538902 ms;中位数 108.538902 ms。其运行器和 Node 版本不能校准标准托管通道。
+
+请求历史的 CI 期望值为 190 ms,向上取整并高于该观测范围。实际执行的中位数预算为 `ceil(190 × 1.25) = 238 ms`;CI 测量不再应用共享的参考机器 2× 系数。这与 [Session 重开 63 ms 预算](../../../../benchmarks/session-open/session-open.bench.ts)的直接 CI 校准方法一致,而非将 M4 参考值重新标注为托管证据。238 ms 预算仍低于独占原版实现的 M4 中位数 246.130875 ms。
+
+确定性对照调用与计时场景相同的 `assertRequestHistoryBudget` 断言。它们接受已记录的托管中位数和最大值(185.042397 ms),拒绝已记录的原版 M4 中位数,并拒绝由 248, 250, 252, 251, 249 ms 输入得到的合成 250 ms 中位数。合成输入模拟显著回归,并非运行时测量。回放已记录数值验证的是断言,而非新的托管运行。接受对照在校准前以 175 ms 预算失败;三个对照和五个请求冻结行为测试在 238 ms 预算下均通过。
+
 ## 考虑过的替代方案
 
 **`Object.isFrozen` 为真时立即返回。** 已冻结根对象不能证明其后代已冻结。在共享辅助函数中使用此捷径会削弱所有调用方,包括恢复与投影路径。

+ 37 - 5
benchmarks/agent-continuation/agent-continuation.bench.ts

@@ -24,10 +24,13 @@ const TOOL_CONTINUATION_BUDGET_MS = Math.ceil(EXPECTED_TOOL_CONTINUATION_CI_MS *
 /** Standard two-CPU hosted CI catalog median is 858.364 ms; 900 ms is the rounded expectation. */
 const EXPECTED_CATALOG_CI_MS = 900
 const CATALOG_BUDGET_MS = Math.ceil(EXPECTED_CATALOG_CI_MS * PERFORMANCE_BUDGET_HEADROOM)
+/** Two-CPU ubuntu-24.04 / Node 24.20 samples span 182.161–185.042 ms; rounded CI expectation. */
+const EXPECTED_REQUEST_HISTORY_CI_MS = 190
+const REQUEST_HISTORY_BUDGET_MS = Math.ceil(EXPECTED_REQUEST_HISTORY_CI_MS * PERFORMANCE_BUDGET_HEADROOM)
 const EXPECTED_RETAINED_HEAP_MB = 23
 const WORKERS = join(import.meta.dirname, '..', '.dsh-build', 'agent-continuation')
 
-type Scenario = keyof typeof EXPECTED_MS | 'catalog' | 'tool-continuation' | 'request-history'
+type Scenario = 'request-history' | 'catalog' | 'tool-continuation' | keyof typeof EXPECTED_MS
 type Report = ContinuationReport | CatalogReport | ProfileReport
 
 function workerName(scenario: Scenario): string {
@@ -96,6 +99,34 @@ describe('standard hosted baseline request-history calibration', () => {
   })
 })
 
+function assertRequestHistoryBudget(value: number): void {
+  expect(value).toBeLessThanOrEqual(REQUEST_HISTORY_BUDGET_MS)
+}
+
+describe('standard hosted request-history calibration', () => {
+  it('accepts the recorded two-CPU samples above the historical budget', () => {
+    const recorded = [183.355397, 184.468253, 185.042397, 182.160790, 182.924728]
+    const recordedMedian = median(recorded)
+
+    expect(recordedMedian).toBe(183.355397)
+    expect(recordedMedian).toBeGreaterThan(ciTimeBudget(70))
+    assertRequestHistoryBudget(recordedMedian)
+    assertRequestHistoryBudget(Math.max(...recorded))
+    expect(REQUEST_HISTORY_BUDGET_MS).toBe(238)
+  })
+
+  it('rejects a synthetic material request-history regression', () => {
+    const regressionMedian = median([248, 250, 252, 251, 249])
+    expect(() => assertRequestHistoryBudget(regressionMedian)).toThrow()
+  })
+
+  it('rejects the recorded original implementation on the M4 reference', () => {
+    const originalMedian = median([249.050708, 238.275291, 242.172084, 250.093166, 246.130875])
+    expect(originalMedian).toBe(246.130875)
+    expect(() => assertRequestHistoryBudget(originalMedian)).toThrow()
+  })
+})
+
 describe('continuing tool-heavy Sessions with large histories', () => {
   let scratch: string | undefined
   const sources = new Map<Scenario, string>()
@@ -124,16 +155,17 @@ describe('continuing tool-heavy Sessions with large histories', () => {
         finally { await rm(root, { recursive: true, force: true }) }
       }
       const totalMs = samples.map(sample => sample.totalMs)
-      const budgetMs = scenario === 'catalog' ? CATALOG_BUDGET_MS
-        : scenario === 'tool-continuation' ? TOOL_CONTINUATION_BUDGET_MS
-          : scenario === 'request-history' ? BASELINE_REQUEST_BUDGET_MS : ciTimeBudget(EXPECTED_MS[scenario])
+      const budgetMs = scenario === 'request-history' ? REQUEST_HISTORY_BUDGET_MS
+        : scenario === 'catalog' ? CATALOG_BUDGET_MS
+          : scenario === 'tool-continuation' ? TOOL_CONTINUATION_BUDGET_MS : ciTimeBudget(EXPECTED_MS[scenario])
       const retainedHeapBudgetMb = EXPECTED_RETAINED_HEAP_MB * PERFORMANCE_BUDGET_HEADROOM
       console.log(JSON.stringify({
         benchmark: 'agent-continuation/' + scenario, workload: WORKLOAD,
         samples, totalMs: { min: Math.min(...totalMs), median: median(totalMs), max: Math.max(...totalMs) },
         budgetMs, ...(scenario === 'tool-continuation' ? { retainedHeapBudgetMb } : {}),
       }))
-      expectTotalWithinBudget(median(totalMs), budgetMs)
+      if (scenario === 'request-history') assertRequestHistoryBudget(median(totalMs))
+      else expectTotalWithinBudget(median(totalMs), budgetMs)
       if (scenario === 'tool-continuation') {
         expect(median((samples as ContinuationReport[]).map(sample => sample.retainedHeapMb)))
           .toBeLessThanOrEqual(retainedHeapBudgetMb)