瀏覽代碼

docs(perf): confirm repeated frontend CI calibration

Tianyi Cui 4 周之前
父節點
當前提交
00e452aa03

+ 2 - 2
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md
-2026-09-06-frontend-performance-budgets.md: 7136ff614dcbe47511acff9619e4529d30ae0d87
-2026-09-06-frontend-performance-budgets.zh.md: 60b4071ce7f6712931d7bdc7b17f7bcbe78dca9d
+2026-09-06-frontend-performance-budgets.md: 9cbe5980336ad4b0f2a3fb2f7cbe70757885f82e
+2026-09-06-frontend-performance-budgets.zh.md: 6df0d00285d60ce1ea0ec9d7284e7a73924459ab

+ 15 - 15
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md

@@ -22,7 +22,7 @@ Reconnect uses three fresh compiled plain-Node children. Each creates a 100,000-
 
 ## Calibration
 
-Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at product revision `925e012340`, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. Source reference constants retain the original allowances while CI calibration is pending; the bounded-observer 261.60 ms paging median exceeds its 260 ms reference allowance but remains below its 650 ms CI limit; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The shared scale originates in Node CI calibration. The first actual x64 browser run below passes the fixed budgets; a second independent CI run remains pending, so repeated-run browser calibration is not complete.
+Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at product revision `925e012340`, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. Source reference constants retain the original allowances after two passing CI runs; the bounded-observer 261.60 ms paging median exceeds its 260 ms reference allowance but remains below its 650 ms CI limit; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The shared scale originates in Node CI calibration. Both actual x64 browser runs below pass the fixed budgets on unchanged benchmark code; this supplies repeated-run evidence for these runners, not a universal browser speed ratio.
 
 | Endpoint | Measured median | Reference allowance | CI limit |
 |---|---:|---:|---:|
@@ -38,23 +38,23 @@ Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149
 
 Draft typing spans 124.97–504.96 ms across the three isolated samples; the reference remains 500 ms and the scaled CI limit covers that observed spread; the median is not a per-keystroke bound. No budget is an environment override. Temporary zero allowances exercise every rejection path; these negative controls prove enforcement, not an optimization or a historical regression. A separate control waits for the final reply marker before typing and fails the actual-input overlap assertion. The compact synthetic JSONL is 3,262,577 bytes; all three corrected samples report an overlapping trusted input event and end after the 241st rendered turn-tail.
 
-### First actual CI run
+### Actual CI runs
 
-[Run 34020120425, benchmark job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) passes the complete benchmark inventory at `6d1ba089e5052680961825c08aa4de19b4fe137a`. The runner is `VM-7-113-ubuntu-ci-19` in `dsh-selfhosted-ci`, using x64 Node 24.19.0 and Chromium 149.0.7827.55. The following medians use three fresh samples per scenario and leave the local reference table and source budgets unchanged.
+[Run 34020120425, benchmark job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) passes the complete benchmark inventory at `6d1ba089e5052680961825c08aa4de19b4fe137a`. The runner is `VM-7-113-ubuntu-ci-19` in `dsh-selfhosted-ci`, using x64 Node 24.19.0 and Chromium 149.0.7827.55. [Attempt 2, benchmark job 101453296071](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101453296071) also passes the complete inventory at the same commit, on `VM-7-113-ubuntu-ci-25` with the same Node and Chromium versions. The following medians use three fresh samples per scenario in each run and leave the local reference table and source budgets unchanged.
 
-| Endpoint | First CI median |
-|---|---:|
-| Browser open | 303.066 ms |
-| Slowest older page | 432.979 ms |
-| First Trajectory | 298.608 ms |
-| First reply | 740.265 ms |
-| Stream main-thread task | 1614.594 ms |
-| Draft typing | 932.746 ms |
-| Complete response | 1677.882 ms |
-| Reconnect replacement | 29.232 ms |
-| Reconnect retained heap | 23.028 MiB |
+| Endpoint | First CI median | Second CI median |
+|---|---:|---:|
+| Browser open | 303.066 ms | 284.726 ms |
+| Slowest older page | 432.979 ms | 413.798 ms |
+| First Trajectory | 298.608 ms | 267.372 ms |
+| First reply | 740.265 ms | 658.906 ms |
+| Stream main-thread task | 1614.594 ms | 1035.385 ms |
+| Draft typing | 932.746 ms | 142.148 ms |
+| Complete response | 1677.882 ms | 1641.702 ms |
+| Reconnect replacement | 29.232 ms | 31.674 ms |
+| Reconnect retained heap | 23.028 MiB | 23.028 MiB |
 
-All three browser samples report `inputOverlapped: true` and finish after the 241st rendered turn-tail. Post-GC browser heap is approximately 52.94 MiB with 17,064 DOM elements; both remain diagnostic endpoints. This run supports the existing budgets on this runner, not a universal 2× browser speed ratio. The required second independent CI run is pending; no budget is relaxed and no product optimization is claimed.
+All six browser samples report `inputOverlapped: true` and finish after the 241st rendered turn-tail. Post-GC browser heap is approximately 52.94 MiB in the first run and 53.00 MiB in the second, with 17,064 DOM elements in both; these remain diagnostic endpoints. Both runs support the existing budgets on these runners, not a universal 2× browser speed ratio. Draft-typing medians vary from 932.746 ms to 142.148 ms because the endpoint measures the entire typed draft, including scheduling and Playwright actionability, rather than a per-key latency guarantee. No budget is relaxed and no product optimization is claimed.
 
 ## Alternatives considered
 

+ 15 - 15
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.zh.md

@@ -22,7 +22,7 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 ## 校准
 
-在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和产品版本 `925e012340` 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。源码参考常量在 CI 校准待完成期间保留原额度;受限观察器的分页中位数 261.60 ms 超过 260 ms 参考额度,但仍低于 650 ms CI 限制;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。共享倍率源自 Node CI 校准。下述首次实际 x64 浏览器运行通过固定预算;第二次独立 CI 运行仍待完成,因此浏览器重复运行校准尚未完成。
+在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和产品版本 `925e012340` 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。源码参考常量在两次 CI 运行通过后保留原额度;受限观察器的分页中位数 261.60 ms 超过 260 ms 参考额度,但仍低于 650 ms CI 限制;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。共享倍率源自 Node CI 校准。下述两次实际 x64 浏览器运行在基准代码不变的情况下均通过固定预算;这提供这些 runner 的重复运行证据,而非普遍适用的浏览器速度比。
 
 | 终点 | 实测中位数 | 参考额度 | CI 限制 |
 |---|---:|---:|---:|
@@ -38,23 +38,23 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 三个隔离样本中的草稿键入时间为 124.97–504.96 ms;参考额度保持 500 ms,缩放后的 CI 限制覆盖观察到的波动;中位数不是单次按键上限。预算不能通过环境变量覆盖。临时零额度覆盖每条拒绝路径;这些负向对照证明预算执行,而非优化或历史回归。另一项对照在键入前等待最终回复标记,实际输入重叠断言因此失败。紧凑合成 JSONL 为 3,262,577 字节;三个修正样本均报告重叠的真实输入事件,并在第 241 个 turn-tail 渲染后结束。
 
-### 首次实际 CI 运行
+### 实际 CI 运行
 
-[运行 34020120425,基准 job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) 在 `6d1ba089e5052680961825c08aa4de19b4fe137a` 上通过完整基准清单。runner 为 `dsh-selfhosted-ci` 中的 `VM-7-113-ubuntu-ci-19`,使用 x64 Node 24.19.0 和 Chromium 149.0.7827.55。下列中位数来自每个场景的三个全新样本,本地参考表和源码预算保持不变。
+[运行 34020120425,基准 job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) 在 `6d1ba089e5052680961825c08aa4de19b4fe137a` 上通过完整基准清单。runner 为 `dsh-selfhosted-ci` 中的 `VM-7-113-ubuntu-ci-19`,使用 x64 Node 24.19.0 和 Chromium 149.0.7827.55。[第 2 次执行,基准 job 101453296071](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101453296071) 在相同 commit 上也通过完整清单,runner 为 `VM-7-113-ubuntu-ci-25`,Node 和 Chromium 版本相同。下列中位数来自每次运行中每个场景的三个全新样本,本地参考表和源码预算保持不变。
 
-| 终点 | 首次 CI 中位数 |
-|---|---:|
-| 浏览器打开 | 303.066 ms |
-| 最慢更早分页 | 432.979 ms |
-| 首次 Trajectory | 298.608 ms |
-| 首段回复 | 740.265 ms |
-| 流式主线程任务 | 1614.594 ms |
-| 草稿键入 | 932.746 ms |
-| 完整回复 | 1677.882 ms |
-| 重连替换 | 29.232 ms |
-| 重连保留 heap | 23.028 MiB |
+| 终点 | 首次 CI 中位数 | 第二次 CI 中位数 |
+|---|---:|---:|
+| 浏览器打开 | 303.066 ms | 284.726 ms |
+| 最慢更早分页 | 432.979 ms | 413.798 ms |
+| 首次 Trajectory | 298.608 ms | 267.372 ms |
+| 首段回复 | 740.265 ms | 658.906 ms |
+| 流式主线程任务 | 1614.594 ms | 1035.385 ms |
+| 草稿键入 | 932.746 ms | 142.148 ms |
+| 完整回复 | 1677.882 ms | 1641.702 ms |
+| 重连替换 | 29.232 ms | 31.674 ms |
+| 重连保留 heap | 23.028 MiB | 23.028 MiB |
 
-三个浏览器样本均报告 `inputOverlapped: true`,并在第 241 个 turn-tail 渲染后结束。强制 GC 后浏览器 heap 约为 52.94 MiB,DOM 元素为 17,064 个;两者仍为诊断终点。该运行支持此 runner 上的现有预算,而不证明普遍适用的 2× 浏览器速度比。必需的第二次独立 CI 运行仍待完成;没有放宽预算,也不声称产品优化。
+六个浏览器样本均报告 `inputOverlapped: true`,并在第 241 个 turn-tail 渲染后结束。强制 GC 后浏览器 heap 首次运行约为 52.94 MiB,第二次约为 53.00 MiB,两次 DOM 元素均为 17,064 个;这些仍为诊断终点。两次运行支持这些 runner 上的现有预算,而不证明普遍适用的 2× 浏览器速度比。草稿键入中位数从 932.746 ms 变化到 142.148 ms,因为该终点测量整个草稿键入,包含调度和 Playwright 可交互性等待,而非单次按键延迟保证。没有放宽预算,也不声称产品优化。
 
 ## 考虑过的替代方案