Bläddra i källkod

docs(perf): record first frontend CI calibration

Tianyi Cui 1 månad sedan
förälder
incheckning
82cd35467a

+ 2 - 2
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md
-2026-09-06-frontend-performance-budgets.md: 0df2e9c90b6192a7ededd445c24e35c5b9b6e7dc
-2026-09-06-frontend-performance-budgets.zh.md: e1cc7314df90009cce7f994b8d5dc0ca0f9e5807
+2026-09-06-frontend-performance-budgets.md: 7136ff614dcbe47511acff9619e4529d30ae0d87
+2026-09-06-frontend-performance-budgets.zh.md: 60b4071ce7f6712931d7bdc7b17f7bcbe78dca9d

+ 19 - 1
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md

@@ -22,7 +22,7 @@ Reconnect uses three fresh compiled plain-Node children. Each creates a 100,000-
 
 ## Calibration
 
-Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at product revision `925e012340`, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. Source reference constants retain the original allowances while CI calibration is pending; the bounded-observer 261.60 ms paging median exceeds its 260 ms reference allowance but remains below its 650 ms CI limit; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The existing scale comes from Node CI calibration, not a measured x64 browser comparison; browser-specific runner calibration remains an explicit gap.
+Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149.0.7827.55, at product revision `925e012340`, establish the baseline below. An isolated repeat follows a complete workflow smoke. Each browser sample reports raw endpoint values and every page; the paging verdict uses the median of the sample maxima. Reconnect reports all child measurements. Source reference constants retain the original allowances while CI calibration is pending; the bounded-observer 261.60 ms paging median exceeds its 260 ms reference allowance but remains below its 650 ms CI limit; the shared 2× time scale and 1.25× variance allowance produce CI limits. Memory uses only variance allowance. The shared scale originates in Node CI calibration. The first actual x64 browser run below passes the fixed budgets; a second independent CI run remains pending, so repeated-run browser calibration is not complete.
 
 | Endpoint | Measured median | Reference allowance | CI limit |
 |---|---:|---:|---:|
@@ -38,6 +38,24 @@ Three-sample medians on the arm64 reference machine, Node 24.19 and Chromium 149
 
 Draft typing spans 124.97–504.96 ms across the three isolated samples; the reference remains 500 ms and the scaled CI limit covers that observed spread; the median is not a per-keystroke bound. No budget is an environment override. Temporary zero allowances exercise every rejection path; these negative controls prove enforcement, not an optimization or a historical regression. A separate control waits for the final reply marker before typing and fails the actual-input overlap assertion. The compact synthetic JSONL is 3,262,577 bytes; all three corrected samples report an overlapping trusted input event and end after the 241st rendered turn-tail.
 
+### First actual CI run
+
+[Run 34020120425, benchmark job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) passes the complete benchmark inventory at `6d1ba089e5052680961825c08aa4de19b4fe137a`. The runner is `VM-7-113-ubuntu-ci-19` in `dsh-selfhosted-ci`, using x64 Node 24.19.0 and Chromium 149.0.7827.55. The following medians use three fresh samples per scenario and leave the local reference table and source budgets unchanged.
+
+| Endpoint | First CI median |
+|---|---:|
+| Browser open | 303.066 ms |
+| Slowest older page | 432.979 ms |
+| First Trajectory | 298.608 ms |
+| First reply | 740.265 ms |
+| Stream main-thread task | 1614.594 ms |
+| Draft typing | 932.746 ms |
+| Complete response | 1677.882 ms |
+| Reconnect replacement | 29.232 ms |
+| Reconnect retained heap | 23.028 MiB |
+
+All three browser samples report `inputOverlapped: true` and finish after the 241st rendered turn-tail. Post-GC browser heap is approximately 52.94 MiB with 17,064 DOM elements; both remain diagnostic endpoints. This run supports the existing budgets on this runner, not a universal 2× browser speed ratio. The required second independent CI run is pending; no budget is relaxed and no product optimization is claimed.
+
 ## Alternatives considered
 
 **Use the Node fold as paint evidence.** Rejected because it never performs DOM mutation, layout, or browser scheduling. The focused reconnect case likewise makes no GUI speed claim.

+ 19 - 1
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.zh.md

@@ -22,7 +22,7 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 ## 校准
 
-在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和产品版本 `925e012340` 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。源码参考常量在 CI 校准待完成期间保留原额度;受限观察器的分页中位数 261.60 ms 超过 260 ms 参考额度,但仍低于 650 ms CI 限制;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。现有倍率来自 Node CI 校准,并非实测 x64 浏览器对比;浏览器专用 runner 校准仍是明确缺口。
+在 arm64 参考机器、Node 24.19、Chromium 149.0.7827.55 和产品版本 `925e012340` 上,三个样本的中位数建立下表基线。完整工作流 smoke 后执行一次隔离重复测量。每个浏览器样本报告原始终点数据和每一页;分页判定使用各样本最大值的中位数。重连报告全部子进程测量。源码参考常量在 CI 校准待完成期间保留原额度;受限观察器的分页中位数 261.60 ms 超过 260 ms 参考额度,但仍低于 650 ms CI 限制;共享的 2× 时间倍率和 1.25× 方差余量产生 CI 限制。内存仅使用方差余量。共享倍率源自 Node CI 校准。下述首次实际 x64 浏览器运行通过固定预算;第二次独立 CI 运行仍待完成,因此浏览器重复运行校准尚未完成。
 
 | 终点 | 实测中位数 | 参考额度 | CI 限制 |
 |---|---:|---:|---:|
@@ -38,6 +38,24 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 三个隔离样本中的草稿键入时间为 124.97–504.96 ms;参考额度保持 500 ms,缩放后的 CI 限制覆盖观察到的波动;中位数不是单次按键上限。预算不能通过环境变量覆盖。临时零额度覆盖每条拒绝路径;这些负向对照证明预算执行,而非优化或历史回归。另一项对照在键入前等待最终回复标记,实际输入重叠断言因此失败。紧凑合成 JSONL 为 3,262,577 字节;三个修正样本均报告重叠的真实输入事件,并在第 241 个 turn-tail 渲染后结束。
 
+### 首次实际 CI 运行
+
+[运行 34020120425,基准 job 101451135853](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34020120425/job/101451135853) 在 `6d1ba089e5052680961825c08aa4de19b4fe137a` 上通过完整基准清单。runner 为 `dsh-selfhosted-ci` 中的 `VM-7-113-ubuntu-ci-19`,使用 x64 Node 24.19.0 和 Chromium 149.0.7827.55。下列中位数来自每个场景的三个全新样本,本地参考表和源码预算保持不变。
+
+| 终点 | 首次 CI 中位数 |
+|---|---:|
+| 浏览器打开 | 303.066 ms |
+| 最慢更早分页 | 432.979 ms |
+| 首次 Trajectory | 298.608 ms |
+| 首段回复 | 740.265 ms |
+| 流式主线程任务 | 1614.594 ms |
+| 草稿键入 | 932.746 ms |
+| 完整回复 | 1677.882 ms |
+| 重连替换 | 29.232 ms |
+| 重连保留 heap | 23.028 MiB |
+
+三个浏览器样本均报告 `inputOverlapped: true`,并在第 241 个 turn-tail 渲染后结束。强制 GC 后浏览器 heap 约为 52.94 MiB,DOM 元素为 17,064 个;两者仍为诊断终点。该运行支持此 runner 上的现有预算,而不证明普遍适用的 2× 浏览器速度比。必需的第二次独立 CI 运行仍待完成;没有放宽预算,也不声称产品优化。
+
 ## 考虑过的替代方案
 
 **用 Node 折叠作为绘制证据。** 拒绝,因为它不执行 DOM 修改、布局或浏览器调度。聚焦重连用例同样不声称 GUI 提速。