Sfoglia il codice sorgente

test(perf): respect browser failover provisioning policy

Tianyi Cui 4 settimane fa
parent
commit
c51c16cdac

+ 2 - 2
.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.md
-2026-07-24-web-gui-browser-e2e-lane.md: a996b380a55fc33f44cfdc2e31e179bc11f40be3
-2026-07-24-web-gui-browser-e2e-lane.zh.md: 1a413b6b8277ac696c6b728de740494bbf056695
+2026-07-24-web-gui-browser-e2e-lane.md: 07a37ada9c2a43f04612048f9bff6b22d022ec40
+2026-07-24-web-gui-browser-e2e-lane.zh.md: 668e612712175821d6ad123ce364a7cb96e01272

+ 1 - 1
.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.md

@@ -93,4 +93,4 @@ Surveyed AI-chat/agent web UIs and mocking layers (LibreChat, vercel/ai-chatbot
 
 ## Consequences
 
-The web surface gains its record-once/replay-forever tier: the real chromium → SSE → apiproxy → loop → tools → persistence chain runs keylessly in ~10-30s, deterministic across repeat runs, with fixtures owned and re-recordable by the lane itself. Costs accepted: every intentional conversation-UI change ends with a keyless `DSH_SNAPSHOT=refresh` (golden churn is reviewed diff, anchors keep semantic green); the aria format is Playwright-owned — the one committed snapshot format the repo does not control — so playwright version bumps must be deliberate bump-and-refresh commits (the dependency floats `^1.49.0` in `apps/web/package.json`; pin exactly if churn bites); replay's first-call-order binding constrains scenarios to one prompting session each, with the consumption assertion as the tripwire; `compaction-basic` shares the session's replay cursor and stays inert only under the published 128k catalog window; and the required consumer job pays for Chromium provisioning and one browser run so the PR that changes the assembled UI owns its expected-output diff. The opt-in performance lane preserves a repeatable diagnostic workload without adding host-sensitive duration or memory expectations to CI; performance regressions remain a manually interpreted signal until the repository owns a calibrated benchmark environment.
+The web surface gains its record-once/replay-forever tier: the real chromium → SSE → apiproxy → loop → tools → persistence chain runs keylessly in ~10-30s, deterministic across repeat runs, with fixtures owned and re-recordable by the lane itself. Costs accepted: every intentional conversation-UI change ends with a keyless `DSH_SNAPSHOT=refresh` (golden churn is reviewed diff, anchors keep semantic green); the aria format is Playwright-owned — the one committed snapshot format the repo does not control — so playwright version bumps must be deliberate bump-and-refresh commits (the dependency floats `^1.49.0` in `apps/web/package.json`; pin exactly if churn bites); replay's first-call-order binding constrains scenarios to one prompting session each, with the consumption assertion as the tripwire; `compaction-basic` shares the session's replay cursor and stays inert only under the published 128k catalog window; and the required consumer job pays for Chromium provisioning and one browser run so the PR that changes the assembled UI owns its expected-output diff. The opt-in performance lane preserves a repeatable, threshold-free diagnostic workload whose measurements require manual interpretation. The separate required [frontend performance benchmarks](2026-09-06-frontend-performance-budgets.md) enforce calibrated budgets in the isolated benchmark CI job; they do not add thresholds to the manual inventory.

+ 1 - 1
.agents/notes/implemented/testing/2026-07-24-web-gui-browser-e2e-lane.zh.md

@@ -93,4 +93,4 @@ Web GUI 以一条真实组装链交付——chromium 页面 → client 插件 bu
 
 ## 后果
 
-Web 表面获得了录制一次/永久回放的层级:真实 chromium → SSE → apiproxy → 循环 → 工具 → 持久化的链路以约 10-30 秒无密钥运行,重复运行结果确定,fixture 由车道自身持有并可重录。接受的成本:每次有意的会话 UI 变更都以一次无密钥 `DSH_SNAPSHOT=refresh` 收尾(预期输出变动是受评审的 diff,锚断言保住语义绿色);aria 格式归 Playwright 所有——仓库唯一不受自己控制的提交快照格式——因此 playwright 版本升级必须是刻意的升级加刷新提交(依赖在 `apps/web/package.json` 中浮动为 `^1.49.0`;若变动伤人则改为精确锁定);回放的首次调用顺序绑定把每个场景限制为至多一个发起提示的会话,消费断言是绊线;`compaction-basic` 与会话共享回放游标,仅在目录中发布的 128k 上下文窗口下保持闲置;必需的消费方任务承担 Chromium 供给与一次浏览器运行的成本,使改动组装后 UI 的 PR(Pull Request)持有相应的预期输出 diff。按需启用的性能车道保留了可重复的诊断工作负载,又不会向 CI 添加受 host 差异影响的时长或内存预期;在仓库拥有经校准的基准测试环境之前,性能回归仍是需要人工解读的信号。
+Web 表面获得了录制一次/永久回放的层级:真实 chromium → SSE → apiproxy → 循环 → 工具 → 持久化的链路以约 10-30 秒无密钥运行,重复运行结果确定,fixture 由车道自身持有并可重录。接受的成本:每次有意的会话 UI 变更都以一次无密钥 `DSH_SNAPSHOT=refresh` 收尾(预期输出变动是受评审的 diff,锚断言保住语义绿色);aria 格式归 Playwright 所有——仓库唯一不受自己控制的提交快照格式——因此 playwright 版本升级必须是刻意的升级加刷新提交(依赖在 `apps/web/package.json` 中浮动为 `^1.49.0`;若变动伤人则改为精确锁定);回放的首次调用顺序绑定把每个场景限制为至多一个发起提示的会话,消费断言是绊线;`compaction-basic` 与会话共享回放游标,仅在目录中发布的 128k 上下文窗口下保持闲置;必需的消费方任务承担 Chromium 供给与一次浏览器运行的成本,使改动组装后 UI 的 PR(Pull Request)持有相应的预期输出 diff。按需启用的性能车道保留可重复、无阈值的诊断工作负载,其测量需要人工解读。独立的必需[前端性能基准](2026-09-06-frontend-performance-budgets.zh.md)在隔离的基准 CI job 中执行经校准的预算;它们不向手动清单添加阈值。

+ 2 - 2
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.i18n.yaml

@@ -2,5 +2,5 @@
 # side as of the last confirmed-consistent state. Both languages carry equal authority;
 # after editing either side, bring the other along and re-record with:
 #   pnpm run verify-translation-pairing --write .agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md
-2026-09-06-frontend-performance-budgets.md: 9cbe5980336ad4b0f2a3fb2f7cbe70757885f82e
-2026-09-06-frontend-performance-budgets.zh.md: 6df0d00285d60ce1ea0ec9d7284e7a73924459ab
+2026-09-06-frontend-performance-budgets.md: 0474025b42c1f08bd964972081f2b612f3cbd8d8
+2026-09-06-frontend-performance-budgets.zh.md: 533c5784260674ed5f6834c83d172ccd4c01a96e

+ 1 - 1
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.md

@@ -72,4 +72,4 @@ All six browser samples report `inputOverlapped: true` and finish after the 241s
 
 The benchmark layer changes no product implementation or user-visible behavior. It adds approximately fifteen seconds of local browser/reconnect execution plus Web build and browser provisioning to the existing isolated CI lane. A fresh browser discards previous caches, but each workflow deliberately retains its own loaded history and previously activated Trajectory during continuation.
 
-The baseline is independently mergeable and protects current performance; optimization layers tighten budgets only with repeated measurements and focused semantic tests. It does not cover sidebar cardinality, an hours-long soak, GPU presentation, real model latency, a published Host launch, or reconnect rendering. The manual Web diagnostic and existing functional browser tests retain those separate responsibilities. The existing Session performance note remains active because it owns Node calibration and persistence rationale; this note extends rather than supersedes it.
+The baseline is independently mergeable and protects current performance; optimization layers tighten budgets only with repeated measurements and focused semantic tests. It does not cover sidebar cardinality, an hours-long soak, GPU presentation, real model latency, a published Host launch, or reconnect rendering. The [Web browser lane](2026-07-24-web-gui-browser-e2e-lane.md) retains its separate threshold-free manual diagnostics and functional browser tests; calibrated required measurements belong to this benchmark lane. The existing Session performance note remains active because it owns Node calibration and persistence rationale; this note extends rather than supersedes it.

+ 1 - 1
.agents/notes/implemented/testing/2026-09-06-frontend-performance-budgets.zh.md

@@ -72,4 +72,4 @@ Node 对话折叠很快,并不能证明浏览器能绘制长对话或在流式
 
 基准层不改变产品实现或用户可见行为。它在现有隔离 CI lane 中增加约十五秒的本地浏览器与重连执行,以及 Web 构建和浏览器安装成本。全新浏览器丢弃此前的缓存,但每个工作流刻意在续接期间保留自身已加载历史和曾激活的 Trajectory。
 
-基线可独立合并并保护现有性能;优化层只有在重复测量与聚焦语义测试支持下才收紧预算。它不覆盖侧栏数量级、数小时 soak、GPU 显示、真实模型延迟、发布版 Host 启动或重连渲染。手动 Web 诊断和现有功能浏览器测试继续各负其责。现有 Session 性能记录保持活跃,因为它拥有 Node 校准和持久化理由;本记录扩展而不替代它。
+基线可独立合并并保护现有性能;优化层只有在重复测量与聚焦语义测试支持下才收紧预算。它不覆盖侧栏数量级、数小时 soak、GPU 显示、真实模型延迟、发布版 Host 启动或重连渲染。[Web 浏览器车道](2026-07-24-web-gui-browser-e2e-lane.zh.md)保留独立的无阈值手动诊断与功能浏览器测试;经校准的必需测量由本基准车道负责。现有 Session 性能记录保持活跃,因为它拥有 Node 校准和持久化理由;本记录扩展而不替代它。

+ 7 - 1
.github/workflows/ci.yml

@@ -206,9 +206,15 @@ jobs:
       - name: Install (immutable)
         run: pnpm install --frozen-lockfile
 
-      - name: Install benchmark browser
+      - name: Install benchmark browser and hosted dependencies
+        if: vars.DSH_CI_FAILOVER_LINUX != 'selfhosted' || github.event.pull_request.user.login == 'dependabot[bot]'
         run: pnpm --filter @deepseek-ai/dsh-benchmarks exec playwright install --with-deps chromium
 
+      # The persistent VM image owns Linux system packages; do not run apt here.
+      - name: Install benchmark browser on the failover VM
+        if: vars.DSH_CI_FAILOVER_LINUX == 'selfhosted' && github.event.pull_request.user.login != 'dependabot[bot]'
+        run: pnpm --filter @deepseek-ai/dsh-benchmarks exec playwright install chromium
+
       - name: Run performance benchmarks
         env:
           DSH_GATE_VERBOSE: '1'

+ 1 - 1
benchmarks/long-session-browser/long-session.bench.ts

@@ -87,7 +87,7 @@ it('opens, pages, navigates and streams into a 240-turn browser history', async
             await page.getByRole('row').last().waitFor()
           })
           await page.getByRole('tab', { name: 'Chat', exact: true }).click()
-          await page.waitForFunction(selector => document.querySelectorAll(selector).length === 240, TAIL)
+          await page.waitForFunction(({ selector, expected }) => document.querySelectorAll(selector).length === expected, { selector: TAIL, expected: HISTORY_TURNS })
           const composer = page.locator('[data-composer-input][contenteditable="true"]').last()
           await composer.fill('Continue the synthetic review and summarize the validation. '.repeat(30))
           const cdp = await page.context().newCDPSession(page)

+ 10 - 0
scripts/ci-workflow.spec.ts

@@ -211,6 +211,16 @@ describe('CI workflow', () => {
     expect(aggregate.needs).toContain('node-24-bench')
     expect(node24Bench.name).toBe('node 24 / benchmarks')
     expect(node24Bench.env).toBeUndefined()
+    expect(node24Bench.steps).toContainEqual({
+      name: 'Install benchmark browser and hosted dependencies',
+      if: "vars.DSH_CI_FAILOVER_LINUX != 'selfhosted' || github.event.pull_request.user.login == 'dependabot[bot]'",
+      run: 'pnpm --filter @deepseek-ai/dsh-benchmarks exec playwright install --with-deps chromium',
+    })
+    expect(node24Bench.steps).toContainEqual({
+      name: 'Install benchmark browser on the failover VM',
+      if: "vars.DSH_CI_FAILOVER_LINUX == 'selfhosted' && github.event.pull_request.user.login != 'dependabot[bot]'",
+      run: 'pnpm --filter @deepseek-ai/dsh-benchmarks exec playwright install chromium',
+    })
     expect(node24Bench.steps).toContainEqual({
       name: 'Run performance benchmarks',
       env: { DSH_GATE_VERBOSE: '1' },