|
|
@@ -37,7 +37,7 @@ Normal-heap mode performs a fixed pair of explicit garbage collections after Hos
|
|
|
|
|
|
A small untimed fixture prerequisite verifies current migration, message preservation, immutable V0 bytes, and successor reopen. Worker failures retain the first and last ten stderr lines, or fatal heap diagnostics, so setup rejection remains distinguishable from a budget breach. The timed performance cases do not duplicate semantic assertions owned by functional tests; it requires only that the target call completes and reaches its measured endpoint. The Client-fold benchmark continues to use the real `ConversationNodeAssembler` and every Chat Definition, and requires both the large window's absolute time and its scaling relative to the small window to remain below fixed budgets.
|
|
|
|
|
|
-Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Except for current-generation `open`, source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. Current-generation `open` uses a directly measured standard-runner expectation of 50 ms with only the 1.25× headroom, rounded up to a 63 ms budget. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
|
|
|
+Budgets are calibrated per measured endpoint. Two repeated Node 24.19 x64 CI runs differ by at most 5.2% in their medians; their CPU-heavy wall times are 1.95–2.06× the Node 24.18 arm64 reference run. Except for current-generation `open` and first-open Agent resume, source constants record expected reference-machine durations; `ciTimeBudget()` multiplies them by the measured 2× CI time scale and 1.25× variance headroom. Current-generation `open` uses a directly measured standard-runner expectation of 50 ms with only the 1.25× headroom, rounded up to a 63 ms budget. First-open Agent resume uses a reviewed 562 ms hosted limit. The retained-heap and Client-fold scaling budgets use only the 1.25× headroom because neither is a wall-clock duration. The 128 MB completion check remains an independent transient-allocation limit. The resulting first-open time limits, constrained-heap checks, and Client-fold limits all reject the known regressions. Pre-stack commit `0d7ea53743e273930a31e9e2b6ca682f21dd4ca5` is the fixed calibration and review reference; CI does not check out or execute the historical repository. Budgets are reviewed source constants and have no environment-variable override.
|
|
|
|
|
|
## Calibration evidence
|
|
|
|
|
|
@@ -56,7 +56,9 @@ The pre-stack implementation keeps V0 as its current format, so first open does
|
|
|
|
|
|
The [standard two-CPU run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34023970384/job/101461539961) at `ca3ffe95dac2c55eefeb16ed9b61067bbd19ee90` uses Node 24.20.0 x64 and Ubuntu image `20260831.293.1`. Its five current-generation `open` samples are 49.2, 47.4, 49.1, 48.6, and 48.1 ms: median 48.6 ms, maximum 49.2 ms. The rounded 50 ms CI expectation gives a 63 ms limit without reapplying the 2× machine scale. The log identifies two available CPUs but not their model; it does not isolate hardware from the Node-version change. This is endpoint-specific runner calibration, not evidence of an application optimization or a new reference-machine measurement. Every other benchmark passes its existing budget. Deterministic controls reject the observed median at the historical 30 ms limit, accept it at 63 ms, reject a synthetic 75 ms reopen median, and reject a synthetic 4,000 ms first-open duration at its unchanged 550 ms limit. These controls verify budget enforcement, not a measured new regression.
|
|
|
|
|
|
-A cold-verifier packaging change removes runtime workspace-module loading without changing these budgets or the measured endpoint. On macOS arm64, Node 24.18.0, the same 127,400-event fixture at `ac48359b195558806ee5a2286697074fd1a52815` takes 164.2, 162.4, 159.7, 149.3, and 167.7 ms for first writable resume (median 162.4 ms). Bundling the verifier through the workspace build gives 119.8, 120.9, 121.9, 122.3, and 121.8 ms (median 121.8 ms, 25% lower). Retained heap stays at 5.4 MB; median peak RSS changes from 144.9 to 143.7 MB. Reopen medians are 27.5 and 27.1 ms, and all 16 Session cases, including the 128 MB completion checks, pass. A CPU profile attributes part of the old verifier cost to module resolution and compilation. The isolated-package built-worker test fails on the original worker because its workspace imports cannot resolve, and passes with the bundled worker, including rejection of an incorrect event count. These local results do not establish Linux runner timing; the existing 450 ms CI gate remains the acceptance check.
|
|
|
+A cold-verifier packaging change removes runtime workspace-module loading without changing these budgets or the measured endpoint. On macOS arm64, Node 24.18.0, the same 127,400-event fixture at `ac48359b195558806ee5a2286697074fd1a52815` takes 164.2, 162.4, 159.7, 149.3, and 167.7 ms for first writable resume (median 162.4 ms). Bundling the verifier through the workspace build gives 119.8, 120.9, 121.9, 122.3, and 121.8 ms (median 121.8 ms, 25% lower). Retained heap stays at 5.4 MB; median peak RSS changes from 144.9 to 143.7 MB. Reopen medians are 27.5 and 27.1 ms, and all 16 Session cases, including the 128 MB completion checks, pass. A CPU profile attributes part of the old verifier cost to module resolution and compilation. The isolated-package built-worker test fails on the original worker because its workspace imports cannot resolve, and passes with the bundled worker, including rejection of an incorrect event count. These local results do not establish Linux runner timing; those cases use the 450 ms CI limit.
|
|
|
+
|
|
|
+The [hosted run at `a7884138be`](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34265057987/job/102192211510) includes the bundled verifier and reports first-open Agent-resume samples of 454.2, 454.8, 455.4, 457.8, and 459.8 ms: median 455.4 ms against 450 ms. The code-equivalent [preceding run](https://github.com/deepseek-harness/deepseek-harness/actions/runs/34263062688/job/102185561214) reports a 436.2 ms median; only the bilingual request-history README and its pairing record differ between those heads. The reviewed ceiling is 562 ms, `floor(450 × 1.25)`, a 24.89% increase that leaves 23.4% above the observed 455.4 ms median. Five fresh M4 Pro / Node 24.19 samples span 150.07–158.22 ms with a 152.57 ms median. A bounded main-thread profile identifies no obvious small optimization; it excludes verifier-thread CPU and does not establish the cause of hosted variation. Controls reject the recorded median at 450 ms, accept it at 562 ms, and reject 600 ms. The bundled-verifier improvement, workload, other time budgets, shared scaling, and memory limits remain intact.
|
|
|
|
|
|
The calibrated source budgets are:
|
|
|
|
|
|
@@ -69,7 +71,7 @@ The calibrated source budgets are:
|
|
|
| Projection | 14 ms | 35 ms |
|
|
|
| First-open first history | 220 ms | 550 ms |
|
|
|
| Current-generation first history | 48 ms | 120 ms |
|
|
|
-| First-open Agent resume | 180 ms | 450 ms |
|
|
|
+| First-open Agent resume | 180 ms (historical reference) | 562 ms |
|
|
|
| Current-generation Agent resume | 40 ms | 100 ms |
|
|
|
| Agent retained heap | 26.1 MB | 33 MB |
|
|
|
| Client-fold absolute time | 16 ms | 40 ms |
|