Ver Fonte

docs(benchmarks): record the CG-30 A/B — deterministic win, no behavioural regression

Primary evidence is deterministic: on django, query.py rendered 2.12x its budget
on main and 1.49x with the bound, and the freed bytes reach the files below it
(+2,104 chars of source in the same five files). gin is a true control — the two
builds emit byte-identical explore output there, which is what makes its agent-run
deltas variance by construction.

Also records the harness trap that voided the first two batches: ab-new-vs-baseline
swaps src/ to the baseline ref mid-run, so a commit made while it runs captures
baseline sources. Check the "changed:" line before believing any run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Colby McHenry há 1 mês atrás
pai
commit
0d014a6582
1 ficheiros alterados com 108 adições e 0 exclusões
  1. 108 0
      docs/benchmarks/explore-oversize-member-ab-cg30.md

+ 108 - 0
docs/benchmarks/explore-oversize-member-ab-cg30.md

@@ -0,0 +1,108 @@
+# Agent A/B — bounded oversize cluster member (task CG-30)
+
+**Date:** 2026-08-06 · **New:** `bugfix/CG-30` · **Baseline:** `main` @ `d6d1728` ·
+**Harness:** `scripts/agent-eval/ab-new-vs-baseline.sh`, `--model sonnet --effort high`,
+**both arms codegraph-on**, CLI blocked (0 contamination in every run),
+`CODEGRAPH_NO_PROMPT_HOOK=1`.
+
+CG-30 bounds how far a cluster's top member may overshoot what its file may spend: past 1.5x it
+is windowed on whole lines instead of emitted whole. The risk the A/B exists to price is the one
+CLAUDE.md names — a section that is no longer sufficient sends the agent to Read, and one or two
+of those teach it to stop calling codegraph at all.
+
+**Verdict: no regression, and the deterministic win is unambiguous.** The behavioural bar holds
+(Read/Grep ~0, no abandonment, allocation efficiency 100% on the repo where the bound engages),
+the cost is a ~10% median duration increase on django inside overlapping ranges, and the one
+allocation-miss call the new arm produced is matched by two recall-miss calls in the baseline.
+
+> **Harness note, recorded because it cost a re-run:** `ab-new-vs-baseline.sh` checks the engine
+> out at the BASELINE ref while its baseline arm runs and restores it on exit. **Do not commit
+> while it is running** — a commit made mid-run captures baseline sources. The first django/gin
+> batches were void for exactly this reason (their `changed:` line listed only
+> `explore-diagnostics.ts`, i.e. both arms ran identical retrieval code) and were re-run. Check
+> that line before believing any A/B in this harness.
+
+---
+
+## Deterministic measurement — where the bound actually engages
+
+Same index, same query, both builds. This is the primary evidence; the agent runs below only
+price the risk.
+
+**django** — `codegraph explore "How does a QuerySet turn into SQL and fetch rows from the
+database?"`
+
+| | baseline | new |
+|---|---|---|
+| `django/db/models/query.py` | 7,784 chars on a 3,669 budget — **2.12x** | 5,464 — **1.49x**, windowed |
+| `django/contrib/admin/filters.py` | 3,633 (inherited 2,271 spendable) | **8,057** (inherited 9,160) |
+| source delivered | 17,929 chars, 5 files | **20,033** chars, 5 files |
+
+The reported CG-30 signature, reproduced on a public repo and then closed: the rank-#1 file took
+2.12x its budget, and the files below it inherited the shortfall. Bounding it hands those bytes
+straight down the rank order — the response carries the same five files and 2,104 more chars of
+actual source.
+
+**gin (control)** — the two builds produce **byte-identical** explore output (13,457 chars) for
+the route-dispatch query. Nothing in gin is oversize enough for the bound to engage (max
+observed 0.94x of spendable), which is exactly what a control should show — and it means every
+gin number in the agent table below is run-to-run variance, not the change.
+
+**Fixture** — `__tests__/fixtures/oversize-member-ts`, three report builders competing for one
+envelope, each a single long function:
+
+| File | baseline | new |
+|---|---|---|
+| `monthly.ts` (24.5K, one ~490-line function) | 12,391 chars on a 3,334 budget — **3.7x** | 4,941 — **1.48x**, windowed |
+| `quarterly.ts` (11.4K, one ~200-line function) | **dropped** — `budget-clusters`, no headroom left | 4,004 delivered |
+| response | 19,223 chars, 3 files | 15,852 chars, 4 files |
+
+The two rows are the same defect from both sides: a member bigger than the file's share eats the
+envelope, and a member bigger than the whole response ceiling makes the file vanish. Pinned by
+`__tests__/explore-oversize-member.test.ts` (9 tests; 4 fail on `main`).
+
+## Agent runs
+
+| | django new | django base | gin new | gin base | excalidraw new | excalidraw base |
+|---|---|---|---|---|---|---|
+| runs | 5 | 5 | 3 | 3 | 2 | 2 |
+| duration (s) | 39 [36–71] | 35 [35–60] | 39 [37–51] | 34 [28–46] | 52 [43–60] | 41 [40–42] |
+| tool calls | 3 [3–10] | 4 [3–23] | 4 [3–4] | 3 | 4 [3–5] | 4 [3–4] |
+| codegraph calls | 2 [2–3] | 2 [0–3] | 2 [2–3] | 2 | 3 [2–4] | 3 [2–3] |
+| Read | 0 [0–5] | 0 [0–13] | 0 [0–1] | 0 | 0 | 0 |
+| Grep/Glob | 0 | 0 | 0 | 0 | 0 | 0 |
+| occupancy share | 33.1% [30.7%–49.5%] | 34.4% [29.3%–47.3%] | 28.9% [26.4%–37.3%] | 30.1% [28.6%–31.9%] | 43.0% | 39.3% |
+| allocation efficiency | 100.0% | 98.6% | 88.5% | 98.3% | 85.8% | 90.5% |
+
+django is pooled over two batches (n=2 + n=3). Questions: django "How does a QuerySet turn into
+SQL and fetch rows from the database? Trace the flow end to end."; gin "How does a registered
+route handler get invoked for an incoming HTTP request?…"; excalidraw "How does updating an
+element re-render the canvas on screen?…".
+
+**Sufficiency, pooled per call — the bar that matters.** django: new 1 "Read a file we returned"
+in 10 answered calls (the allocation-miss signal a window would trip first) against the
+baseline's 1 "Read a file we did not return" + 1 Grep in 10 — a shift in miss type, not an
+increase. gin: 1 allocation miss in 7 against 0 in 6, on a repo where the two builds emit
+identical bytes, so it is variance by construction. excalidraw: 0 misses in either arm.
+
+**Where the new arm looks worse, and why it is not read as a regression:**
+
+- *django duration, ~10% slower median.* Ranges overlap (36–71 vs 35–60) at n=5, and one
+  baseline run lost its codegraph attach entirely (0 codegraph calls, 13 Reads, 23 tool calls),
+  which distorts that arm's spread in both directions.
+- *gin allocation efficiency 88.5% vs 98.3%.* The builds are byte-identical on gin. This is the
+  metric's documented relativity — attribution is by citation and the agent's follow-up queries
+  differ per run — not an effect of the change.
+- *excalidraw occupancy/duration.* Call-count noise: one of the two new-arm runs made a 4th
+  explore call where the baseline made 2–3, and duration, envelope and occupancy all follow it.
+  Per-call envelope is flat (20,015 vs 19,446 chars/call), Read/Grep stay 0, tool calls match.
+  CLAUDE.md's own worked example records 3–10 codegraph calls on this prompt.
+
+## Caveat carried forward
+
+The `self-query` probe fixture in `scripts/agent-eval/allocation-fixtures.json` flips to FAIL
+under this change. Allocation is unchanged between arms and `tools.ts` delivers the identical
+8,282 chars in both — what changed is that an over-reserved incidental file now *delivers*
+instead of being cut by the hard-ceiling truncation, which is what its previous PASS depended
+on. Recorded as that fixture's `afterCG30` block. The over-reservation itself is epic CG-24's
+subject; it should not be answered by loosening this bound.