Explorar o código

docs: the occupancy baseline says our residual is higher — write that down (CG-13)

Fills the empty RESULTS placeholder with the 2026-08-05 campaign
(bgjob-6d357cd2: 7 repos x 2 arms x 4 runs x 3 turns, 137 min).

The finding is not the flattering one. Retrieval residual is 82% HIGHER
with codegraph and share-of-context 27% higher, on all seven repos —
vscode 67k resident against 18k. At the same time six of seven
without-arms *process* more total tokens (gin 660k vs 290k) while
leaving less behind. Both are true: one dense verbatim payload stays
resident where many small Read/Grep results evict. This corroborates
issue #1500 on our own harness; the aggregator used to print it as
"-82% lower with codegraph" until the sign bug at 520ed9d.

Also:

- States the regime everywhere. This ran claude-sonnet-5 / 3-turn; the
  README's table is Opus 4.8 / single-question. Measured 24/23/20/84
  against the published 60/69/20/89 — model and turn count, not
  contamination. Records the two inversions honestly (vscode processes
  98% more tokens, django costs 17% more) and that 4 of 28 with-arm
  sessions still touched Read.
- Corrects the "Settled" section, which claimed a 7-repo baseline
  existed before one did, and adds the unclaimed Opus rerun.
- Records the contamination gate: 0 CLI calls returned output in 56
  sessions, but 29 attempts were blocked — 26 of 28 without-arm
  sessions tried. no-cli-shim.sh is load-bearing, not precautionary.
- Records the secondary readings as absolute, not before/after: 86.7%
  allocation efficiency pooled over 110 calls, read-of-a-file-we-
  returned 2%, explore-again 73% and ambiguous by construction.

README.md is deliberately untouched — restating its numbers from sonnet
3-turn data would be wrong. A proposed README paragraph is drafted at
the end of the benchmark doc for the maintainer to accept or reject.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Colby McHenry hai 1 mes
pai
achega
5dd4db68cd

+ 4 - 1
docs/benchmarks/agent-eval-feedback-metrics.md

@@ -54,7 +54,10 @@ table still renders against whichever arm's logs are already in `$AGENT_EVAL_OUT
 
 **A campaign — `bench-readme.sh`.** The 7 README repos, three turns each,
 `RUNS` per arm, through `run-all.sh` — so every run in a campaign carries all
-three metrics. Aggregate with `parse-bench-readme.mjs`.
+three metrics. Aggregate with `parse-bench-readme.mjs`. One has been run:
+[the 2026-08-05 baseline](residual-context-occupancy.md#baseline-the-7-readme-repos)
+(sonnet, 3 turns, 4 runs/arm) — read its regime box before comparing anything to
+it, and note that it is **not** the regime the README's table was published in.
 
 **A log you already have.** `parse-run.mjs <run.jsonl> [run.tN.jsonl …]` prints
 the three blocks for any stream-json log; `--brief` drops the numbered call

+ 213 - 4
docs/benchmarks/residual-context-occupancy.md

@@ -168,18 +168,185 @@ without-arm may have been using codegraph.**
 
 ## Baseline: the 7 README repos
 
-<!-- RESULTS -->
+### The regime — read this before any number below
+
+| | This campaign | README's published table |
+|---|---|---|
+| Model | **`claude-sonnet-5`** | **Claude Opus 4.8** |
+| Session shape | **3 turns** (README question + 2 in-flow follow-ups) | **1 question** |
+| Runs | 4 per arm × 7 repos = **56 sessions** | 4 per arm, median |
+| Ran | 2026-08-05, 137 min (`bgjob-6d357cd2`), raw under `/tmp/ab-readme` | 2026-07-21 |
+
+**These two are not comparable, and the difference is model + turn count — not
+contamination.** Sonnet is the deliberate floor model for this harness
+(`CLAUDE.md`: an affordance that lands on Sonnet generalizes up; one that only
+works on Opus does not generalize down). Three turns is what makes occupancy
+chargeable at all. Both choices move the efficiency numbers, so the throughput
+row below reads *lower* than the README's and neither figure invalidates the
+other. Settling whether the published Opus figures still hold needs a matched
+**Opus 4.8, single-question** rerun; that is deliberately out of scope here.
+
+Reproduce with:
+
+```bash
+CORPUS=/tmp/codegraph-corpus scripts/agent-eval/bench-readme.sh   # RUNS=4 CG_TURNS=3
+node scripts/agent-eval/parse-bench-readme.mjs /tmp/ab-readme
+```
+
+### The finding: codegraph's residual is 82% HIGHER, on all seven repos
+
+```
+repo        turns W→WO   final ctx W→WO   residual W→WO       % of ctx W→WO   % of window W→WO
+vscode      12.5/44      113k→53k         67k→18k  (+276%)    59.7%→36.0%     33.7%→9.0%
+excalidraw  9/32         87k→57k          43k→25k   (+71%)    49.5%→47.6%     21.5%→12.5%
+django      7.5/16.5     60k→51k          18k→10k   (+71%)    29.3%→19.9%      8.8%→5.1%
+tokio       9/33.5       87k→64k          45k→31k   (+45%)    52.1%→50.5%     22.7%→15.7%
+okhttp      6/14         61k→59k          20k→16k   (+27%)    33.2%→27.6%     10.1%→8.0%
+gin         6/15.5       56k→49k          15k→8k    (+79%)    26.3%→16.7%      7.3%→4.1%
+alamofire   10/31        76k→65k          34k→32k    (+7%)    44.7%→50.6%     16.9%→15.8%
+
+AVERAGE: retrieval residual 82% HIGHER with codegraph · share-of-context 27% HIGHER
+```
+
+W = codegraph's responses still resident. WO = Read + Grep/Glob + Bash results
+still resident. `turns` is median assistant turns per session.
+
+**Seven of seven.** There is no repo where codegraph leaves less behind. On
+vscode it leaves **67k tokens resident against the without-arm's 18k** — a third
+of a 200k window, gone before turn 4 starts. The only near-tie is Alamofire
+(+7%), and it is a tie because that arm's share-of-context is actually *lower*
+(44.7% vs 50.6%), not because the residual is small.
+
+**Both things are true at once.** On six of seven repos the without-arm
+*processes* far more total tokens than the with-arm — gin 660k vs 290k, okhttp
+704k vs 302k — while leaving *less* behind. Throughput and stock are different
+quantities and they point opposite ways here:
+
+- codegraph front-loads **one large verbatim payload** (2 explore calls on gin,
+  each tens of thousands of dense source characters) and that payload **stays
+  resident** for every turn after it;
+- Read/Grep/Bash churn **many small results** (gin: ~6 reads + ~5 bash per run),
+  most of which are re-derivation the agent then discards, and which evict.
+
+**This corroborates issue [#1500](https://github.com/colbymchenry/codegraph/issues/1500)
+on our own harness.** The reporter's complaint was exactly this axis, and until
+this campaign we had no measurement that could see it. Note for anyone reading
+git history: the aggregator originally printed this as "-82% *lower* with
+codegraph" — a sign bug, fixed at `520ed9d`. The honest number is the entire
+point of the metric; do not soften it.
+
+**Fixed overhead.** codegraph's tool schema + MCP instructions cost **+546 tok**
+of context before any tool is called (median with-arm `ctxBase` minus median
+without-arm `ctxBase`, averaged over repos). Paid whether or not the agent ever
+calls codegraph. Small — the residual, not the schema, is where the context goes.
+
+### Throughput in the same campaign (sonnet · 3 turns)
+
+Reported for completeness and because the occupancy finding only means anything
+read against it. **These are not the README's numbers and must not be quoted as
+such.**
+
+```
+repo        time W→WO        tools W→WO   tokens W→WO (saved)   cost W→WO (saved)
+vscode      2m 59s→1m 59s    8→60         949k→478k  (-98%)     $1.21→$1.62  (25%)
+excalidraw  1m 45s→2m 1s     5→44         557k→869k   (36%)     $0.78→$1.03  (24%)
+django      1m 4s→1m 30s     3→13         366k→686k   (47%)     $0.52→$0.45 (-17%)
+tokio       1m 47s→4m 35s    5→43         574k→793k   (28%)     $0.67→$1.25  (47%)
+okhttp      49s→1m 25s       2→11         302k→704k   (57%)     $0.35→$0.48  (27%)
+gin         1m 6s→1m 43s     2→12         290k→660k   (56%)     $0.38→$0.43  (12%)
+alamofire   1m 35s→1m 47s    6→29         545k→870k   (37%)     $0.67→$1.30  (49%)
+
+AVERAGE saved: cost 24% · tokens 23% · time 20% · tool calls 84%
+```
+
+| | this campaign (sonnet, 3-turn) | README (Opus 4.8, 1-question) |
+|---|---|---|
+| cost saved | **24%** | 60% |
+| tokens saved | **23%** | 69% |
+| time saved | **20%** | 20% |
+| tool calls saved | **84%** | 89% |
+
+Tool-call reduction and wall-clock survive the regime change almost intact; the
+cost and token savings roughly halve. Two repos invert outright — **vscode
+processes 98% *more* tokens with codegraph** (8 calls of dense source against a
+without-arm that mostly greps), and **django costs 17% more**. Neither is hidden
+here. The with-arm is also not read-free in this regime: 4 of 28 with-arm
+sessions still touched Read (vscode run4 `rd5 bs7`, tokio run2 `rd3 bs2`,
+django run4 `rd1`, alamofire run2 `rd1`), against the README's "zero file reads
+on all seven repos" under Opus.
+
+### Contamination gate: clean, and the channel is real
+
+**0 CLI calls returned output in any of the 56 sessions.** The aggregate is
+uncontaminated and no run was dropped.
+
+But **29 attempts were blocked** — 26 in the without-arm (in **26 of its 28
+sessions**) and 3 in the with-arm. Ninety-three percent of without-arm sessions
+tried to reach codegraph through Bash and were stopped by the sanitized PATH +
+PreToolUse hook (`no-cli-shim.sh`). That is not a hypothetical channel the
+harness guards out of caution; it is the agent's *default* move once it notices
+`.codegraph/` in the tree. **`no-cli-shim.sh` is load-bearing** — without it this
+campaign would have been codegraph-over-CLI vs codegraph-over-MCP, exactly as the
+earlier 14-of-15 pass was (see the section above). Check the contamination row
+before believing any number from this harness.
+
+### Secondary readings — absolute, not before/after
+
+There is **no baseline-build arm in this campaign** — every number below is the
+current build's absolute reading on these questions. Allocation efficiency in
+particular is *relative* (attribution is by citation): it compares builds on the
+same question and says nothing on its own about waste. For a before/after
+allocation A/B see [`explore-allocation-ab-1500.md`](explore-allocation-ab-1500.md).
+
+```
+repo        calls  again    read-ret  read-miss  grep    MOVED ON   alloc eff  envelope
+vscode      26     21 81%   0  0%     1  4%      1  4%   3 12%      63.2%      442k
+excalidraw  18     14 78%   0  0%     0  0%      0  0%   4 22%      94.7%      338k
+django      11      6 55%   1  9%     0  0%      0  0%   4 36%      96.2%      186k
+tokio       16     12 75%   0  0%     0  0%      1  6%   3 19%      92.9%      343k
+okhttp       8      4 50%   0  0%     0  0%      0  0%   4 50%      97.9%      151k
+gin          8      4 50%   0  0%     0  0%      0  0%   4 50%      99.0%      116k
+alamofire   23     19 83%   1  4%     0  0%      0  0%   3 13%      89.0%      306k
+
+POOLED (110 answered explore calls):
+  explore again 73% · Read a file we returned 2% · Read a file we did NOT return 1%
+  · Grep/Glob 2% · moved on / answered 23%
+POOLED allocation efficiency: 86.7% over 110 calls / 1.9M chars
+```
+
+- **Allocation efficiency 86.7%** pooled. vscode is the outlier at 63.2% — the
+  largest envelope (442k chars) and the lowest citation share, which is where an
+  allocation change would show up first.
+- **`Read a file we returned` = 2%** (2 of 110). Right file, wrong bytes is
+  nearly absent; the allocation misses this metric was built to catch are not
+  what is driving vscode's number.
+- **Recall misses** are 3% total (1 read-miss, 2 grep).
+- **`explore again` = 73%** and is **ambiguous by construction** — it is
+  indistinguishable between "the first call was insufficient" and "the agent is
+  working through a 3-turn session and this is turn 2's first call." In a 3-turn
+  regime that ambiguity is much larger than it was single-turn; treat the
+  high-`again` repos (alamofire 83%, vscode 81%) as unresolved, not as failures.
 
 ---
 
 ## What this settles, and what it does not
 
-**Settled.** The metric exists, it is measured rather than estimated, it runs over
-multi-turn sessions — the regime where occupancy is actually charged — and there
-is a baseline across the 7 README repos to compare future changes against.
+**Settled.** The metric exists, it is measured rather than estimated, and it runs
+over multi-turn sessions — the regime where occupancy is actually charged. As of
+2026-08-05 there is a baseline across the 7 README repos (above) to compare
+future changes against, and it says codegraph's residual is **higher**, on every
+repo. (Before that campaign this section claimed such a baseline existed when it
+did not; it does now, and it is one regime — `claude-sonnet-5`, 3 turns — not a
+general result.)
 
 **Not settled, and deliberately not claimed:**
 
+- **The README's efficiency figures.** The baseline above ran sonnet / 3 turns;
+  the README published Opus 4.8 / single-question. The gap between 24/23/20/84
+  and 60/69/20/89 is regime, not regression, and this campaign cannot tell you
+  which way the published numbers have moved. That needs a matched **Opus 4.8,
+  single-question** rerun. Out of scope here, and `README.md` was deliberately
+  left untouched.
 - **A different host.** The reporter was in Cursor. We measure Claude Code.
   Window size, system prompt, and compaction policy all differ, so the *share*
   numbers do not transfer host to host; the ratio between the arms is the part
@@ -203,3 +370,45 @@ is a baseline across the 7 README repos to compare future changes against.
   *enough* — that is [explore sufficiency](explore-sufficiency.md), which every
   run now prints alongside this block — nor about how much of the returned bytes
   the answer actually used (CG-9).
+
+---
+
+## Proposed README wording — for the maintainer, not applied
+
+`README.md` is **deliberately untouched by this work.** Its benchmark table is
+Opus 4.8 / single-question and nothing measured here can restate it. What follows
+is a *proposal*: the occupancy finding as an honest counterweight to the
+efficiency table, phrased so it does not depend on the sonnet-vs-Opus regime for
+its claim. Accept, reject, or rewrite — this is not a pending edit.
+
+Suggested placement: immediately after the "A note on cost" paragraph (README
+line ~195), as a second `>` note under the same table.
+
+> **A note on context.** The efficiency table above measures *throughput* —
+> tokens processed, tools called, dollars spent to reach one answer. It does not
+> measure what is still sitting in the window afterward, and on that axis
+> CodeGraph costs more, not less. Across the same seven repos in multi-turn
+> sessions, CodeGraph's responses leave **~80% more retrieval context resident**
+> at the end of a session than the file-reading agent's do — on VS Code, 67k
+> tokens against 18k. The mechanism is the same one that makes it fast:
+> CodeGraph returns one dense, verbatim payload that answers the question and
+> then stays in the window, where a grep-and-read agent churns many small results
+> that get evicted. Fewer tokens *processed* and a larger persistent *footprint*
+> are both real. If you are running long sessions in a small window, budget for
+> it. Measured, per-repo:
+> [`docs/benchmarks/residual-context-occupancy.md`](docs/benchmarks/residual-context-occupancy.md).
+
+Three notes on the drafting, if it gets edited:
+
+1. **No percentages from this campaign's efficiency table appear in it.** "~80%
+   more resident" is the occupancy ratio between arms, which is the part that
+   travels across models and hosts; the 24/23/20/84 throughput figures are
+   sonnet-3-turn-specific and must not go near the README.
+2. **It concedes the point rather than framing it as a feature.** That is
+   deliberate — `CLAUDE.md`'s "honesty in the product is load-bearing" applies to
+   the README before it applies to a product screen, and a reader who hits #1500
+   in their own session and finds the README silent on it trusts nothing else in
+   the table.
+3. **The share numbers (33.7% of a 200k window) are Claude Code's** and do not
+   transfer to another host, so the draft quotes absolute tokens and the arm
+   ratio only.