The numbers ledger
Every number, with its receipt
Every number I've cited, with the primary source, the date and how I checked it. ✓ means I read it at the source. Self-reported means a vendor said so and nobody audited it. Reuse any of it under CC BY 4.0. Just say where you got it.
Download: numbers.json · numbers.csv
| Number | What it measures | Source | Checked | Track · issue |
|---|---|---|---|---|
| >20% | Extra inference cost from repository context files (AGENTS.md), with no general gain in task success | Gloaguen et al., ETH Zurich, arXiv:2602.11988 | ✓verified at source | CKContext & knowledge Issue #01 |
| ~66% | Share of Cursor's system prompt trimmed | Cursor blog | self-reported | CKContext & knowledge Issue #01 |
| 7% | Cursor's total token reduction, "without degrading quality" (method not shown) | Cursor blog | self-reported | CKContext & knowledge Issue #01 |
| 22–25 of 27 | Top-four model scores on enterprise questions where the rules are stated | Era by Eon, arXiv:2609.30055 | ✓verified at source | CKContext & knowledge Issue #01 |
| ≤6 of 24 | Scores of four of six models when the answer hinges on a hidden fact in contradictory records | Era by Eon, arXiv:2609.30055 | ✓verified at source | CKContext & knowledge Issue #01 |
| 82% | Estimated share of LLM migrations made after the model's shutdown date (22,555 commits, 17,703 GitHub repos) | Kim, arXiv:2609.31288 | ✓verified at source | UMUnderstanding & modernisation Issue #01 |
| 94% | LLM apps with the model ID hard-coded | Kim, arXiv:2609.31288 | ✓verified at source | UMUnderstanding & modernisation Issue #01 |
| 89% vs 13% | Post-shutdown migrations under Anthropic's 60 to 114-day notices vs OpenAI's one-year notice | Kim, arXiv:2609.31288 | ✓verified at source | UMUnderstanding & modernisation Issue #01 |
| 8% | LLM migrations that switched model provider | Kim, arXiv:2609.31288 | ✓verified at source | UMUnderstanding & modernisation Issue #01 |
| 94.98% | Best strict governance-gate success by an LLM (Gemini 3.8 Flash; GPT-5.6 Luna 83.29%, DeepSeek v4.1 Flash 74.18%; 300 synthetic projects, 899 runs) | Canale, arXiv:2609.29345 | ✓verified at source | POPeople & operating model Issue #01 |
| 1,700 of 1,700 | Governance gates passed by a deterministic rules baseline | Canale, arXiv:2609.29345 | ✓verified at source | POPeople & operating model Issue #01 |
| 0.5% | Lost-ack writes duplicated by instructed frontier-model agents when a read-back was available | Li, LIMBO, arXiv:2609.29095 | ✓verified at source | ABAgents that build Issue #01 |
| 56–74% | Lost-ack writes duplicated by frontier-model agents with no read-back | Li, LIMBO, arXiv:2609.29095 | ✓verified at source | ABAgents that build Issue #01 |
| 28% → 4% | Agent duplicate writes without vs with an idempotency key on every write | Li, LIMBO, arXiv:2609.29095 | ✓verified at source | ABAgents that build Issue #01 |
| 40 of 111 | Agent approval records where the fields a human sees left effects uncovered | Agent Approval Laundering, arXiv:2609.28586 | ✓verified at source | VTVerification & trust Issue #01 |
| 1,000 PRs in 7 days | PRs merged by Kiro's team with 3 engineers and 50+ concurrent agent sessions (no quality metrics disclosed) | Kiro blog | self-reported | ABAgents that build Issue #01 |
Last updated . Reuse with credit under CC BY 4.0: “The Dabbawala Protocol, thedabbawalaprotocol.com”.