The Mumbai dabbawala carries one tiffin to your desk, not your mother's entire kitchen, and your coding agent deserves the same courtesy.
I've watched this movie from close up. On a customer project, the team was under pressure to show the client that AI was speeding things up, so it built a prompt library. What went in was decided by gut feel: someone liked a prompt, in it went. The only yardstick was story points, which barely moved, and story points were never going to tell anyone much about AI gains. That's like judging the dabbawala by the weight of the tiffin.
One prompt in that library did work: an accessibility one. It was long, but not one line of it was a tour of the codebase. Every line was a rule the agent couldn't have guessed from the code:
We all do this. The agent slips, so we add a paragraph, then a directory tour, then the full history of the monolith. The file grows like the loft in a Mumbai flat: everything goes up, nothing comes down.
Here's the number that should stop you. A team at ETH Zurich tested coding agents on SWE-bench tasks and on real repos with context files their own developers had committed. Providing those files "does not generally improve task success rates, while increasing inference cost by over 20% on average." Repository overviews, the thing model providers keep recommending, were "not helpful." The instructions inside the files? Those were followed well. That's my accessibility prompt in one sentence.
Extra inference cost from context files, with no general gain in task success.
That study is from February, and everyone keeps ignoring it. This fortnight Cursor says it trimmed about 66% of its system prompt and cut total tokens by 7% "without degrading quality" (self-reported, method not shown). And a new enterprise benchmark shows where the real gap sits. When a question states the rules, the top four models score 22 to 25 out of 27. When the answer hinges on a hidden fact buried in contradictory records, four of six models get 6 of 24 or fewer (Era by Eon, arXiv:2609.30055, 24 Sep).
So the context worth writing is the stuff the code can't tell the agent: the rule nobody wrote down, the gotcha from that outage in 2019. Not a tour of src/ the agent can read for itself in two seconds.
The fair pushback comes from context-driven teams (Tessl makes the case well): shared context files are where workflow knowledge should live. Agreed on where. Not on how much.
What would change my mind: a team that A/B tested their AGENTS.md on private repos and showed overviews raise success at equal or lower cost. Send it to me.
Measure it, trim it, measure again. Then tell me your context file is helping.
When the Model Retires: your AI stack is already legacy
Eight out of ten teams found out their AI model had retired the way you find out your phone plan expired: when the call drops.
A new study mined 22,555 commits across 17,703 GitHub repos that moved off deprecated OpenAI, Anthropic and Google models. An estimated 82% of those migrations happened after the shutdown date, once the app was already failing. 94% of the apps had the model ID hard-coded. Notice length mattered a lot: 89% migrated after shutdown under Anthropic's 60 to 114-day notices, against 13% under OpenAI's one-year notice. Only 8% switched provider.
Why you should care: your shiny agent stack is tomorrow's legacy estate, and it's rotting faster than your COBOL. Put model IDs in config, with an owner and a retirement date.
Verdict: AI apps age like milk, not wine, and most teams don't even keep the expiry date on the fridge.
The Last Human Gate: when automating governance adds work
The AI passed 95% of governance checks. A plain old rules script passed 100%, and nobody put that on a slide.
This paper turns enterprise review gates into executable contracts and runs 899 model runs over 300 synthetic projects. Strict gate success: 94.98% for Gemini 3.8 Flash, 83.29% for GPT-5.6 Luna, 74.18% for DeepSeek v4.1 Flash. A deterministic baseline passed all 1,700 gates. The authors warn that automating most cases "may paradoxically increase total labor" once you count exceptions and upkeep. The projects are synthetic, so read it as a strong signal, not a verdict on your org.
Why you should care: before you put an LLM on a review gate, check whether a boring rules engine already does the job perfectly, and count the human hours for the 5% it gets wrong.
Verdict: "Automate 95%" sounds like savings until someone has to hunt for the 5%, forever.
Your Ticket Is the Prompt: smell it before the agent eats it
We used to write vague tickets because a good developer would walk over and ask. The agent won't walk over. It'll just guess, confidently, at 2 a.m.
I've seen this one land. On a project, a story said a new section should be "identical" to an existing section of the UI. The business meant: make it look the same. What got built was identical everything, screen to backend, because "identical" is a very honest word. The twist: an agent wrote that story, and developers paired with agents built it. The question "identical how?" only came up when the rework did.
That ticket wasn't a conversation starter anymore. It was the prompt.
A new paper injected requirement "smells" (vague words, tangled structure, contradictions) into the specs for four applications and had LLMs write the code. More smells went with lower functional correctness, measured by test suites. Cleaning them up helped but didn't guarantee correct code (Villamizar et al., arXiv:2609.29208, 24 Sep). The abstract gives no headline percentage, so I won't invent one. Laurie Voss puts it plainly: "The bottleneck moves to the description of the problem."
So here's the lint I'd run on every ticket an agent touches. It takes ten minutes.
- Lexical pass (words). Hunt vague terms (fast, simple, relevant, user-friendly, etc.), weak modals (should, may, might) and comparisons with no baseline ("faster than before").
- Syntactic pass (shape). Find passive voice that hides the actor ("the order is cancelled" by whom?), dangling references ("as discussed", "see above") and missing conditions (when? for which users?).
- Semantic pass (meaning). Look for contradictions, missing error and edge cases, and anything you can't test.
- Add three things. Two or three concrete examples in Given/When/Then, explicit non-goals, and one "must never happen" line.
- Write one failing test first. Turn one example into a test before the agent starts. As Frisinger puts it, "The prose spec is a claim. The test is a receipt."
- Measure. Count smells per ticket at refinement, then compare first-pass agent PR acceptance for linted and unlinted tickets. Two sprints will show you a trend.
Run that on my "identical" story. Step 2 flags a reference with no boundary. Step 4 adds the one line that would have saved the rework: "Non-goal: the backend."
"The system should quickly show relevant offers to users on the checkout page."
"On checkout, show up to 3 offers the logged-in customer is eligible for, ranked by discount value. Given a cart of ₹1,200 and offers A (10% off over ₹1,000) and B (flat ₹200 off over ₹1,500), show A only. If the offers service fails, show checkout with no offer banner. Non-goal: personalised ranking. Must never: show an offer the customer can't redeem."
The tracker, one row per ticket: ticket ID, lexical smells, syntactic smells, semantic smells, examples added (Y/N), non-goals (Y/N), must-never line (Y/N), failing test first (Y/N), agent PR accepted first pass (Y/N), rework commits.
What I'd do: lint every ticket an agent touches and write one failing test before it starts. What I wouldn't: buy a "spec generator" to write the ticket for me. My "identical" story came out of an agent. A smelly ticket written faster is still smelly.
The agent can't read your mind. Stop making it try.
Don't Hit Pay Twice: exactly-once writes for agents that retry
Every Indian who's used UPI knows the rule. The payment says "pending", your thumb hovers over Pay again, and you stop, because you check the UTR first. Agents don't know that rule. They hit Pay again.
A new study built a sandbox called LIMBO with six services and 12 injected faults (late commits, redelivered messages, partial batches) and ran 25,930 episodes across nine models and three production harnesses. When the agent could read back whether a write had happened, instructed frontier models duplicated a lost-ack write only 0.5% of the time. Without a read-back: 56 to 74%. Idempotency keys on every write cut duplication from 28% to 4%.
Read that again. The fix isn't a smarter model. It's the tool contract.
- Classify every tool by effect. Read-only, idempotent write (a PUT with full state), or non-idempotent write (charge, send, deploy, create). MCP's tool
annotationsfield is one place to declare it. - Mint the idempotency key in the harness, not the model. Derive it from stable IDs:
hash(run_id, step_id, tool, canonical_args). The model never sees it, so a retry reuses the same key. - Send the key on every non-idempotent write. Services that honour keys (Stripe's
Idempotency-Keyheader, the IETF httpapi draft) return the original result on a replay instead of acting twice. - On a timeout or 5xx, read back before you retry. Ask the service: does a charge exist with this key?
- No read path? Don't retry. Escalate. Park the step as "unknown outcome" and hand it to a human or a reconciler.
- Bound the in-flight window. Under late commits, the paper shows no verify-only policy guarantees exactly-once unless you wait out a known maximum commit delay before deciding the write failed.
- Keep an effect ledger. Record
(key, tool, args_hash, status)before and after each call, and reconcile it against service state at the end of the run.
# Illustrative sketch: rename to fit your harness import hashlib, json, time def idem_key(run_id, step_id, tool_name, args): blob = json.dumps({"r": run_id, "s": step_id, "t": tool_name, "a": args}, sort_keys=True) return hashlib.sha256(blob.encode()).hexdigest()[:32] def safe_write(tool, args, ctx, max_commit_delay=30, attempts=3): key = idem_key(ctx.run_id, ctx.step_id, tool.name, args) # args frozen at first attempt for _ in range(attempts): ledger.record(key, tool.name, "requested") try: result = tool.call(args, idempotency_key=key) ledger.record(key, tool.name, "confirmed") return result except (Timeout, ServerError): if not tool.has_read_path: ledger.record(key, tool.name, "unknown") raise NeedsHuman(f"{tool.name} outcome unknown, key={key}") time.sleep(max_commit_delay) # wait out late commits found = tool.lookup(idempotency_key=key) if found: ledger.record(key, tool.name, "confirmed") return found # nothing landed inside the bound: safe to replay with the same key ledger.record(key, tool.name, "unknown") raise NeedsHuman(f"{tool.name} failed {attempts} times, key={key}")
- Keys only help if the service honours them. Plenty of internal APIs don't. Put a dedupe table at the gateway.
- Canonicalise args carefully. A timestamp or a float formatted differently makes every retry a "new" key.
- Keys expire. Providers drop them after a window, so a very late retry becomes a fresh write.
- Freeze args at the first attempt. If the model rephrases the args on retry, the hash changes and the key is useless.
- Waiting costs latency. The in-flight bound slows the unhappy path. That's the price of not charging twice.
- Batches need per-item keys, not one key for the whole batch.
When I'd use it: the day an agent gets a tool that spends money, sends a message or deploys. When I wouldn't: read-only tools. Making grep idempotent is solving a problem you don't have.
Check the UTR before you pay again.
"Don't worry, the agent asks before it runs anything. A human approves every command."
Evidence for- Approval gates are standard in the major agents and getting stronger. GitHub added "proof of presence" for high-impact actions on 24 Sep (changelog).
- An "approval laundering" paper checked 111 approval/trace pairs. In 40, the fields a human sees left effects uncovered. The authors prove no record-only policy can be both safe and permissive (arXiv:2609.28586, 23 Sep).
- When agents can't verify a write, frontier models duplicate it 56 to 74% of the time. A click at the start doesn't cover the retry at the end.
- Vendors are shipping sandboxes: GitHub brought local sandboxing to the Copilot app on 23 Sep. That's the quiet admission.
The human approves the command; the damage is in what the command sets off. Sandboxes and effect checks keep you safe. The click keeps you busy.
1,000 PRs in a week: five questions the headline didn't answer
A thousand PRs in a week is a great headline. Let's ask what it left out.
Kiro's team says it merged 1,000 PRs in 7 days, with 3 engineers building the system and 50+ concurrent agent sessions (self-reported, Kiro, 18 Sep). The write-up on how they got there, from a single session to a supervisory "Crew" agent, is generous and worth your time. What it doesn't disclose:
- What share got rolled back?
- What was the median review time?
- How many were one-line dependency bumps?
- How many defects reached users?
- What did the 3 engineers stop doing?
This isn't a knock on Kiro. It's a knock on all of us for clapping at the number before asking for the denominator.
Open your AGENTS.md or system prompt right now. What's one line you'd delete, and what's the one line you'd fight to keep?
Measured an AI rollout properly and found something nobody wanted to hear? Reply with the idea. The best ones run here, with your name on them.
On LinkedIn, the algorithm decides if you see the next tiffin. On Kit, it just arrives.
Every second Tuesday. One email. Unsubscribe in one click, no hard feelings.Ved
Views are my own, not my employer's. Issue text published under CC BY 4.0: reuse with credit.