An estimated 82% of migrations off deprecated models happened after the shutdown date, once the app was already failing. 94% of the apps had the model ID hard-coded.
matters
AI apps age like milk, not wine, and most teams don't even keep the expiry date on the fridge.
When a question states the rules, the top four models score 22 to 25 out of 27. When the answer hinges on a hidden fact buried in contradictory records, four of six models get 6 of 24 or fewer.
Strict gate success: 94.98% for Gemini 3.8 Flash, 83.29% for GPT-5.6 Luna, 74.18% for DeepSeek v4.1 Flash. A deterministic baseline passed all 1,700 gates.
matters
"Automate 95%" sounds like savings until someone has to hunt for the 5%, forever.
When the agent could read back whether a write had happened, instructed frontier models duplicated a lost-ack write only 0.5% of the time. Without a read-back: 56 to 74%. Idempotency keys on every write cut duplication from 28% to 4%.
Of 111 approval/trace pairs, in 40 the fields a human sees left effects uncovered. The authors prove no record-only policy can be both safe and permissive.
Providing context files "does not generally improve task success rates, while increasing inference cost by over 20% on average." Repository overviews, the thing model providers keep recommending, were "not helpful."