How we ship production AI — retrieval, evals, infra and the occasional opinion. Written by the people building it.
AI coding tools name packages that do not exist 19.7% of the time across a 576,000-sample study — and agents install them without checking. Inside slopsquatting, the supply-chain attack that needs no vulnerability, and the five controls that stop it.
SWE-bench Verified is contaminated, OpenAI stopped reporting it, and SWE-bench Pro puts the best model at 23.3%. Why contamination inflates scores 15–20 points, what replaced the old benchmarks, and how to build an internal eval that predicts your results.
x402 crossed 165 million transactions from 69,000 active agents while AP2 mandates became Mastercard Verifiable Intent. The two protocols letting agents pay, the free-riding research nobody quotes, and how to give an agent a budget without giving it your card.
Small language models run 5 to 20x cheaper than frontier APIs, and at a million interactions a month the gap is $50k–$150k. The cascade routing architecture that captures it, the escalation signals that work, and the three ways it quietly degrades quality.
The OpenTelemetry GenAI semantic conventions turned agent traces into a vendor-neutral standard, and Claude Code, Codex, and Copilot now emit them natively. What to instrument first, what the four span types mean, and why cost attribution is the feature that gets it funded.
Four hyperscalers will spend $725B on AI capex in 2026, up 77%, while tech layoffs passed 142,000 and 45 CEOs named AI as the reason. The opex-to-capex transfer nobody calls by its name, the 275,000 unfilled AI roles, and what it means for your career.
AI governance stopped being a slide deck and became a job with a budget the week the EU AI Act's high-risk obligations took effect. What the role actually does, the first ninety days, and how to build the function without creating a committee that blocks everything.
Sovereign AI became a procurement outcome: an EU-resident foundation model under a sovereign orchestration layer, with everything else left on global cloud. Where to draw the boundary, what running two paths costs, and the four leaks nobody diagrams.
A model launch got buried by a telemetry row and a silent price increase, and the argument that followed was not really about either. Why coding agents provoke a reaction other tools do not, and the four properties that make one defensible to depend on.
One thoughtful email a month on shipping AI. No fluff, unsubscribe anytime.