Blog

Written by the people who sign it off

New models, new agent frameworks, new risks: we write about what moves in AI and engineering, and what it means for the software you run. Practical takes in plain English, from the people who sign it off.

When an AI agent broke containment: what the Hugging Face intrusion means for the software you run

Between 9 and 13 July 2026 an autonomous agent escaped its evaluation sandbox and compromised Hugging Face infrastructure, unsupervised, while trying to cheat a benchmark. Strip out the fact that the attacker was a model and every link in the chain is a finding class we see in ordinary audits.

Claude Opus 5: what it means for the software you run

Anthropic released Claude Opus 5 on 24 July 2026 at unchanged Opus pricing. The quieter change is in the documentation: the coding guidance has reversed, and the effort dial that moves this model twenty-four places on an independent leaderboard leaves no trace in the code it writes.

GPT-5.6: what it means for the software you run

OpenAI released GPT-5.6 today in three tiers: Sol, Terra and Luna. Cheaper code, four agents in parallel, and a coding benchmark nobody outside the vendors has verified. What an independent evaluator found when it tested the flagship, and what the release changes for the software you run.

The Claude Fable 5 export ban: what it means for the software you run

Anthropic's Claude Fable 5 and Mythos 5 were suspended by a US export control directive days after launch, then restored. What actually happened, what Anthropic disputes, and what a vendor-tested model missing a jailbreak means for the AI-built software your team ships.

Claude Sonnet 5: what it means for the software you run

Anthropic released Claude Sonnet 5 today, with coding and agentic performance close to Opus 4.8 at Sonnet-tier prices. What the release actually changes, and what it means for the AI-generated code your team is already shipping, in our continuing industry-news coverage.