It was an exciting week. In a single day we got Anthropic's Opus 5.5 and OpenAI's GPT-6 Sol and Luna. Opus 5.5 has become my default for daily work. The rest of the week's news ran on two tracks: AI agents going off-script on routine tasks, and a push to make AI cheaper and easier to use.
AI agents are going off-script on routine tasks - OpenAI told dozens of organizations, including governments and universities, that its AI agents may have bypassed their security controls or disrupted their services during its own testing and training. An agent looking up Australian medicine spending got around blocks on a government statistics site, and another pulled US Census data using login credentials it found online. These agents were doing ordinary data lookups, not hacking tests. Google and Anthropic have reported similar incidents in their own evaluations, and Australia's cyber security agency issued an alert recommending strong access controls, log monitoring, and incident response testing for AI agent scenarios.
AI models got smarter and cheaper in the same week - Anthropic's Claude Opus 5.5 costs about 40% less to run than its predecessor in Anthropic's tests and tops Artificial Analysis's independent ranking of model intelligence. OpenAI's GPT-6 Sol and Luna cost half as much as the versions they replace. Jev, a new low-cost decision model from TypeSafe AI, returns probabilities for answer options you supply, at a fraction of the cost of leading chat models. Cheaper AI tends to expand what companies automate, so whether the total bill shrinks depends on which work goes to which model.
Spread thin, AI spending doesn't show up in the P&L - McKinsey studied 20 companies that consistently turned AI into financial results. Two-thirds focused on three business areas or fewer, chosen because a small improvement there shows up in the P&L. The risk for most companies isn't experimentation itself but experimentation without focus or follow-through: pick the one or two areas where a 5% to 10% improvement would show in results, point the experiments there, and turn what repeats into shared practice.
Meta's Muse is easy to set up, and hungry for data - Meta's personal AI agent reached number one on Apple's US App Store by needing almost no setup. WIRED found data collection for AI training on by default, and Amazon blocked Muse from shopping on its site over how it identifies itself and handles logins.
Also in this issue:
- The Brief - Agents going off-script, same-day price cuts, Jev for small decisions, McKinsey on focus, and Muse's rise and data concerns.
- In brief - A federal appeals court sided with the Pentagon against Anthropic; AI labs report new scientific results; a test found chatbots still get personal-finance questions wrong; a chatbot-planned climb on Mount Shasta ended in a rescue.
- Interstitials / Overheard - A Claude time-zone exchange that didn't inspire confidence.
- Meanwhile... - Two severely paralyzed people used a brain implant to communicate through an on-screen avatar using words and gestures at the same time.
- What I'm Consuming - A guide to getting the most from Opus 5.5, a hands-on Jev demo, MIT Technology Review on AI's trillion-dollar infrastructure gamble, HBR on AI agents doing business with each other, and the FT on Adam Smith versus Schumpeter on job creation.
- After Hours - Wes Anderson's The Phoenician Scheme and Barbara Tuchman's The Zimmermann Telegram.