The 30th edition of Quanta Bits looks at a crowded week for frontier AI. Anthropic released Claude Fable 5.1 and Mythos 5.1, OpenAI released GPT-6 Astra, and Google introduced Gemini 3.8 Flash. The vendors reported large gains, but the practical question is where each model fits in an operating system of tools, costs, controls, and human judgment.
Two Frontier Models in One Week - The main feature. Anthropic reports substantial improvements for Fable 5.1 on multistep computer tasks and business workflows, along with lower costs for reusing previously processed material. Those API economics do not automatically make subscriptions cheaper or expand included usage. Access rules, weekly allowances, and paid credits still determine what users actually spend.
OpenAI reports a slightly higher result on the same computer-task benchmark and says Astra can handle a much broader range of computer work. It also rated the model critical for cybersecurity under its own framework and deployed restrictions around the most advanced capabilities. Its safety report adds an important complication: Astra's written explanations are becoming less useful for spotting unwanted behavior, even when monitoring its actions and outputs still provides signals.
The headline results have not yet been independently replicated. Prices, plan rules, and vendors' stated safety classifications are verifiable; improved reliability and total cost still need to be measured on the work a company actually performs. Early use suggests both models are better than their predecessors, particularly for computer tasks, while also consuming subscription allowances faster.
For operations leaders, the evaluation should include what happens when one agent delegates to several others, which tasks justify a more expensive model, whether the result can be tested or otherwise verified, and how much human correction remains. A spending limit enforced by the workflow is more dependable than an instruction asking the model to stay within a budget.
Also in this issue:
- The Brief - New models promise harder work with less help, delegated agents can increase total AI spending, useful work matters more than usage, financial authorities and insurers are preparing for AI-assisted attacks, and blind benchmark testing may reduce question contamination.
- Patterns & Signals - Companies are connecting AI spending to defined jobs and measurable outcomes, while security planning is expanding from conventional attacks to agent permissions, stolen sessions, and credential exposure.
- Signals to Watch - Nvidia agreed to buy Hugging Face, AT&T reports reducing coding costs by routing simpler work to cheaper models, and Google DeepMind is piloting protected double-blind model evaluations.
- Interstitials / Overheard - A joke about disruptions at Grok, Codex, and Claude asks whether all the systems are secretly running on the same server.
- Meanwhile... - Researchers improved vision in aged mice by replacing a fatty acid whose production declines with age, though a human treatment remains far away.
- What I'm Consuming - Managing parallel AI agents, the hidden work of supervising them, digital sovereignty, Ramp's internal coding agent, the week's model stack, and an AI-assisted product-management workflow.
- Term of the Week - Credential proxying, where a service holds an outside system's access token and attaches it to an agent's request without exposing the credential directly to the agent.
- After Hours - Fiasco: The Battle for Boston, an audio documentary about school segregation, court-ordered busing, and the city's political history.
The right response to a stronger model is a controlled comparison, not a wholesale switch. Give competing models the same recurring job, measure result quality, review and correction time, and total cost, then decide which model deserves that responsibility.