What Landed
Nate B. Jones argues Opus 5.5 is a real efficiency step, not a benchmark bump. It is priced at $4/$20 per million input/output tokens, 20% below Opus 5, and Anthropic claims typical workloads cost about 40% less once fewer retries and tool calls are counted. GitHub, Lovable and Spotify report fewer steps for the same work. Anthropic also named instruction adherence as a top fix, after complaints that Opus 5 acknowledged instructions and then ignored them. Matthew Berman covered OpenAI Dev Day:
- Dots, an always-on agent with standing access to email, docs and calendar.
- GPT-6.1 Soul at $2/$10.
- Ultrafast, a Cerebras-accelerated tier that costs 6x more.
- Codex moving fully to cloud execution.
Why It Matters
Jones's lasting point is methodological. Price per token misleads because retries, file re-reads and correction cycles drive real cost. The 40% figure is an Anthropic claim backed by launch partners, so it has to be checked against each customer's own workloads. OpenAI's pricing now runs on two curves: cheap models keep getting cheaper while fast inference gets more expensive. Its new $500 Pro tier also gives worse marginal value than the $200 tier (25x base limits versus 10x), so blanket upgrades need a workload case. The enterprise issue with Dots is governance more than capability. It has standing access to mail and calendar, and usage billing starts as soon as it drives a ChatGPT or Codex thread.
Worth Raising With Customers
- Build a standing internal cost-per-task benchmark on real recurring workloads, rerun it on every release, and log failures as carefully as successes. With flagship releases now about 18 days apart, rate-card comparisons are stale before procurement closes.
- Every overnight or long-running agent needs explicit stop conditions and a defined done state. Without them, the model keeps going and token spend has no ceiling.
- Judge ambient agents (Dots and its peers) first on how much data they can reach and how their billing is tied to other products, then on capability.