Wispr Flow turns your voice into clean, ready-to-send writing — speak naturally, it strips the filler and fixes the punctuation. I've used it daily since February 2026 to build this newsletter. Read the full review →
Anthropic released Claude Opus 4.7 on 16 April 2026 and Claude Opus 4.8 on 28 May 2026. Two flagship models, 42 days apart.
Coverage of that second launch put the gap at 41 days. It is 42 — April has 30 days, so 14 remain after the 16th, and May contributes 28. Nobody was harmed by the missing day, but it is a fair preview of how carefully the rest of the numbers around these releases have been handled.
The number that isn't there
The most-quoted benchmark figure for Opus 4.8 is a precise percentage on SWE-bench Verified, repeated confidently across write-ups. It does not appear anywhere on Anthropic's announcement page for Opus 4.8. Not in the table, not in a footnote. We checked the page directly for the term.
What the page does carry, in its running text, is four benchmark claims — a Super-Agent benchmark, CursorBench, a Legal Agent Benchmark, and 84% on Online-Mind2Web. Every one of them sits inside a testimonial from a named customer, not in Anthropic's own account of what it measured.
That is not a scandal. But a number that reached you through one customer's evaluation harness is a different class of evidence from one a lab measured and published with its methodology attached.
Watch on demand: Tabs + PwC break down how finance teams are operationalizing usage-based pricing — without the manual overhead.
What actually changed
Anthropic's own release note describes Opus 4.8 as having "the same set of tools and platform features as Claude Opus 4.7." Both models list at $5 per million input tokens and $25 per million output. Both carry a 1M-token context window, 128k maximum output, and a January 2026 knowledge cutoff.
The real differences are in the plumbing, and there are two of them:
Tool calls got cheaper to set up. Every request including tools carries a system prompt Anthropic adds automatically. On Opus 4.7 it was 675 tokens with a tool choice of auto or none; on Opus 4.8 it is 290. That is 385 fewer input tokens on every tool-using request — a 57% cut in that overhead, before your own prompt is counted.
The prompt-caching floor dropped to 1,024 tokens on Opus 4.8, down from 2,048 on Opus 4.7 — both figures sit in Anthropic's prompt-caching documentation. Prompts too short to cache before became cacheable, and a cache hit costs 10% of the standard input price.
That is the whole list. Two further differences circulated widely, and neither holds up. Adaptive thinking is not a 4.8 feature — Anthropic's thinking documentation lists it for Opus 4.7, Opus 4.6 and Sonnet 4.6 as well, and on all of them thinking stays off until you set the thinking type to adaptive.
And effort defaulting to high is not a change — Anthropic's effort documentation gives the API default as high for Opus 4.7 in the same words it uses for 4.8, and the Opus 4.8 announcement itself says that effort level spends a similar number of tokens as Opus 4.7's default, with better performance. What is genuinely new is a user-facing control: alongside the model selector on claude.ai and in Cowork, people can now choose how much effort Claude puts into a response.
We nearly printed both of those as fact. They reached us the way that SWE-bench percentage does — through coverage, confidently, sourced to nobody. One of them was written down here as a direct quotation that appears in no Anthropic document we can find. Which is this edition's own lesson, arriving uninvited.
The trap in "the price didn't change"
Anthropic's pricing docs note that Claude 4.7 and later models use a newer tokenizer producing "approximately 30% more tokens for the same text." Opus 4.6, 4.7 and 4.8 all list at $5 and $25 — a flat line across three releases. But the document that cost a million tokens on 4.6 costs roughly 1.3 million on 4.7.
That step happened at 4.7, not 4.8, so it is not a difference between these two models. It is a very good illustration that "the price didn't change" and "the bill didn't change" are different sentences — and that a benchmark chart would never have told you either way.
Three things worth ten minutes
Check whether you set effort explicitly; if not, you are running a spending default you never chose. Revisit any decision that your prompts were too short to cache — that was reached against a higher minimum. And measure tokens rather than documents, because any cost model built before Opus 4.7 understates what the same work costs today.
One last thing. Both of these models are now filed under "Legacy models" in Anthropic's own documentation. Opus 4.7 held the flagship slot for 42 days. Twelve days after Opus 4.8 arrived, Claude Fable 5 landed; fifty-seven days after it, Claude Opus 5.
In a cycle moving this fast, a benchmark chart is a photograph of a single afternoon — and half the ones you will be shown were taken by somebody else. The release notes and pricing tables are dull and unglamorous. They are also the part that tells you what it costs to live there.
The site version adds the full dated timeline, a proof table checking four claims from the coverage against Anthropic's documentation, and an interactive calculator that works out what that 385-token difference is worth at your own request volume.

