For about a year, using AI well meant making two decisions: pick a model, then decide whether to switch on "thinking" — the mode where the model reasons through a problem before it answers.
That advice is now wrong, in a way you can check in about ten seconds. On Claude Sonnet 5 and Opus 5, asking for extended thinking manually does not get you deeper reasoning. It gets you an error.
What actually changed
Anthropic is blunt about it in its own documentation. Sonnet 5 runs with adaptive thinking on by default, and manual extended thinking is removed — it returns a 400 error, the same as on Opus 4.8 and Opus 4.7. The same release also stopped accepting temperature, top_p and top_k. The knobs did not get better. They got taken away.
What replaced them is a single parameter: effort, with five levels — low, medium, high, xhigh, max. High is the default, and setting high is exactly identical to setting nothing at all.
Turn AI into Your Income Engine
Ready to transform artificial intelligence from a buzzword into your personal revenue generator?
HubSpot’s groundbreaking guide "200+ AI-Powered Income Ideas" is your gateway to financial innovation in the digital age.
Inside you'll discover:
A curated collection of 200+ profitable opportunities spanning content creation, e-commerce, gaming, and emerging digital markets—each vetted for real-world potential
Step-by-step implementation guides designed for beginners, making AI accessible regardless of your technical background
Cutting-edge strategies aligned with current market trends, ensuring your ventures stay ahead of the curve
Download your guide today and unlock a future where artificial intelligence powers your success. Your next income stream is waiting.
The part worth understanding is what effort touches. It is not a thinking dial. It affects all tokens in a response: the answer, the thinking, and the tool calls. Lower the effort and the model does not just think less — it makes fewer tool calls, skips the preamble, gets terser. Raise it and it explains its plan first, calls more tools, and writes you a summary of what it did.
It is a suggestion rather than a ceiling. The docs describe effort as a behavioral signal, not a strict token budget: at low effort on a genuinely hard problem, Claude still thinks, just less than it would at high. You cannot accidentally switch off its brain.
The difference nobody mentions
Both major labs now use the same five words. They disagree about where to start.
Claude Sonnet 5 and Opus 5 default to high.
OpenAI's GPT-5.6 defaults to medium — and adds a none level below low, which Claude does not offer.
That is not trivia. "Just use the defaults" buys you meaningfully more deliberation on Claude than on ChatGPT. And if you have been comparing the two head-to-head on default settings, you have not been comparing like with like.
The framework
Model tier tracks how much judgment the task needs. Effort tracks how expensive it is to be wrong.
Those are independent, and collapsing them into a single "how hard is this?" question is the mistake almost everyone makes. The genuinely useful question is a different one: if this answer were subtly wrong, when would I find out?
Immediately — code that either compiles or does not, a number you will sanity-check, copy you are going to read line by line anyway. Low or medium effort. You are the verification step, so do not pay a model to be one too.
Eventually — a report someone acts on next week, an analysis feeding a decision. High effort, the default. Leave it alone.
Maybe never — architecture you will build on for months, a ten-step chain where step seven quietly depends on step three, anything where you are the only reviewer and might not catch a subtle error. Xhigh or max, on the best tier you have.
That last one is where people systematically underspend, for an uncomfortable reason: a confidently wrong answer looks exactly like a right one.
Two footnotes that save real time. On Opus 5, effort controls thinking volume but not visible response length — if you want shorter answers, ask for shorter answers. And if you rely on prompt caching, pick an effort level at the start of a conversation and keep it there, because changing it mid-thread throws the cache away.
The honest shortcut
Most people need none of this. Run everything on your mid-tier model at the default, and escalate only when an answer actually disappoints you. Selection works far better as a reaction to a bad result than as a prediction you make in advance — the minutes spent staring at the dropdown usually cost more than simply re-running the task one tier up.
The tell that this shift is real and not cosmetic: Anthropic now describes Claude Fable 5's lower effort settings as often exceeding xhigh performance on previous models. The floor keeps rising underneath everyone. Which is the actual lesson here — the habit that wins was never memorizing this year's settings, it is noticing when the question itself has changed.
The full version on the site adds the side-by-side table of every control both labs expose, an interactive picker that names the tier and effort level for your specific task, and a free copy-paste prompt — The Effort Audit — that interviews you about one recurring job and tells you exactly where you are overpaying.
Read less. Know more.
Morning Brew delivers the biggest stories in business, finance, and tech in about 5 minutes — with just enough personality to keep things interesting.
Join 4,000,000+ professionals who start their mornings a little smarter.



