Does “near Fable intelligence at half the cost” survive independent testing?
That claim comes from Anthropic’s own reported results. It is plausible, but third-party evaluations across real workloads are the only way to know how broadly it holds.
A FORWARD FUTURE FIELD GUIDE / CLAUDE OPUS 5
Yes—if careful, multi-step coding, analysis, or agent work is a meaningful part of your job.
Why: Opus 5 is a meaningful step up from 4.8 for that work at the same base API price, and its effort controls let you spend less on routine tasks without giving up deeper reasoning when a task needs it. Keep Sonnet for high-volume general work, and validate Opus against your actual workflows before a broad rollout.
01 / THE CAVEAT
That claim comes from Anthropic’s own reported results. It is plausible, but third-party evaluations across real workloads are the only way to know how broadly it holds.
Opus 5 comes close to Mythos 5 at finding vulnerabilities, according to Anthropic, but remains far behind at developing exploits. Finding and exploiting are different capabilities; coverage should not collapse them into one claim.
Anthropic says biology-related requests blocked on Fable 5 will now route to Opus 5 instead of Opus 4.8. That is a real policy change, not merely a model-spec update.
02 / THE X READ
Early reactions are not benchmarks. They are useful, however, for the details a launch chart cannot show: taste, code quality, speed, and where the model feels different in practice.
Theo on Claude Opus 5
Boris Cherny on Claude Opus 5
Elon Musk on Claude Opus 5
shirish on Claude Opus 5
ashen on Claude Opus 5
Posts selected for reach and useful first-hand perspective; views expressed are their authors’ own.
03 / THE CHANGELOG
Not a new price tier. Not a new context-window story. Opus 5 is Anthropic’s bid to make frontier-level work more adjustable—and more usable every day.
Anthropic reports state-of-the-art results on Frontier-Bench and GDPval-AA. It also says Opus 5 remains behind Mythos 5 on cybersecurity.
Use low through max effort to balance depth and token use. Fast mode is a speed option—around 2.5× faster at twice the base price—not another effort level.
Anthropic expects cyber classifiers to intervene about 85% less often than Fable 5’s, while still blocking binary-based scanning, penetration testing, and exploit generation.
Anthropic reports gains of 10.2 percentage points on internal organic-chemistry tasks and 7.7 points on protein-function prediction versus Opus 4.8.
Developers can change tools mid-conversation without invalidating prompt cache, and opt into automatic safety fallbacks on the API.
04 / PERFORMANCE
The pattern in Anthropic’s release data is clear: Opus 5 is trying to offer near-Fable performance at a lower task cost. Independent replication is still the thing to watch.
Anthropic says Opus 5 scores three times higher than the next-best model on this novel-problem-solving evaluation.
Pass rate at the same cost per task versus the next-best model, according to Anthropic.
Anthropic says Opus 5 exceeds Fable 5’s best result at just over a third of its cost.
A release-day decision guide, built to be updated as the model landscape changes.
Named evaluations in the launch: Frontier-Bench v0.1, CursorBench 3.2, AA Coding Agent Index, ARC-AGI 3, GDPval-AA v2, OSWorld 2.0, HLE, AutomationBench, and DeepSearchQA.
05 / SHOW, DON’T TELL
Anthropic shipped live artifacts with the launch. Try them here instead of taking a screenshot’s word for it.
An interactive airflow simulation. The interesting part is not just that it renders—it gives you a real control surface to explore.
Open full screen ↗A zoomable cell illustration built as an explorable artifact, a stronger test of visual explanation than a static output.
Open full screen ↗WATCH / FORWARD FUTURE VIDEO
Watch the companion video, then use this guide to compare the headline claims with the practical tradeoffs.
06 / GETTING REAL VALUE
Use low effort for drafting, classification, extraction, and well-bounded edits. Move up when the model must plan, debug across files, verify a result, or recover from ambiguity.
Anthropic says prompt caching can cut costs by up to 90%. Put stable instructions, reference material, and tools early in your request structure.
For offline enrichment, review, or transformation jobs, batch processing can reduce cost by 50%—a lever more teams should use before changing models.
Fast mode is worth considering where responsiveness changes the workflow. It is not a shortcut to better reasoning, and it costs twice the base rate.
07 / PROMPT COOKBOOK
These prompts are starting points, not magic spells. Include your constraints, source material, desired output, and how you want the work checked.
Trace this production error across this repository. First map the likely data flow and competing root-cause hypotheses. Then propose the smallest fix, add a regression test, and explain what evidence would falsify your conclusion.Plan this research task as an agent workflow. Separate facts we can verify from assumptions, define a stop condition for each step, and ask for approval before any irreversible external action.Turn this operating data into a three-scenario model. State every assumption, preserve formulas rather than hard-coding outputs, flag missing inputs, and run a sensitivity analysis on the two assumptions that matter most.Review this proposed experimental interpretation. List the strongest alternative explanations, the controls that would distinguish them, and which conclusions are supported versus merely plausible.Redline this agreement from the customer’s perspective. Group changes by business risk, explain the practical consequence of each clause, and preserve commercially reasonable language where possible.Build this responsive page from the brief. Before you finish, test the key flows at desktop and mobile widths, identify anything below the fold or off-screen, and fix those issues before summarizing the implementation.08 / QUICK ANSWERS
Use Opus 5 when the cost of a wrong answer, shallow plan, or failed multi-step task is materially higher than the model cost. Sonnet remains the better default for high-volume, less demanding work.
Anthropic says Opus 5 has the same base API price as Opus 4.8. Test your prompts, tool flow, latency expectations, and safety fallback behavior before a broad migration.
Anthropic’s announcement describes Fast mode’s speed and pricing, but does not specify rate-limit changes. Check the current Claude Platform documentation for your account and tier.
Anthropic lists $5 per million input tokens and $25 per million output tokens—the same base API price as Opus 4.8. Fast mode is priced at twice the base rate.
KEEP THE SIGNAL
Join our community of 800k professionals getting the AI news that actually matters.
One email. Five minutes.
Get ahead of 99% of the world.