A FORWARD FUTURE FIELD GUIDE / CLAUDE OPUS 5

SHOULD YOU
SWITCH?

Yes—if careful, multi-step coding, analysis, or agent work is a meaningful part of your job.

Why: Opus 5 is a meaningful step up from 4.8 for that work at the same base API price, and its effort controls let you spend less on routine tasks without giving up deeper reasoning when a task needs it. Keep Sonnet for high-volume general work, and validate Opus against your actual workflows before a broad rollout.

RELEASEDJUL 24, 2026
API PRICE$5 / $25 PER MTOK
DEFAULTCLAUDE MAX
FAST MODE~2.5× SPEED · 2× PRICE

01 / THE CAVEAT

What is still unclear.

Does “near Fable intelligence at half the cost” survive independent testing?

That claim comes from Anthropic’s own reported results. It is plausible, but third-party evaluations across real workloads are the only way to know how broadly it holds.

What does the cyber gap actually mean?

Opus 5 comes close to Mythos 5 at finding vulnerabilities, according to Anthropic, but remains far behind at developing exploits. Finding and exploiting are different capabilities; coverage should not collapse them into one claim.

How will the safeguard change work in practice?

Anthropic says biology-related requests blocked on Fable 5 will now route to Opus 5 instead of Opus 4.8. That is a real policy change, not merely a model-spec update.

02 / THE X READ

What people are saying after trying it.

Early reactions are not benchmarks. They are useful, however, for the details a launch chart cannot show: taste, code quality, speed, and where the model feels different in practice.

Posts selected for reach and useful first-hand perspective; views expressed are their authors’ own.

03 / THE CHANGELOG

What actually changed.

Not a new price tier. Not a new context-window story. Opus 5 is Anthropic’s bid to make frontier-level work more adjustable—and more usable every day.

01

A new top line on coding and knowledge-work evals

Anthropic reports state-of-the-art results on Frontier-Bench and GDPval-AA. It also says Opus 5 remains behind Mythos 5 on cybersecurity.

02

Effort settings, plus a separate Fast mode

Use low through max effort to balance depth and token use. Fast mode is a speed option—around 2.5× faster at twice the base price—not another effort level.

03

More permissive safeguards than Fable 5

Anthropic expects cyber classifiers to intervene about 85% less often than Fable 5’s, while still blocking binary-based scanning, penetration testing, and exploit generation.

04

A visible life-sciences jump

Anthropic reports gains of 10.2 percentage points on internal organic-chemistry tasks and 7.7 points on protein-function prediction versus Opus 4.8.

05

Two useful platform betas

Developers can change tools mid-conversation without invalidating prompt cache, and opt into automatic safety fallbacks on the API.

ANTHROPIC-REPORTED RESULTS

04 / PERFORMANCE

Benchmarks are a map. Not the territory.

The pattern in Anthropic’s release data is clear: Opus 5 is trying to offer near-Fable performance at a lower task cost. Independent replication is still the thing to watch.

ARC-AGI 3

Anthropic says Opus 5 scores three times higher than the next-best model on this novel-problem-solving evaluation.

AUTOMATIONBENCH~1.5×

Pass rate at the same cost per task versus the next-best model, according to Anthropic.

OSWORLD 2.0~⅓

Anthropic says Opus 5 exceeds Fable 5’s best result at just over a third of its cost.

Which model belongs on the job?

A release-day decision guide, built to be updated as the model landscape changes.

MODELPRICE POSITIONBEST FIT
OPUS 4.8$5 / $25 per MTokExisting integrations, proven workflows
SONNET 5Check current pricingHigh-volume general work
FABLE 5Higher-cost frontier tierMaximum intelligence when budget is secondary

Named evaluations in the launch: Frontier-Bench v0.1, CursorBench 3.2, AA Coding Agent Index, ARC-AGI 3, GDPval-AA v2, OSWorld 2.0, HLE, AutomationBench, and DeepSearchQA.

05 / SHOW, DON’T TELL

Two things Opus 5 built itself.

Anthropic shipped live artifacts with the launch. Try them here instead of taking a screenshot’s word for it.

DEMO 01

Wind tunnel

An interactive airflow simulation. The interesting part is not just that it renders—it gives you a real control surface to explore.

Open full screen ↗
DEMO 02

Interactive cell

A zoomable cell illustration built as an explorable artifact, a stronger test of visual explanation than a static output.

Open full screen ↗

WATCH / FORWARD FUTURE VIDEO

Our Opus 5 breakdown.

Watch the companion video, then use this guide to compare the headline claims with the practical tradeoffs.

06 / GETTING REAL VALUE

Use the expensive thinking deliberately.

01

Start low. Escalate with evidence.

Use low effort for drafting, classification, extraction, and well-bounded edits. Move up when the model must plan, debug across files, verify a result, or recover from ambiguity.

02

Keep the prompt cache warm.

Anthropic says prompt caching can cut costs by up to 90%. Put stable instructions, reference material, and tools early in your request structure.

03

Batch work that can wait.

For offline enrichment, review, or transformation jobs, batch processing can reduce cost by 50%—a lever more teams should use before changing models.

04

Use Fast mode for the same job, sooner.

Fast mode is worth considering where responsiveness changes the workflow. It is not a shortcut to better reasoning, and it costs twice the base rate.

Read Anthropic’s prompting guidance

07 / PROMPT COOKBOOK

Give it a real job.

These prompts are starting points, not magic spells. Include your constraints, source material, desired output, and how you want the work checked.

ADVANCED CODING

Debug the system, not the symptom.

Trace this production error across this repository. First map the likely data flow and competing root-cause hypotheses. Then propose the smallest fix, add a regression test, and explain what evidence would falsify your conclusion.
AI AGENTS

Build in checkpoints.

Plan this research task as an agent workflow. Separate facts we can verify from assumptions, define a stop condition for each step, and ask for approval before any irreversible external action.
FINANCIAL MODELING

Make the model auditable.

Turn this operating data into a three-scenario model. State every assumption, preserve formulas rather than hard-coding outputs, flag missing inputs, and run a sensitivity analysis on the two assumptions that matter most.
LIFE SCIENCES

Make uncertainty explicit.

Review this proposed experimental interpretation. List the strongest alternative explanations, the controls that would distinguish them, and which conclusions are supported versus merely plausible.
LEGAL REDLINING

Focus on the actual risk.

Redline this agreement from the customer’s perspective. Group changes by business risk, explain the practical consequence of each clause, and preserve commercially reasonable language where possible.
FRONTEND BUILDS

Ask it to inspect its own work.

Build this responsive page from the brief. Before you finish, test the key flows at desktop and mobile widths, identify anything below the fold or off-screen, and fix those issues before summarizing the implementation.

08 / QUICK ANSWERS

Questions to settle before rollout.

When should I use Opus 5 instead of Sonnet 5?

Use Opus 5 when the cost of a wrong answer, shallow plan, or failed multi-step task is materially higher than the model cost. Sonnet remains the better default for high-volume, less demanding work.

Do existing Opus 4.8 API integrations need to change?

Anthropic says Opus 5 has the same base API price as Opus 4.8. Test your prompts, tool flow, latency expectations, and safety fallback behavior before a broad migration.

Does Fast mode change my rate limits?

Anthropic’s announcement describes Fast mode’s speed and pricing, but does not specify rate-limit changes. Check the current Claude Platform documentation for your account and tier.

How much does Opus 5 cost?

Anthropic lists $5 per million input tokens and $25 per million output tokens—the same base API price as Opus 4.8. Fast mode is priced at twice the base rate.

KEEP THE SIGNAL

Wake up smarter about AI

Join our community of 800k professionals getting the AI news that actually matters.
One email. Five minutes.
Get ahead of 99% of the world.