Skip to the tests
Forward Future / Model field tests

Opus 5.5.
Put to the test.

We gave Opus 5.5 cities to build, games to make, and a story to animate. Here’s what came back.

Explore all eight tests
Matthew, Alex & Brian
Hands-on builds. Original recordings.

People, pets, cars, and traffic. Matthew’s San Francisco recreation brings Unreal Engine and Jev together.

Original post
Behind the tests

The work.
And the context.

A collection of hands-on experiments from the Forward Future team’s early access to Opus 5.5.

These projects cover different prompts, tools, and amounts of iteration. The recordings show what the team shared; they aren’t a controlled benchmark or a claim that every build worked on the first try.

Follow the process

Brian’s animation took about 12 hours to reach its initial result, followed by further requests. Tidekeeper emerged from a bracket of twelve puzzle games before the winner received more development.

Go to the source

Every test includes its original X post. The card game and puzzle collection also have public links, so you can try them yourself.

Explore Matthew’s Astra review
The reported numbers

Benchmarks.

Anthropic’s launch comparison, published September 22, 2026. These are reported results, separate from our eight hands-on tests.

Highlighted = highest score. — = unreported. Scroll horizontally to compare all five models.

Published benchmark results
EvaluationOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.0Highest score: 66.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 (Main)Highest score: 54.4%50.3%48.0%53.3%47.5%
CursorBench 4.0Highest score: 57.8%51.8%46.6%41.7%
GDPval-AA v2.1EloHighest score: 18461735170815421588
AutomationBench40.0%31.4%26.9%Highest score: 41.4%28.8%
Humanity’s Last ExamWith toolsHighest score: 67.7%65.6%63.6%57.2%
Terminal-Bench-Science 0.158.7%52.6%29.0%Highest score: 64.6%22.4%
OSWorld 2.0PartialHighest score: 81.8%80.7%74.0%
ChartographyWith toolsHighest score: 89.0%88.4%83.4%
Evaluation settings & caveats

Opus 5.5 uses adaptive thinking at max effort, except Terminal-Bench 4.0: xhigh for Opus, high for Astra. Production safeguards were enabled; interventions routed cybersecurity tasks to Opus 4.8, and biology/frontier-model work to Opus 5.

Zapier ran AutomationBench without fallbacks; interventions counted as failures. Terminal-Bench standard error: ±2.6 points for Opus 5.5, ±1.6–2 for other Claude models. Science: ±3.5–5 points. Anthropic’s reproduced Opus 5 scores differ slightly from public leaderboards. See the source for full methodology.

Source: Anthropic’s Opus 5.5 announcement

What it costs

Pricing.

Claude API list prices in USD per million tokens. Standard input and output are 20% cheaper than Opus 5.

Standard input / 1M tokens$4
Standard output / 1M tokens$20
Cache reads / 1M tokens$0.20
Opus 5.5 · USD per 1M tokens
ProcessingInputOutput
Standard$4$20
Fast mode Research preview$8$40
Batch API Asynchronous$2$10
Standard cache writes / 1M tokens5-minute cache $51-hour cache $8

Fast mode is available on the first-party Claude API and cannot be combined with Batch. Prices above use global routing; US-only inference adds 10%. Tool charges and partner-platform pricing can differ. API usage is billed separately from Claude subscriptions.

Source: Claude Platform pricing · Checked September 22, 2026

Keep looking forward

More tests. Less guesswork.

Follow the next round of AI experiments at Forward Future.

Explore Forward Future