A recreation of San Francisco in Unreal Engine, with people, pets, cars, and traffic powered by Jev.
Opus 5.5.
Put to the test.
We gave Opus 5.5 cities to build, games to make, and a story to animate. Here’s what came back.
Explore all eight testsPeople, pets, cars, and traffic. Matthew’s San Francisco recreation brings Unreal Engine and Jev together.
Original postAlex’s Dark Souls recreation, built with Opus 5.5 and Unreal Engine.
A flight simulator from Alex’s early-access experiments with Opus 5.5.
Alex takes on Mario Maker in another of his Opus 5.5 game-building tests.
A rocket roguelike rounds out Alex’s four-game Opus 5.5 showcase.
Opus 5.5 added new features to Brian’s card game and edited the trailer. Watch it, then try the game yourself.
Brian described a music-driven stickman cartoon about looting a labyrinth. An initial result took about 12 hours, followed by more requests to refine it.
Twelve puzzle games entered a bracket judged on design, fun, clarity, and satisfying gameplay. Brian then asked Opus to develop the winner: Tidekeeper.
The work.
And the context.
A collection of hands-on experiments from the Forward Future team’s early access to Opus 5.5.
These projects cover different prompts, tools, and amounts of iteration. The recordings show what the team shared; they aren’t a controlled benchmark or a claim that every build worked on the first try.
Brian’s animation took about 12 hours to reach its initial result, followed by further requests. Tidekeeper emerged from a bracket of twelve puzzle games before the winner received more development.
Every test includes its original X post. The card game and puzzle collection also have public links, so you can try them yourself.
Benchmarks.
Anthropic’s launch comparison, published September 22, 2026. These are reported results, separate from our eight hands-on tests.
Highlighted = highest score. — = unreported. Scroll horizontally to compare all five models.
| Evaluation | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | Highest score: 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (Main) | Highest score: 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | Highest score: 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1Elo | Highest score: 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench | 40.0% | 31.4% | 26.9% | Highest score: 41.4% | 28.8% |
| Humanity’s Last ExamWith tools | Highest score: 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | Highest score: 64.6% | 22.4% |
| OSWorld 2.0Partial | Highest score: 81.8% | 80.7% | 74.0% | — | — |
| ChartographyWith tools | Highest score: 89.0% | 88.4% | 83.4% | — | — |
Evaluation settings & caveats
Opus 5.5 uses adaptive thinking at max effort, except Terminal-Bench 4.0: xhigh for Opus, high for Astra. Production safeguards were enabled; interventions routed cybersecurity tasks to Opus 4.8, and biology/frontier-model work to Opus 5.
Zapier ran AutomationBench without fallbacks; interventions counted as failures. Terminal-Bench standard error: ±2.6 points for Opus 5.5, ±1.6–2 for other Claude models. Science: ±3.5–5 points. Anthropic’s reproduced Opus 5 scores differ slightly from public leaderboards. See the source for full methodology.
Pricing.
Claude API list prices in USD per million tokens. Standard input and output are 20% cheaper than Opus 5.
| Processing | Input | Output |
|---|---|---|
| Standard | $4 | $20 |
| Fast mode Research preview | $8 | $40 |
| Batch API Asynchronous | $2 | $10 |
Fast mode is available on the first-party Claude API and cannot be combined with Batch. Prices above use global routing; US-only inference adds 10%. Tool charges and partner-platform pricing can differ. API usage is billed separately from Claude subscriptions.
Source: Claude Platform pricing · Checked September 22, 2026
More tests. Less guesswork.
Follow the next round of AI experiments at Forward Future.