Find the right model
for the thing you’re doing.

Choose the job. We’ll put our current pick at the top, show you the alternatives, and tell you exactly when the ruling was last reviewed.

Start here

What do you want to do?

More use cases

The current verdict

Best match for your task

Reviewed weekly

All current verdicts

Locked: not close Leaning: clear pick, live race Contested: personal choice

Best model for coding

Contested

Claude Fable 5.

As of · Reviewed Jul 17, 2026

Under review: Jul 19 - Fable's promotional subscription access ends

Runner-up
GPT-5.6 Sol - benchmark-competitive, but one week of real-world use so far
Budget pick
Claude Sonnet 5 - intro $2/$10 per 1M tokens through Aug 31 (as of Jul 17)
Full verdict

Why

Fable 5 performs well on multi-file refactors and long-horizon changes that require consistent architectural context. During its 19-day export-control suspension, we used Opus 4.8 and restored Fable as the coding pick when access returned on July 1. GPT-5.6 Sol has competitive agentic-coding results and a four-agent ultra mode, but it has not yet accumulated a month of daily production use. Fable costs $10/$50 per 1M tokens, so Sonnet 5 remains the lower-cost option.

The receipts

What would change our mind

  • If Sol's real-world reliability matches its benchmark numbers after a month of daily use, this flips.
  • If the July 19 promo cliff puts Fable out of subscription reach for most readers, the budget pick becomes the practical ruling.

Ruling history

  1. Jul 1, 2026Re-flipped to Fable 5 the day it was restored (US export controls lifted Jun 30).
  2. Jun 12, 2026Reverted to Opus 4.8 when a US export-control order pulled Fable 5 offline worldwide.
  3. Jun 10, 2026Flipped from Opus 4.8 to Fable 5 after initial testing favored Fable on long-horizon coding tasks.

Best model for writing

Leaning

Claude Sonnet 5.

As of · Reviewed Jul 17, 2026

Under review: Sep 1 - intro pricing steps up to $3/$15

Runner-up
GPT-5.6 Terra - the best structured long-form outliner, stiffer prose
Budget pick
Muse Spark 1.1 - $1.25/$4.25 (as of Jul 17), competitive on short-form work
Full verdict

Why

Benchmarks capture writing worse than any other category, so this ruling leans hardest on daily use: Sonnet 5 drafts in something close to a human register, cuts filler on its own, and takes edit instructions literally instead of rewriting around them. Three weeks of newsletter drafts that needed one pass instead of three sold us - the 96.2% GPQA Diamond score is beside the point. Two footnotes the leaderboards bury: the "roughly cost-neutral" migration claim hides a new tokenizer that emits 1.0–1.35x as many tokens as Sonnet 4.6, and the intro $2/$10 pricing steps up to $3/$15 on September 1. Neither changes the quality verdict; both change the invoice.

The receipts

  • Vellum LLM Leaderboard - Sonnet 5 tops GPQA Diamond at 96.2%, verified Jul 17
  • Forward Future - three weeks of production newsletter drafting since Jun 30
  • Anthropic - tokenizer migration notes and Sep 1 price schedule, as of Jul 17

What would change our mind

  • If the September 1 price step lands and Terra closes the register gap in our blind drafting comparisons, this flips on value.

Ruling history

  1. Jun 30, 2026Flipped from Opus 4.8 the day Sonnet 5 shipped at 60%-off intro pricing. Better drafts at a fifth of the cost.

Best model for deep research

Leaning

GPT-5.6 Sol.

As of · Reviewed Jul 17, 2026

Runner-up
Claude Fable 5 - writes the better final memo, narrower sweep
Budget pick
DeepSeek V4 Pro - 48.2% on HLE at open-weight prices
Full verdict

Why

Sol's ultra mode runs four parallel agents, and it's the first "go read everything and come back" workflow that returns with sources we didn't have to re-verify line by line. Fable 5 still writes the better final synthesis, but Sol wins where research actually fails: coverage. In our first week it surfaced primary documents the others missed and flagged its own uncertainty instead of papering over it. At $5/$30 per 1M tokens - and ultra mode burns tokens - this is a professional tool, not a casual one; the free-tier answer is the everyday-chat verdict. One week is a short window, hence Leaning, and the flip condition below is doing real work.

The receipts

What would change our mind

  • If a month of use shows ultra-mode agents compounding errors on long chains, this reverts to Fable 5.
  • If Anthropic ships a comparable parallel-research mode, we re-run the shootout that week.

Ruling history

  1. Jul 10, 2026Flipped from Fable 5 the day after GPT-5.6's public release. Parallel agents beat a better writer at this job.
  2. Jun 10, 2026Flipped from Opus 4.8 to Fable 5 at launch.

Best everyday chat model

Leaning

Claude Sonnet 5.

As of · Reviewed Jul 17, 2026

Runner-up
ChatGPT - best app ecosystem; GPT-5.5 Instant is still its default
Budget pick
Meta One AI - $7.99/mo plan in testing (as of Jul 17); too new to rule
Full verdict

Why

Sonnet 5 became Claude's Free and Pro default on June 30. In our testing, it handled email, summaries, and explanations more consistently than ChatGPT's default. GPT-5.5 Instant remains ChatGPT's everyday model; GPT-5.6 Sol is available through reasoning settings on eligible paid plans. ChatGPT keeps the runner-up slot for its voice, memory, and app integrations. Meta One AI's proposed $7.99 plan will be evaluated when it becomes broadly available.

The receipts

  • Anthropic - Sonnet 5 as Free/Pro default, Jun 30
  • OpenAI - ChatGPT default-model documentation, as of Jul 17
  • Forward Future - our everyday-task comparison notes

What would change our mind

  • If OpenAI makes GPT-5.6 the everyday ChatGPT default, this likely flips the same week.
  • If Meta One AI ships broadly at $7.99 with acceptable quality, the value math changes for casual users.

Ruling history

  1. Jun 30, 2026Flipped from ChatGPT (GPT-5.5 Instant) when Sonnet 5 became Claude's default. The free tiers are no longer close.

Best budget / high-volume model

Contested

GLM 5.2.

As of · Reviewed Jul 17, 2026

Under review: Sep 1 - Sonnet 5's intro pricing expires, reshuffling the tier

Runner-up
GPT-5.6 Luna - $1/$6 (as of Jul 17) with the OpenAI ecosystem
Budget pick
Grok 4.5 cached input at $0.50 - for repeat-context workloads; pricing doubles past 200k tokens
Full verdict

Why

GLM 5.2 costs $0.95/$3 per 1M tokens, runs at 347 tokens/sec, and scores 54.7% on Humanity's Last Exam. Other current options include GPT-5.6 Luna at $1/$6, Muse Spark 1.1 at $1.25/$4.25 with $20 in introductory credits, Sonnet 5 at $2/$10 through August 31, and Grok 4.5 with $0.50 cached input. Grok pricing doubles beyond 200k input tokens. Every price here is as of July 17, 2026.

The receipts

What would change our mind

  • Any price cut from Luna or Muse Spark puts them within a re-run.
  • Sonnet 5's September 1 step-up removes it from this tier entirely - calendared.

Ruling history

  1. Jul 16, 2026Issued at launch. The tier re-opened Jul 8–9 when Grok 4.5, Luna, and Muse Spark 1.1 all shipped within 48 hours.

Best open-weight model

Leaning

GLM 5.2.

As of · Reviewed Jul 17, 2026

Under review: Jul 27 - Kimi K3 full weights scheduled for release

Runner-up
Kimi K3 - API live now; full weights scheduled for Jul 27
Budget pick
DeepSeek V4 Flash - 51.6% on HLE at the lowest tier price
Full verdict

Why

GLM 5.2 keeps the ruling because its weights are available, its independent results are established, and its hosted price is $0.95/$3. Kimi K3 launched on July 16 with 2.8T parameters, native vision, a 1M-token context window, and API pricing of $3/$15. Its full weights and technical report are scheduled for July 27. K3 moves into the runner-up position while the verdict remains under review.

The receipts

What would change our mind

  • Kimi K3 shipping downloadable weights on July 27 with independent results that hold near its launch claims.
  • Any of the three changing license terms - open-weight verdicts are license verdicts too.

Ruling history

  1. Jul 17, 2026Held GLM 5.2 after Kimi K3's API launch. K3 moves directly to runner-up; re-review scheduled for the full weight release.
  2. Jul 16, 2026Issued at launch. GLM 5.2 over Kimi K2.6 by a nose, on speed and price rather than capability.

Best coding agent / dev tool

Contested

Cursor.

As of · Reviewed Jul 17, 2026

Runner-up
Claude Code - the better pure agent if you live in a terminal
Budget pick
Claude Code bundled with Claude Pro at $20/mo
Full verdict

Why

The honest answer this month: the harness matters more than the model, because the model lead keeps changing hands. Cursor wins by being model-agnostic at exactly that moment - Fable 5, GPT-5.6 Sol, and Grok 4.5 (which was co-trained with Cursor) all slot in the day they ship, so a verdict flip in the coding category doesn't force a tool change. Claude Code is the better pure agent if you're terminal-native and all-in on Anthropic models. The Contested tag is structural: the category is red-hot, and the pending ~$60B SpaceXAI acquisition of Cursor could change what Cursor is - see the flip conditions.

The receipts

What would change our mind

  • If the SpaceXAI acquisition closes and Cursor deprioritizes non-Grok models, this flips to Claude Code.
  • If a terminal agent ships true multi-model support at Cursor's polish level, the race re-opens.

Ruling history

  1. Jul 16, 2026Issued at launch. Model-agnosticism is the deciding trait in a month with three frontier releases.

Best model for agents & automation

Leaning

Claude Sonnet 5.

As of · Reviewed Jul 17, 2026

Under review: Sep 1 - the price step-up hits pipelines hardest

Runner-up
Grok 4.5 - 500K context and $0.50 cached input, but eight days old
Budget pick
GLM 5.2 - $0.95/$3 for the steps that don't need a frontier model
Full verdict

Why

Agents multiply every model trait by a thousand calls: cost, latency, and how a model fails matter more than peak intelligence. Sonnet 5 is the current sweet spot - fast, obedient on structured output, and $2/$10 intro pricing that makes long chains affordable. Two operational notes: the new tokenizer's 1.0–1.35x token multiplier compounds across a pipeline, so budget for the top of that range, and the September 1 step-up to $3/$15 is calendared below. Grok 4.5's 500K context and $0.50 cached input make it the one to watch for long-context automation - but it shipped July 8, and we don't rule on a week.

The receipts

  • Artificial Analysis - latency, throughput, and price standings
  • Anthropic - Sonnet 5 tokenizer notes and price schedule, as of Jul 17
  • Forward Future - our own production pipelines run on Sonnet-class models daily

What would change our mind

  • The September 1 price step could flip this to GLM 5.2 on cost alone - calendared.
  • A month of Grok 4.5 proving reliable on long-context chains puts it in a head-to-head.

Ruling history

  1. Jul 16, 2026Issued at launch. Sonnet 5 over Grok 4.5 on track record, not capability.

Best image model

Leaning

Gemini 3 Pro Image.

As of · Reviewed Jul 17, 2026

Runner-up
Muse Image - shipped last week; striking photoreal, no track record yet
Budget pick
Muse Image on Meta's $20 free API credits (as of Jul 17)
Full verdict

Why

For work - social assets, brand visuals, editorial art - the three traits that matter are instruction-following on layout, legible text rendering, and consistent style across a series. Gemini 3 Pro Image is still the most controllable model on all three, which is why it runs in our own production pipeline every day; that's the receipt that matters most. Muse Image shipped last week alongside Spark 1.1 and the early photoreal results are genuinely striking, but a week is not a track record, and series consistency is exactly the thing that takes a month of daily use to judge. This is the most watch-this-space verdict on the board.

The receipts

What would change our mind

  • A month of Muse Image holding up on text rendering and series consistency flips this.

Ruling history

  1. Jul 16, 2026Issued at launch. Muse Image noted as the live challenger, seven days into its existence.

Best $20 consumer subscription

Contested

Claude Pro.

As of · Reviewed Jul 17, 2026

Under review: Jul 19 - Fable 5's promotional subscription access ends

Runner-up
ChatGPT Plus - includes GPT-5.6 Sol, buried behind reasoning settings
Budget pick
Meta One AI - $7.99/mo plan in testing (as of Jul 17)
Full verdict

Why

For $20 a month, Claude Pro currently gets you the best everyday default (Sonnet 5) plus promotional access to Fable 5 - the top of our coding board - through July 19. That last clause is exactly why this verdict carries a scheduled review three days out: when the promo lapses, the calculus may flip to ChatGPT Plus, which does include GPT-5.6 Sol even if you have to dig through reasoning settings to use it. The dinner-party version: if you code at all, Claude Pro; if you mostly want the app ecosystem - voice, memory, integrations - ChatGPT Plus; and if you're price-first, wait for Meta One AI to leave testing. Ask us again on the 19th. We mean that literally.

The receipts

  • Anthropic - plan tiers and the Jul 19 promo end date, as of Jul 17
  • OpenAI - ChatGPT Plus model access, as of Jul 17
  • Forward Future - we pay for all of these; subscription notes in the newsletter

What would change our mind

  • The July 19 promo lapse - calendared, and this verdict gets re-ruled that day.
  • Meta One AI's plans going GA at $7.99/$19.99 resets the value baseline for casual users.

Ruling history

  1. Jul 16, 2026Issued at launch, with a known three-day shelf life. That's the honest shape of this category right now.

Specific-use pick

Best for structured long-form & outlining

Leaning

GPT-5.6 Terra.

As of · Reviewed Jul 17, 2026

Use for
Long reports, presentations, ordered outlines
Category
Writing overall still goes to Claude Sonnet 5
Runner-up here
Claude Sonnet 5 - better final prose and editing
Full verdict

Why

Terra is the structured long-form specialist in the writing category. When the job is outlining a report, building a slide narrative, or holding a long ordered argument together, it beats Sonnet 5 on structure. Sonnet still wins everyday drafting and final prose - that is why the category card stays with Sonnet, and why this is a specific-use pick rather than a category flip.

The receipts

  • Writing verdict - Terra listed as runner-up / structured long-form specialist
  • OpenAI - GPT-5.6 Terra positioning, as of Jul 17
  • Forward Future - outlining and long-form draft comparisons since Jul 9

What would change our mind

  • If Sonnet closes the outlining gap in our blind structure comparisons, this specific-use pick collapses back into the writing verdict.
  • If Terra's prose register matches Sonnet on final drafts, the category card itself is in play.

Ruling history

  1. Jul 9, 2026Issued as a specific-use pick when GPT-5.6 Terra went public. Category writing ruling stayed with Sonnet 5.

Specific-use pick

Best for voice, memory & app integrations

Leaning

ChatGPT.

As of · Reviewed Jul 17, 2026

Use for
Voice, persistent memory, and the broader app ecosystem
Category
Everyday chat overall still goes to Claude Sonnet 5
Runner-up here
Claude Sonnet 5 - better default model for general chat
Full verdict

Why

ChatGPT keeps the specific-use pick when the product around the model matters more than the default model itself: voice, persistent memory, and app integrations. Sonnet 5 remains the everyday chat ruling for email, summaries, and explanations. That split is deliberate - model quality and product surface are different jobs.

The receipts

What would change our mind

  • If Claude ships a comparable voice + memory + apps surface, this specific-use pick collapses into the chat verdict.
  • If OpenAI makes GPT-5.6 the everyday ChatGPT default, the category card may flip the same week.

Ruling history

  1. Jun 30, 2026Kept as the ecosystem / voice / memory pick when Sonnet 5 took the everyday chat ruling.

Forward Future Stack

What we're using today

Leaning
Daily driverClaude Sonnet 5
CodingClaude Fable 5
Cheap bulkGLM 5.2
ImageGemini 3 Pro Image

As of · Reviewed Jul 17, 2026

Full verdict

Why

Power users don't pick one model; they run a stack. This is ours: Sonnet 5 for everyday work and first drafts, with introductory pricing of $2/$10 through August 31. Fable 5 for complex engineering work where long-horizon reasoning justifies its $10/$50 rate. GLM 5.2 for high-volume classification, extraction, and batch summarization at $0.95/$3. Gemini 3 Pro Image for visual work that requires layout control and consistent styling.

What would change our mind

  • Sonnet 5's September 1 price step-up ($3/$15) could hand the daily-driver slot to a cheaper model.
  • A month of GPT-5.6 Sol proving out could collapse the coding and research slots into one.

Ruling history

  1. Jul 16, 2026Issued at launch. Muse Image showed competitive early results, but had only been available for seven days.

The change feed

Every ruling change across every verdict.

  1. Jul 17, 2026 Note

    Kimi K3 enters as the open-weight runner-up after its API launch. Full weights are scheduled for July 27; the verdict is under review until they are actually downloadable.

  2. Jul 16, 2026 Launch

    The Verdicts goes live. All ten verdicts plus the FF Stack issued or re-affirmed; the ledger below is seeded from our internal tracking since June 9.

  3. Jul 10, 2026 Flip

    Deep research: Claude Fable 5 → GPT-5.6 Sol. Ultra mode's four parallel agents beat a better writer at coverage.

  4. Jul 9, 2026 Note

    GPT-5.6 Sol/Terra/Luna go public; Meta ships Muse Spark 1.1 at $1.25/$4.25, its first paid model. Budget and subscription verdicts placed under review. Also noted, because most coverage botched it: GPT-5.5 Instant remains ChatGPT's everyday default.

  5. Jul 8, 2026 Note

    Grok 4.5 ships at $2/$6 with 500K context and $0.50 cached input (pricing doubles past 200k input tokens - the footnote leaderboards bury). Watching for agents and budget.

  6. Jul 1, 2026 Flip

    Coding: Opus 4.8 → Claude Fable 5, re-flipped the day the model was restored. Export controls lifted June 30.

  7. Jun 30, 2026 Flip

    Writing: Opus 4.8 → Claude Sonnet 5, shipped at 60%-off intro pricing. Everyday chat: ChatGPT → Claude, as Sonnet 5 becomes the Free/Pro default.

  8. Jun 12, 2026 Revert

    Coding: Claude Fable 5 → Opus 4.8. A US export-control order pulled Fable offline worldwide three days after launch. The category churn is the product.

  9. Jun 10, 2026 Flip

    Coding: Opus 4.8 → Claude Fable 5, one day after the first public Mythos-class model launched.

Scheduled reviews

Methodology

How a ruling gets made, in plain language.

What we weigh

Published benchmarks, preferring independent harnesses (Vellum, Artificial Analysis) over vendor-published charts. Hands-on daily use - we run these models in production, on camera, and in this codebase. Price, always with an as-of date and any known expiration. Reliability: how a model fails matters as much as how it scores.

What we ignore

Saturated benchmarks where the models score so closely that the differences are not meaningful. Vendor-only claims that no independent harness has reproduced. Launch-day vibes - we don't rule on a week of use, which is why several runner-ups on this board are newer than the models that beat them.

The pledge

No sponsored verdicts. Ever. Sponsorship can live in clearly labeled adjacent placements, never in a ruling. One compromised verdict would kill this entire asset, and we know it.

Conflicts disclosure

Forward Future carries sponsor and partner relationships across its newsletter, YouTube channel, and live show - some with companies whose models appear on this board. Sponsors get no input into rulings, no preview of flips, and no notice before publication. The proof is the ledger: the history on every verdict shows we've flipped toward and away from every lab, on dates you can check against their sponsorship calendars.

Review cadence

Every verdict is reviewed weekly; ruling changes ship within 24 to 48 hours of a major release; scheduled reviews are calendared the day a dated event becomes known.

Pre-registered flips

Every verdict states in advance what would change our mind. When a flip happens, it isn't "they changed their minds" - it's the thing we said would happen, happening. Check any verdict's history against its flip conditions and hold us to it.

FAQ

How often are the verdicts reviewed?

Weekly, every one of them, and ruling changes ship within 24 to 48 hours of a major release. Each verdict carries a last-reviewed date even when the ruling holds. Verdicts affected by dated events - pricing cliffs, availability rollouts - display a scheduled review date in advance.

What do the confidence tags mean?

Locked: not close. Leaning: clear pick, live race. Contested: reasonable people disagree and the ruling could flip on the next release. The tag is an honest signal when the race is tight, not a hedge.

Do sponsors influence the verdicts?

No. There are no sponsored verdicts, ever - see the pledge and conflicts disclosure in the methodology. The ruling history is the proof: we've flipped toward and away from every major lab, on the public record.

Why trust an editorial verdict over a leaderboard?

Leaderboards are useful data sources, and we cite them on nearly every verdict. We add a dated editorial decision, supporting evidence, pre-registered flip conditions, and a public history of changes.

What happens when a verdict is wrong?

We flip it in public, within 24 to 48 hours of the evidence, and the history ledger records the date and the reason. Every verdict also states in advance what would change our mind, so a flip is a prediction landing, not a walk-back.

The Verdicts — Best AI Model for Every Job, Dated