AI Model Reviews

Claude Opus 5 Review: Features, Pricing, Pros and Cons

Claude Opus 5 review featured image highlighting Anthropic’s recommended model for complex agentic and enterprise work, with a one-million-token context window, adaptive reasoning effort, agentic coding, and pricing at half the cost of Claude Fable 5.1

Open Anthropic's own developer documentation and look for the line that tells you which Claude model to actually use. It does not point at the expensive one. It says, in plain text: start with Claude Opus 5 for most workloads, and reach for the pricier flagship, Claude Fable 5.1, only once your own testing shows Opus 5 falls short. That is unusual for a company to publish about its own $10-per-million-token model, and it is the reason this review exists.

One quick note before we go further. Frontier AI models move fast, and Anthropic released four separate Claude 5-generation models in under two months this year. Treat every number below as accurate as of today, September 12, 2026, and confirm the current lineup, pricing, and access directly on Anthropic's own docs before you commit real budget to any of it.

Claude Opus 5 is Anthropic's mid-tier flagship: a $5-input, $25-output-per-million-token model built for complex agentic coding and enterprise work, priced at exactly half of Claude Fable 5.1 while landing within a point or two of it on most independently verified benchmarks. This post covers what changed, what the benchmarks show once vendor numbers are separated from independent testing, what it actually costs including the cache-pricing details, and where it sits against every other model in the current Claude lineup, a comparison most coverage skips entirely.

Claude Opus 5 at a Glance

Spec

Detail

Maker

Anthropic

Released

July 24, 2026

Model ID / alias

claude-opus-5 (dateless, pinned snapshot)

Positioning

Complex agentic coding and enterprise work

Context window

1,000,000 tokens

Max output

128,000 tokens (up to 300,000 on the Batch API with a beta header)

Input price

$5 per million tokens

Output price

$25 per million tokens

Knowledge cutoff

May 2026 (reliable), June 2026 (training data)

Latency

Moderate

Thinking mode

Adaptive, default effort high

Retirement floor

Not sooner than July 24, 2027

Availability

Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry

What Is Claude Opus 5, and Where Does It Sit in Anthropic's Lineup?

Claude Opus 5 is the fourth Claude 5-generation model Anthropic shipped in under two months, arriving July 24, 2026, after Mythos 5, Fable 5, and Sonnet 5 in June. "Opus" has historically meant Anthropic's most capable, most expensive tier. This generation breaks that pattern on price: Opus 5 launched at the same $5/$25 rate as its predecessor, Opus 4.8, even as it closed much of the capability gap to Fable 5.1, the model that actually holds the "most capable" title now.

That repositioning is the whole story of this release. Anthropic built Opus 5 to be the model most people reach for by default, not a stepping stone to something pricier. Its own docs describe it as built "for complex agentic coding and enterprise work," a tier below Fable 5.1's "demanding reasoning and long-horizon agentic work" but above Sonnet 5's speed-first positioning.

What's Actually New in Opus 5

The headline consumer-facing feature is an effort toggle: you can set how much reasoning effort the model spends on a given request, low, medium, or high, trading cost and latency against capability on a per-call basis rather than switching models entirely. Think of it like choosing shipping speed at checkout, same product, but standard, express, or overnight each cost and arrive differently. That maps directly to the documented effort parameter on Anthropic's platform, and Opus 5 defaults to high.

A few other things worth knowing, kept honestly separated between what is documented and testable versus how Anthropic itself framed the release:

  • Full 1M context window at standard pricing. No long-context premium, a 900,000-token request bills at the same per-token rate as a 9,000-token one, verified directly on Anthropic's pricing page.

  • 300,000 output tokens on the Batch API, with the output-300k-2026-03-24 beta header, more than double the 128,000-token synchronous limit, useful for large async jobs.

  • Self-verification and less iterative refinement. Anthropic and early coverage describe Opus 5 as better at checking and self-correcting its own work without prompting. That is a vendor/product framing, not a benchmark result, so treat it as worth testing on your own tasks rather than a settled fact.

  • Positioned for scientific research, with a specific emphasis on biology-related tasks, again Anthropic's own positioning, not an independently audited claim.

  • Tighter cybersecurity guardrails and an automatic safety fallback. A declined request reportedly routes to another model instead of returning a bare error. Anthropic calls it the most aligned Opus model it has shipped.

What the Benchmarks Actually Say About Claude Opus 5

This section earns its keep, because most model-launch coverage is just the vendor's own chart, retyped. Here it is split cleanly: what Anthropic itself reported, and what independent evaluators found once they ran their own tests.

Anthropic's own reported numbers (vendor-reported, not independently verified unless noted otherwise):

  • ARC-AGI-3: roughly 30%, which Anthropic frames as three times the next best model. Worth noting: this is a modest number next to some rivals' headline ARC-AGI claims elsewhere in the industry, several of which depend on non-standard harnesses. Anthropic's figure here is not one of those inflated ones.

  • IMO 2026: 42 out of 42 problems, gold-medal level, without an external agent harness or tools. The real conditions matter: a 256,000-token output limit, adaptive thinking set to maximum, and at least one solution needing resampling at a lower effort to land. Genuinely strong, but not "we just asked it casually."

  • An internal benchmark Anthropic calls Frontier Code reportedly peaks at medium reasoning effort, around 53%, and declines at higher settings. A useful, counterintuitive point: more thinking is not automatically better on this specific test.

Independent testing, from evaluators with no stake in how Opus 5 performs:

  • Vals AI, an independent evaluator that ran Opus 5 across 27 separate benchmark leaderboards, puts it at number one on its overall Vals Index (67.21%) and number two on its Multimodal Index (73.90%), both within roughly a point of Claude Fable 5.1. Opus 5 ranks first on 14 of those 27 leaderboards, including SWE-bench Verified (97.00%), MMLU Pro (91.59%), and MMMU (89.88%).

  • Artificial Analysis, a separate independent benchmark aggregator, measured Opus 5 at 89.1% on Terminal-Bench 2.1 at maximum effort, in the Terminus 2 harness inside an e2b sandbox, pass@1 averaged over three repeats. That places it second behind GPT-5.6 Sol's 89.5% and ahead of Grok 4.6's 88.4%. Notice that Vals AI's own Terminal-Bench 2.1 run for the same model landed at 84.64%, a real gap on the exact same benchmark name. That is precisely why naming the evaluator and the exact configuration matters more than naming the benchmark alone: two credible labs running "Terminal-Bench 2.1" produced meaningfully different numbers.

  • Mercor's Apex Agents benchmark, which scores economically valued real-work tasks, puts Opus 5 in the lead at 43.5%, marginally ahead of Fable 5.1.

  • Cognition's Apex SWE benchmark, focused specifically on software engineering tasks, shows Fable 5.1 leading by a small margin over Opus 5.

  • Epoch AI's Capabilities Index, a broad composite measure, has Opus 5 trailing both Fable 5.1 and GPT-5.6 on the combined score.

Put together, the picture is more nuanced than "Opus 5 matches the flagship" or "Opus 5 is clearly behind." It depends on the task category and which lab is measuring, exactly the detail a vendor's own launch chart never tells you. We ran the same benchmark-labelling exercise on GPT-6 Astra; and if the real decision in front of you is Claude versus ChatGPT rather than which Claude model, our ChatGPT vs Claude comparison covers that directly.

Claude Opus 5 Pricing: What It Actually Costs

The headline number is simple: $5 per million input tokens, $25 per million output tokens, on the Claude API. But the sticker price is not the whole bill. Here is everything that actually changes what you pay, verified against Anthropic's own pricing page:

Pricing factor

What it does to your bill

Base rate

$5 / MTok input, $25 / MTok output

5-minute cache write

$6.25 / MTok, 1.25x base input

1-hour cache write

$10 / MTok, 2x base input

Cache read (hit)

$0.50 / MTok, 10% of base input, the standard multiplier

Batch API

50% off both directions: $2.50 / MTok input, $12.50 / MTok output

Fast mode (research preview)

$10 / MTok input, $50 / MTok output, double the standard rate

Data residency (inference_geo "us")

1.1x multiplier across every pricing category

Long-context (1M window)

No premium; a 900K-token request bills at the same rate as a 9K one

Fast mode is worth flagging: it is research preview only, first-party API only, on Opus 5 and Opus 4.8 only, and unavailable on the Batch API or any partner cloud. Real, but not a default.

Cache reads on Opus 5 cost the standard 10% of base input, not the discounted rate Fable 5.1 gets. That single detail matters more than it looks, and it sets up the model comparison below.

The Full Claude Model Family, Compared

This is the table most single-model coverage skips, and it is the one that actually decides which model fits a given job. Every current, generally available Claude model, side by side.

Model

API ID

Input

Output

Context

Max output

Knowledge cutoff

Latency

Thinking

Retirement floor

Claude Fable 5.1

claude-fable-5-1

$10 / MTok

$50 / MTok

1M tokens

128K tokens

Jun 2026

Slower

Adaptive, always on

Not sooner than Sep 1, 2027

Claude Opus 5

claude-opus-5

$5 / MTok

$25 / MTok

1M tokens

128K tokens

May 2026

Moderate

Adaptive

Not sooner than Jul 24, 2027

Claude Sonnet 5

claude-sonnet-5

$2 / MTok

$10 / MTok

1M tokens

128K tokens

Jan 2026

Fast

Adaptive

Not sooner than Jun 30, 2027

Claude Haiku 4.5

claude-haiku-4-5-20251001

$1 / MTok

$5 / MTok

200K tokens

64K tokens

Feb 2025

Fastest

Extended

Not sooner than Oct 15, 2026

A few footnotes that change how that table actually plays out, per Anthropic's model deprecation page and its pricing docs:

  • Batch API requests are 50% off across every model in this table, no exceptions.

  • Prompt cache reads cost 10% of base input on every model here except Claude Fable 5.1, where they cost just 2.5%.

  • Data residency (inference_geo "us") adds a 1.1x multiplier on all pricing categories for Claude 4.6-generation models and later, including everything in this table.

  • The 1M context window carries no long-context premium on any of these models. A 900,000-token request bills at the same per-token rate as a 9,000-token one.

There is a fifth current Claude model, Claude Mythos 5.1, Fable 5.1's gated sibling. It shares Fable 5.1's pricing and its 2.5% cache-read discount, but it is invite-only through Anthropic's limited-availability program, not something you can simply start calling through the standard API, so it is left out of the table above rather than presented as an option most readers can actually pick.

Which Claude Model Should You Use?

Anthropic's own docs give a one-line answer: start with Opus 5 for most workloads, move up to Claude Fable 5.1 only once your own evals show Opus 5 at high effort genuinely falls short. Good default. But two details above change the math for specific situations, and they are easy to miss if you only look at sticker prices.

Claude Haiku 4.5 is the outlier in every dimension except price. It is the cheapest model in the lineup, but it is not simply a discounted version of the same product. Its context window is 200K tokens, not 1M. Its max output is 64K tokens, not 128K. Its knowledge cutoff sits at February 2025, over a year behind the rest of the current lineup. And its retirement floor is October 15, 2026, five weeks from today, not deep into 2027 like every other model in this table. Cheapest is not the same thing as "same product, less money."

Claude Fable 5.1's discounted cache reads can flip the naive price comparison entirely. Its base rate is double Opus 5's, but its cache reads cost 2.5% of base input instead of the standard 10%, and Anthropic's own published estimate puts the real-world savings at roughly 25% versus list price on typical workloads, and up to about 45% on heavily agentic ones. Agentic work repeatedly revisits the same codebase, system prompt, tool definitions, and conversation history, exactly the kind of content prompt caching is built to reuse. On a workload like that, Fable 5.1 can end up cheaper in practice than its list price suggests, even against a model with half its sticker rate.

A simple way to actually apply this:

  • Default to Opus 5 for new integrations, agentic coding, and enterprise workloads. It is Anthropic's own starting recommendation, and independent testing largely backs it up.

  • Move up to Fable 5.1 once your own evals show Opus 5 at high effort falls short on your specific task, or your workload is cache-heavy and agentic enough that the 2.5% cache rate meaningfully changes your bill.

  • Drop to Sonnet 5 for high-volume, latency-sensitive, or cost-sensitive production work where the last few points of capability matter less than speed and price.

  • Only use Haiku 4.5 for simple, high-throughput tasks, and only if you can live with the 200K context ceiling and a retirement date that is already five weeks out.

  • Mythos 5.1 is not a practical option for most readers; it shares Fable 5.1's pricing but requires an invite.

Where Claude Opus 5 Is Strong, and Where It Falls Short

Where it's strong:

  • Lands within about a point of Fable 5.1 on Vals AI's independent overall index, at exactly half the price.

  • Ranks first on 14 of 27 independent Vals AI leaderboards, including a 97.00% on SWE-bench Verified.

  • Full 1M token context and 128K max output at the same price Anthropic charged for the previous Opus generation.

  • The effort toggle lets you dial cost against capability per request instead of switching models entirely.

  • 300,000 output tokens on the Batch API, useful for large asynchronous jobs.

  • The most safety-focused Opus model Anthropic has shipped, per the company, with tighter cybersecurity guardrails and an automatic fallback instead of a bare refusal.

Where it falls short:

  • The "matches Fable 5.1" claim does not hold everywhere: Cognition's Apex SWE benchmark and Epoch AI's Capabilities Index both still put Fable 5.1 ahead, particularly on coding-heavy work.

  • Standard 10% cache-read pricing, not Fable 5.1's 2.5%, so cache-heavy agentic workloads may not actually be cheaper on Opus 5 once you run real numbers.

  • Not available on Claude Platform on AWS, only Bedrock, Google Cloud, and Microsoft Foundry.

  • Anthropic's own reporting shows one internal coding benchmark peaking at medium effort and declining at higher settings, so maxing out effort is not automatically right.

  • Fast mode is research preview only, doubles the price, and skips the Batch API and every partner cloud.

  • Moderate latency; Sonnet 5 and Haiku 4.5 both respond faster for latency-sensitive use.

How to Actually Get Claude Opus 5

Claude Opus 5 is available today through the Claude API (model ID or alias claude-opus-5) and through three partner clouds: Amazon Bedrock (anthropic.claude-opus-5), Google Cloud/Vertex AI (claude-opus-5), and Microsoft Foundry (claude-opus-5). It is not currently listed for Claude Platform on AWS, so check Anthropic's own availability table before assuming it is there if that is your setup.

Looking for it inside a claude.ai subscription rather than through the API? The consumer-plan tier mapping was not part of the developer docs checked for this review, so confirm current access directly on claude.ai's pricing page rather than assuming it lines up with API availability.

Who Should Use Claude Opus 5, and Who Should Skip It

Use Claude Opus 5 if:

  • You are starting a new Claude integration and want Anthropic's own recommended default.

  • You are building agentic coding tools or enterprise automations and need strong capability without Fable 5.1's list price.

  • Your work touches scientific or biology-adjacent research, an area Anthropic specifically positions Opus 5 for.

  • You need the full 1M token context window without a long-context pricing penalty.

Skip it, at least as your default, if:

  • Your own evals show Fable 5.1 meaningfully outperforms Opus 5, particularly on coding-heavy work where Cognition's and Epoch AI's numbers still favor Fable 5.1.

  • Your workload is cache-heavy and agentic enough that Fable 5.1's discounted cache reads change which model is actually cheaper.

  • You need the fastest possible response time for a latency-sensitive product; Sonnet 5 or Haiku 4.5 fit better.

  • You specifically need Claude Platform on AWS; Opus 5 is not there yet.

Final Verdict on Claude Opus 5

Claude Opus 5 largely earns the "start here" recommendation Anthropic gives it in its own docs. Independent testing from Vals AI backs the "close to Fable 5.1" claim more often than not, at exactly half the price, with a genuinely useful effort toggle for dialing cost against capability per request. That is a real, verifiable improvement over the previous Opus generation, not just a price cut.

But "close" is not "equal." Cognition and Epoch AI, two independent evaluators, still put Fable 5.1 ahead on coding-heavy work, and Fable 5.1's discounted cache reads shrink the real price gap further on agentic workloads than the sticker prices suggest. The honest recommendation is the one Anthropic itself publishes: default to Opus 5, run your own evals, and only pay double for Fable 5.1 once you have genuinely measured that Opus 5 falls short for what you are building. That is just what "worth it" means once you check the numbers instead of retyping a launch chart.

Curious to see how it performs?

Try

Claude

Now

GOT ANY QUESTIONS LEFT?

Which Claude model should I use?

How much does Claude Opus 5 cost?

Is Claude Opus 5 better than Claude Fable 5.1?

What is Claude Opus 5's context window?