Modelfrontier LLM2mo ago

Claude Opus 4.8 review

Anthropic's most capable Opus-tier model, built for complex agentic coding and enterprise work — a 1M-token context, adaptive thinking, and 128K max output, hosted only via the Claude API.

By Ravi Menon · Models & Benchmarks EditorVerified 2026-07-24
Maker
Anthropic
Launched
May 28, 2026
Pricing
paid
Visit official site
Firstlook

Our verdict

Opus 4.8 is Anthropic's top-tier model, aimed squarely at complex agentic coding and enterprise work: a 1M-token context, adaptive thinking, 128K output, and fast mode when you need speed. The trade-off is cost and lock-in — at $5 / $25 per million tokens it's roughly 2.5x the price of Sonnet 5, and it's hosted-only with no weights to own. If you want maximum capability on hard, long-horizon tasks and the budget follows, it's the pick. We have not run it hands-on yet, so no score.

First look — our read from the docs and sources below; not yet hands-on tested.

Claude Opus 4.8 sits at the top of Anthropic's widely-available lineup. In Anthropic's own words it's the model "for complex agentic coding and enterprise work" — the one you reach for when a task is long, multi-step, and expensive to get wrong. Below it is Sonnet 5 for balanced speed-and-intelligence; above it, in limited availability, is the Fable/Mythos tier.

The spec sheet is deliberately unglamorous and mostly shared with the rest of the current generation: a 1-million-token context window, up to 128K output tokens, a January 2026 knowledge cutoff, and adaptive thinking where the model itself decides how hard to think (tuned by an effort dial that defaults to high). What sets Opus apart isn't a headline number — it's that Anthropic positions it as its most capable model for autonomous, long-horizon agentic work.

Who it's for

Teams doing the hard stuff: large refactors, overnight coding runs, multi-tool agents, and enterprise workflows where capability matters more than the per-token bill. The 1M-token context is billed at standard rates, so whole-repo and long-document tasks don't carry a long-context surcharge. If you need lower latency and can absorb the premium, a fast-mode research preview roughly doubles the price for significantly faster output.

Who should skip it

Anyone cost-sensitive or latency-sensitive. At $5 / $25 per million tokens Opus 4.8 is about 2.5x the price of Sonnet 5, and Sonnet 5 is faster by default and now near-frontier on coding — for most production workloads it's the more economical call. Opus is also closed and hosted-only: there are no weights to download, run offline, or audit, so if ownership or on-prem is a hard requirement, this isn't the model.

This is a first look built from Anthropic's official documentation, not our own evaluation — so no score from us yet. On paper, Opus 4.8 is the capability ceiling of the general Claude lineup; whether you need that ceiling (and its price) over Sonnet 5 is the real decision.

Provider

Provideranthropicclaude-opus-4-8· Proprietary (closed, hosted only)

Specs & key facts

What it isAnthropic's most capable Opus-tier model, for complex agentic coding and enterprise work[src]
Context window1M tokens in · 128K out[src]
Pricing$5 / $25 per MTok (in / out)[src]
ThinkingAdaptive thinking (effort defaults to high); manual extended thinking not supported[src]
Knowledge cutoffJan 2026[src]
LicenseProprietary — hosted only, no downloadable weights[src]

Capabilities

CodingYes (complex agentic coding — its headline use case)
Reasoning / mathYes (adaptive thinking)
Agentic / tool useYes (enterprise agentic work)
Self-hostNo (hosted only)
Hosted APIYes (Claude API)
Commercial useYes (commercial API terms)
Fast modeYes (research preview, premium pricing)

How to use it

  1. 1Call it via the Claude API with the model ID `claude-opus-4-8` — there are no weights to self-host.
  2. 2Leave adaptive thinking on and tune `effort` (it defaults to `high`); use `xhigh` for the hardest coding and agentic tasks.
  3. 3Lean on the 1M-token context for whole-repo, long-document and long-horizon agentic work — it's billed at standard rates.
  4. 4Reach for the fast-mode research preview only when latency matters more than the premium price.

Pricing

Claude API (standard)

$5 / $25 per MTok

$5 per million input tokens, $25 per million output tokens. Cache reads bill at $0.50 / MTok. Hosted only — there are no weights to download.

Fast mode (research preview)

$10 / $50 per MTok

Optional premium tier for significantly faster output on Opus 4.8; billed across the full context window.

Closed, hosted model priced per token on the Claude API. Standard pricing $5 in / $25 out per million tokens; the full 1M-token context window is billed at standard rates (no long-context premium). Verified against Anthropic's docs 2026-07-24.

Pros & cons

Pros

  • Anthropic's most capable model outside the limited-availability Fable/Mythos tier.
  • 1M-token context at standard pricing — no long-context premium.
  • Adaptive thinking with an `effort` dial (defaults to high) for hard reasoning and agentic work.
  • Fast mode available when you need lower latency and will pay the premium.

Cons

  • Expensive: $5 / $25 per MTok, about 2.5x Sonnet 5's standard price.
  • Closed and hosted only — no weights to download, run offline, or audit.
  • Manual extended thinking with a fixed token budget is not supported (adaptive only).
  • Higher latency than Sonnet 5; the speed gap only closes with paid fast mode.

Alternatives

FAQ

Sources

  1. 1.Model ID, positioning, 1M context, 128K max output, adaptive thinking, Jan 2026 knowledge/training cutoff, effort defaults to highhttps://platform.claude.com/docs/en/about-claude/models/overviewVerified 2026-07-24
  2. 2.Pricing $5 / MTok input, $25 / MTok output; cache read $0.50 / MTok; fast mode $10 / $50 per MTokhttps://platform.claude.com/docs/en/about-claude/pricingVerified 2026-07-24
  3. 3.Release date May 28, 2026 (System Card cover date + announcement)https://www.anthropic.com/news/claude-opus-4-8Verified 2026-07-24
  4. 4.Benchmarks (Table 8.1.A): SWE-bench Verified 88.6%, SWE-bench Pro 69.2%, SWE-bench Multilingual 84.4%, Terminal-Bench 2.1 74.6% (high effort), OSWorld-Verified 83.4%, BrowseComp 84.3% single-agent, HLE 49.8% no tools / 57.9% with tools, AutomationBench 15.5%, GPQA Diamond 93.6%https://www.anthropic.com/claude-opus-4-8-system-cardVerified 2026-07-24

More coverage

News & first-looks about this release. Coming soon.
Head-to-head comparisons. Coming soon.