Claude Opus 4.8 replaces earlier Claude Opus versions for teams running long autonomous coding and research loops. Its 1M-token context window handles entire codebases or multi-document research in one pass. Best for agentic coding, enterprise research, and math-heavy reasoning; skip it for real-time voice or on-device deployment.
Claude Opus 4.8 is Anthropic's flagship large language model, scoring 93.6% on GPQA Diamond for graduate-level reasoning. Released May 28, 2026, it adds native computer-use automation for desktop tasks, positioning it for long-horizon agentic coding, research synthesis, and enterprise tool orchestration workflows.
Where it sits
- $10.00/M$ per 1M tokensBlended price (3:1)Lower is better#51 / 64peer median $1.70/Mvendor price, checked by HokAI
- 79 tok/stokens/sOutput speedHigher is better#23 / 39peer median 90 tok/scited: Artificial Analysis
- 88.6%% solvedSWE-bench VerifiedHigher is better#6 / 28peer median 78.3%per source, see benchmark scores
- 93.6%% correctGPQA DiamondHigher is better#8 / 44peer median 88.3%per source, see benchmark scores
Pricier than 84% of the 64 GA models with a published price, in the top third on SWE-bench Verified (rank 6 of 28), and one of 22 that document a zero-data-retention option.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Anthropic · Family: Claude 4
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, pdf, tool-calls · Output: text, tool-calls
About Claude Opus 4.8
Claude Opus 4.8 is Anthropic's most capable generally available model, released on May 28, 2026. It is a dense Transformer with parameter count undisclosed (estimated ~600B based on capability footprint). The model builds on Opus 4.7's foundation with targeted improvements in agentic coding, mathematical reasoning, and behavioral consistency. Positioned as the flagship in Anthropic's Claude Opus family, it is designed to power long-running, multi-step autonomous workflows in enterprise environments, research, and developer tooling. Per-token pricing holds steady from the previous generation, with a new fast mode added alongside it for latency-sensitive work.
On frontier coding benchmarks, Opus 4.8 tops competitors across both dimensions. SWE-bench Verified hits 88.6% (up from 87.6% on Opus 4.7, ahead of GPT-5.5's estimates). SWE-bench Pro (the harder agentic coding benchmark) reaches 69.2%, a jump of 4.9 points from Opus 4.7's 64.3% and 10.6 points ahead of GPT-5.5's 58.6%. On mathematical reasoning (USAMO 2026), the model achieves 96.7%, a stunning leap from Opus 4.7's 69.3%. GPQA Diamond (graduate-level reasoning) stands at 93.6%. Terminal-Bench 2.1 (multi-turn coding and debugging) reaches 74.6%, up from 66.1% on Opus 4.7. Humanity's Last Exam scores 49.8% without tools and 57.9% with tools, the highest in the field. These benchmarks signal dominance in reasoning and code-centric tasks.
The context window is 1M tokens on the Claude API, AWS Bedrock, and Google Vertex AI (200k on Microsoft Foundry). Standard max output is 128k tokens, with a beta 300k output option via the output-300k-2026-03-24 header on Batch API. Long-context recall (needle-in-haystack) is reliable above 100k tokens: the model maintains 99%+ accuracy on fact retrieval deep in the context window. The model architecture also preserves prompt caching at a lower token minimum than Opus 4.7 required, cutting the cost of reusing static system prompts or knowledge bases.
Multimodal inputs include text, images, and PDFs natively. The model processes up to 100 images per request and reads PDF documents directly without extraction. Tool use is native: function calling follows OpenAI-compatible schema with support for parallel tool calls. Computer use (screen reading, keyboard/mouse control) is verified on OSWorld at 83.4%, enabling desktop automation. However, there is no native audio input, audio output, or video processing: audio workflows must pair Opus 4.8 with a separate ASR/TTS model. Structured output (JSON mode) is available via the messages API.
Deployment options span the major cloud platforms. The Claude API (api.anthropic.com) offers direct access with API Key auth. AWS Bedrock hosts the model with IAM-based auth across us-east-1, us-west-2, eu-central-1, ap-northeast-1, and other regions. Microsoft Foundry provides access (200k context only) with Azure identity auth, and Google Vertex AI serves the model with GCP IAM auth. No self-hosting or air-gapped deployment option exists: Opus 4.8 is API-only, proprietary weights. Client libraries ship for Python, JavaScript, TypeScript, Java, Go, and Ruby.
Safety posture is balanced with measurable alignment improvements. Training data cutoff is January 2026. Constitutional AI and RLHF alignment are standard. Anthropic has not yet published a dedicated Opus 4.8 system card (most recent is Opus 4.6, February 2026), but internal evaluations show Opus 4.8 is 4x less likely than Opus 4.7 to overlook code flaws it has written, achieving 0% on 'uncritically reporting flawed results' evaluations, a first for the Opus line. The model refuses clear harms (CSAM, weapons, malware code) and shows improved honesty, with calibrated refusals on edge cases. Inputs and outputs are retained for 30 days for abuse monitoring; zero-retention option available on enterprise plans. SOC 2 Type II, ISO 27001, HIPAA-eligible, and GDPR-compliant.
Behavioral improvements over Opus 4.7 include better long-context handling (fewer compactions, faster recovery), calibrated reasoning effort (aligned behavior at each effort level), and tool triggering (fewer skipped required tool calls). The model shows stronger consistency in long-horizon agentic loops, critical for multi-step workflows running 8+ hours. New feature: Dynamic Workflows in Claude Code (Enterprise/Team/Max plans, research preview) lets Opus 4.8 break a task into up to 1,000 concurrent subagents and check their output before reporting back. The fast mode's price drop makes the faster variant more accessible for latency-sensitive workloads (real-time customer support, instant analytics).
Who should use Opus 4.8: agentic coding teams running long autonomous loops, enterprise research teams analyzing 100K+ documents, mathematical problem-solving workflows, code review and refactoring at scale, customer-support automation requiring multi-turn tool use, and teams leveraging prompt caching on static contexts (system prompts, knowledge bases). Who should avoid: teams needing on-device inference (proprietary, no self-host), voice-first assistants (no native audio), or sub-500ms latency (standard mode ~850ms p50). GPT-5.5 leads on pure speed and real-time constraints. Gemini 3.1 Deep Think excels on mathematical theory. DeepSeek V3 and open-source models (Llama 4) are viable for cost-constrained or air-gapped setups.
Pricing
Standard: $5.00 input, $25.00 output per 1M tokens. Cached input (1024+ token minimum): $0.50 per 1M. Fast mode: $10.00/$50.00 (2.5x faster tokens). Batch API: 50% off ($2.50/$12.50, async 24h).
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.150 | $0.025 | $0.175 |
| Support reply | $0.010 | $0.0075 | $0.018 |
| One coding agent run | $1.00 | $0.500 | $1.50 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- 1M Context Window: Reliably recalls facts and code above 100K tokens. Verified 99%+ needle-in-haystack accuracy. Fully GA on Claude API, Bedrock, Vertex AI.
- Adaptive Thinking: Configurable reasoning depth per request. Balances accuracy and latency. Recommended 2K-8K token budgets for most tasks.
- Prompt Caching: 90% cost savings on cached input (1,024+ token min). Ideal for agentic loops with static system prompts or knowledge bases.
- Computer Use & Tool Orchestration: Screen reading, keyboard/mouse control verified at 83.4% (OSWorld). Parallel tool calls, structured JSON output, function calling.
- Dynamic Workflows: Claude Code (Enterprise/Team/Max, research preview). Plan work, spawn up to 1,000 parallel subagents, verify outputs, report results.
Pros
- Leads GPT-5.5 and Gemini 3.1 Pro on both SWE-bench Verified and the harder SWE-bench Pro agentic coding benchmark.
- Leaps ahead on math reasoning with USAMO 2026, a dramatic jump from the previous generation.
- First Claude model to score a perfect 0% on code-flaw detection evals, sharply cutting the rate of missed bugs versus Opus 4.7.
- 1M context with reliable long-range recall, steep caching discounts, and adaptive thinking for efficient reasoning.
- Pricing holds steady from Opus 4.7's generation, with a notably cheaper fast mode and new Dynamic Workflows support.
Cons
- No native audio or video input; must integrate separate ASR/TTS or vision models.
- Proprietary weights, API-only; cannot self-host, deploy air-gapped, or fine-tune.
- Standard latency ~850ms p50 (slower than GPT-5.5 mini's ~320ms); fast mode trades cost for speed.
- System card not yet published; internal evals strong but third-party red-team data sparse.
- Parameter count undisclosed; estimated ~600B but unconfirmed.
Benchmarks
- Usamo 2026: 96.7 vendor-reported · 28 May 2026
- LMArena Elo: 1412 vendor-reported · 28 May 2026 — Rating from blind human votes on which answer is better.
- GPQA Diamond: 93.6% vendor-reported · 28 May 2026 — PhD-level science questions that are hard to search for, % correct.
- LMArena rank: #2 vendor-reported · 28 May 2026 — Position on the blind human-preference leaderboard; #1 is best.
- SWE-bench Pro: 69.2% vendor-reported · 28 May 2026 — Harder, longer real-repository coding tasks, % solved.
- SWE-bench Verified: 88.6% vendor-reported · 28 May 2026 — Real GitHub issues fixed end to end, % solved.
- Terminal-Bench 2.1: 74.6% vendor-reported · 28 May 2026 — Multi-step tasks completed in a real command line, % solved.
- Humanity's Last Exam: 49.8% vendor-reported · 28 May 2026 — Expert-written questions across many fields, % correct.
- AA Intelligence Index: 72 cited: Artificial Analysis — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- AA blended price: $12.50/M cited: Artificial Analysis — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
- Output speed: 79 tok/s cited: Artificial Analysis — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What are Claude Opus 4.8's pricing plans in 2026?
Claude Opus 4.8 costs $5 per 1M input tokens and $25 per 1M output tokens on standard mode, with cached input dropping to $0.50 per 1M. Fast mode (research preview) roughly doubles the price for quicker output tokens, and Batch API cuts standard pricing by 50% for requests that can run asynchronously.
Can you use Claude Opus 4.8 without paying?
Claude Opus 4.8 does not offer a free plan: access is billed per token, whether through Anthropic's own Claude API, AWS Bedrock, Microsoft Foundry, or Google Vertex AI. Cost-conscious teams can lean on prompt caching for cheaper repeat context, or step down to Claude Sonnet 4.6 for lighter workloads.
What are Claude Opus 4.8's closest competitors?
GPT-5.5 is the pick for real-time, latency-sensitive responses that Opus 4.8's standard mode was not built for. Gemini 3.1 Deep Think has an edge on pure theorem-proving and symbolic math. DeepSeek V3 and open-weight models such as Llama 4 suit teams that need on-device or air-gapped deployment, which Opus 4.8's proprietary API cannot offer.
How does Claude Opus 4.8 compare to GPT-5.5 in 2026?
Claude Opus 4.8 pulls ahead on agentic coding depth, reaching 69.2% on SWE-bench Pro against GPT-5.5's 58.6%, and extends that lead into math reasoning with 96.7% on USAMO 2026. GPT-5.5 answers back with faster response times, better suited to real-time and voice-first products than Opus 4.8's standard mode.
What does it take to start using Claude Opus 4.8?
Get an API key through the Claude API (api.anthropic.com) using API Key authentication, or provision access via an existing AWS Bedrock, Microsoft Foundry, or Google Vertex AI account. Official SDKs cover Python, JavaScript, TypeScript, Java, Go, and Ruby, so a first API call is typically a same-day setup. New users should start on standard mode before testing the pricier fast mode for latency-sensitive workloads.
Top Alternatives
- Claude 4.7 Opus: 4.7 Opus keeps the cheaper fast-mode rate for teams already built around it; 4.8 earns its upgrade with real gains in coding and math benchmarks.
- GPT-5.5: GPT-5.5 wins on sub-500ms real-time responses; Opus 4.8 wins on deeper agentic coding and math reasoning benchmarks.
- Gemini 3.1 Pro: Teams standardized on Google Cloud lean toward Gemini 3.1 Pro; teams chasing higher agentic coding benchmark scores lean toward Opus 4.8.
HokAI guides covering Claude Opus 4.8
- Claude Sonnet 4.6 vs Claude Sonnet 5: What Migrating Changes: Claude Sonnet 4.6 is active with no retirement date, yet Anthropic already calls it legacy. Here's what changes, and what breaks, when you migrate to Sonnet 5.