Best suited to engineering teams running autonomous coding agents and to finance or legal teams processing document sets above 500,000 tokens without chunking. Opus 4.7 now leads on raw benchmarks, so treat Opus 4.6 as a stable option for existing integrations rather than a new build.
Claude Opus 4.6 is Anthropic's February 2026 flagship large language model, scoring 91.3% on GPQA Diamond. It pairs a context window spanning 1 million tokens with adaptive thinking and full computer-use support, aimed at long-horizon agentic coding and multi-step research rather than short conversational replies.
Where it sits
- $10.00/M$ per 1M tokensBlended price (3:1)Lower is better#51 / 64peer median $1.70/Mvendor price, checked by HokAI
- --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
- 80.8%% solvedSWE-bench VerifiedHigher is better#9 / 28peer median 78.3%per source, see benchmark scores
- 91.3%% correctGPQA DiamondHigher is better#14 / 44peer median 88.3%per source, see benchmark scores
Pricier than 84% of the 64 GA models with a published price, in the top third on SWE-bench Verified (rank 9 of 28), and one of 22 that document a zero-data-retention option.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Anthropic · Family: Claude 4
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, pdf, tool-calls · Output: text, tool-calls
About Claude Opus 4.6
Claude Opus 4.6 is Anthropic's fourth-generation flagship large language model, released on February 5, 2026, built on a dense Transformer architecture with extended and adaptive reasoning support. It succeeded Claude Opus 4.5 (November 2025) and sits below the later Claude Opus 4.7 in Anthropic's lineup. Anthropic has not disclosed a parameter count; independent analysts estimate it in the high hundreds of billions. The API model ID is claude-opus-4-6, with a training data cutoff of August 2025 and a reliable knowledge cutoff of May 2025.
Its defining architectural addition is a context window spanning 1 million tokens, made generally available at standard pricing on March 13, 2026, with a 128,000 token max output per synchronous request (300,000 on the Batch API with a beta header). Inputs are text and image, including PDFs, charts, and screenshots; there is no audio or video support, and output is text plus tool-call results only. The model has full computer use capability and native function calling with parallel tool calls.
Safety training combines reinforcement learning from human feedback, reinforcement learning from AI feedback, and Constitutional AI alignment. The system card, running over 200 pages, documents independent red teaming from METR, Apollo Research, and the UK AISI; Anthropic deploys the model under ASL-3 safeguards and describes it as approaching the ASL-4 autonomy threshold. Anthropic classifies Opus 4.6 as a general-purpose AI with systemic risk obligations under the EU AI Act.
Anthropic replaced Opus 4.6 with Claude Opus 4.7 on April 16, 2026, which raised SWE-bench Verified to 87.6% at the same per-token pricing tier. Opus 4.6 remains available as a legacy model with no announced retirement date, though Anthropic recommends new projects build on 4.7 instead.
Pricing
$5.00 per 1M input tokens, $25.00 per 1M output tokens. Cache hits cost $0.50 per 1M tokens (10% of input rate). Cache writes cost $6.25/1M (5-min) or $10/1M (1-hour). Batch API: 50% off ($2.50/$12.50). Fast mode beta: 6x rates ($30/$150). US-only inference adds a 1.1x multiplier.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.150 | $0.025 | $0.175 |
| Support reply | $0.010 | $0.0075 | $0.018 |
| One coding agent run | $1.00 | $0.500 | $1.50 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- 1 Million Token Context Window: Generally available at standard pricing since March 2026. Scores 76% on MRCR v2 retrieval across the full 1M window, versus 18.5% for Sonnet 4.5. No per-token surcharge beyond the standard rate.
- Adaptive Thinking: The model decides on its own how deep to reason based on task difficulty, cutting wasted thinking tokens on easy sub-tasks inside an agentic loop.
- Effort Controls: Four API-level settings (low, medium, high, max) trade intelligence for speed and cost per request without switching models.
- Agent Teams: Multiple independent Opus 4.6 instances run in parallel, each with its own context window, coordinated by a lead agent; introduced in Claude Code in February 2026.
- Computer Use: Reads screens and takes actions (clicks, typing) via the computer use API; Anthropic called it its most accurate computer-use model at release.
- Context Compaction: A beta API feature that auto-summarizes older context so a single conversation can run well past the 1M token window.
- Fast Mode (Beta): Runs the same model at 2.5x output throughput for 6x standard pricing, gated behind a waitlist and unavailable on the Batch API.
Pros
- 80.8% SWE-bench Verified at release, ahead of every other generally available model on real-world software engineering.
- GDPval-AA score is 144 Elo points above GPT-5.2 on economically valuable finance and legal work.
- 1M-token context ships at standard pricing with no long-context surcharge, unlike most rivals' tiered context pricing.
- Lowest over-refusal rate among recent Claude models, per Anthropic's own system card.
Cons
- No native audio or video input or output; voice products need a separate ASR/TTS pipeline.
- Proprietary closed weights rule out self-hosting, fine-tuning, or air-gapped deployment.
- Roughly 1.6-second time-to-first-token, which compounds across long multi-step agentic loops.
- Anthropic has already shipped a faster successor (Opus 4.7), so new projects are steered away from 4.6.
Benchmarks
- MATH: 93% vendor-reported · 05 Feb 2026 — Competition maths problems, % solved.
- AIME 2025: 75.5% vendor-reported · 05 Feb 2026 — Competition-level maths problems from the 2025 exam, % solved.
- ARC-AGI 2: 69.2% vendor-reported · 05 Feb 2026 — Abstract visual puzzles built to resist memorisation, % solved.
- HumanEval: 95% vendor-reported · 05 Feb 2026 — Small programs that must pass hidden tests, % passing.
- GPQA Diamond: 91.3% vendor-reported · 05 Feb 2026 — PhD-level science questions that are hard to search for, % correct.
- SWE-bench Verified: 80.8% vendor-reported · 05 Feb 2026 — Real GitHub issues fixed end to end, % solved.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does Claude Opus 4.6 actually cost?
Claude Opus 4.6 costs $5 per million input tokens and $25 per million output tokens on the standard API. Prompt caching drops input costs to $0.50 per million tokens on a cache hit, while cache writes run $6.25 per million tokens for a 5-minute cache or $10 for a 1-hour cache. The Batch API halves both rates for asynchronous jobs, and a beta Fast Mode charges 6x standard rates for 2.5x faster output.
Can you use Claude Opus 4.6 without paying?
No. Opus 4.6 has no free tier and is billed strictly per token, whether through the direct API or a cloud reseller like Bedrock, Vertex, or Foundry. Developers testing the model without a paid Anthropic account typically start through a cloud provider's own trial credits instead.
Which models compete with Claude Opus 4.6 in 2026?
GPT-5.2 is the closest rival on professional finance and legal knowledge work, where Opus 4.6 holds the GDPval-AA lead. Claude Sonnet 4.6 undercuts it on price for teams that do not need the largest context window. Anthropic's own Claude Opus 4.7 has since replaced it as the recommended choice for new agentic-coding projects.
Is Claude Opus 4.6 better than GPT-5.2?
On the GDPval-AA professional finance and legal knowledge-work benchmark, yes; Opus 4.6 holds the lead there and ships a context window spanning 1 million tokens at standard pricing, with no long-context surcharge. On raw coding ability the picture flips slightly: GPT-5.3-Codex, not GPT-5.2, is the model that edges out Opus 4.6 on SWE-bench Verified, by roughly 1 point. Check OpenAI's own page for current GPT-5.2 pricing, since it moves independently of Anthropic's rates.
What does it take to start using Claude Opus 4.6?
Sign up for an Anthropic API key at console.anthropic.com, or reach the model through AWS Bedrock, Vertex AI on Google Cloud, or Azure via Microsoft Foundry if you already run infrastructure there. The model ID is claude-opus-4-6 on the direct API and anthropic.claude-opus-4-6-v1 on Bedrock. SDKs exist for Python, TypeScript, Java, Go, and Ruby, though Go and Ruby do not support the Foundry deployment path.
Top Alternatives
- Claude Opus 4.7: Pick Opus 4.6 only to keep an existing integration stable; pick Opus 4.7 for new work at 87.6% SWE-bench.
- GPT-5.2: Pick Opus 4.6 for finance and legal knowledge work, where it leads GPT-5.2 by 144 Elo points on GDPval-AA.
- Claude Sonnet 4.6: Pick Sonnet 4.6 for cost-sensitive workloads; pick Opus 4.6 when you need the full 1M-token context window.