Claude Opus 5

Anthropic's flagship model: 1M context, 97% SWE-bench, unchanged Opus price

Claude Opus 5 Review: 1M Context, 97% SWE-bench 2026

Claude Opus 5 is Anthropic's July 2026 flagship model: 97.0% SWE-bench Verified, 94.1 GPQA Diamond, 1M context, priced same as Opus 4.8 at $5/$25 per 1M tokens.

Claude Opus 5 launched July 24, 2026 with a 1 million token context window as the default tier and a new xhigh reasoning effort mode. It replaces Opus 4.8 as Anthropic's flagship Opus model for long-running agentic coding and research workloads.

Claude Opus 5 is Anthropic's flagship large language model, released July 24, 2026, scoring 97.0% on SWE-bench Verified and 94.1 on GPQA Diamond. It ships a 1 million token context window by default and targets long-horizon agentic coding and research work at the same price as Opus 4.8.

Provider: Anthropic · Family: Claude 5

More about Anthropic on HokAI

Context window: 1,000,000 tokens · Max output: 128,000

Input modalities: text, image, pdf, tool-calls · Output: text, tool-calls

About Claude Opus 5

Claude Opus 5 is Anthropic's flagship large language model, released July 24, 2026 as the successor to Claude Opus 4.8. It sits atop Anthropic's Claude 5 lineup alongside Claude Sonnet 5 and Claude Fable 5, built for teams that need the deepest reasoning and the longest agentic runs the company ships. The model surfaced briefly under the codename Honeycomb EAP inside Cursor's model picker on July 9, 2026 before Anthropic pulled it, then completed a full rollout two weeks later. On SWE-bench Verified, Opus 5 scores 97.0%, the highest recorded result published on that benchmark to date. On SWE-bench Pro, a harder agentic-coding suite, it reaches 79.2%, up 10 points from Opus 4.8's 69.2%. GPQA Diamond, the graduate-level science reasoning benchmark, comes in at 94.1%. On ARC-AGI-3 at high reasoning effort, Opus 5 posts 30.16%, roughly 20 times Opus 4.8's score on the same test and about 4 times GPT-5.6 Sol Max. Artificial Analysis puts its composite Intelligence Index at 61 in max-effort mode, well above the 32 median for comparable frontier models. The model ships with a 1 million token context window, both the default and the ceiling, so there is no smaller-context variant to choose. Max output is 128,000 tokens through the standard Messages API, extending to 300,000 tokens through the Message Batches API with a beta header. Extended thinking is on by default, a change from Opus 4.8 where it shipped off by default, and Opus 5 adds a new xhigh reasoning-effort tier above low, medium, high, and max. Opus 5 takes text, image, and PDF input and returns text output, with tool-calling built into both directions. It reads charts, documents, and diagrams and is strongest at replicating UI and frontend visuals when given tools to crop, re-render, and check its own output iteratively rather than in one pass. Tool lists can now be added or removed mid-conversation without breaking the prompt cache. Anthropic also reports that Opus 5 verifies its own work without being asked, so verification instructions written for earlier models can trigger redundant double-checking on Opus 5. Pricing is unchanged from Opus 4.8 at $5 per million input tokens and $25 per million output tokens. Fast Mode runs at roughly 2.5x output speed for $10 input and $50 output per million tokens. Batch API discounts standard pricing 50%, to $2.50 input and $12.50 output per million tokens. Prompt caching now activates from 512 cached tokens, down from 1,024 on Opus 4.8, and cuts input cost up to 90% on repeat-context workloads. Opus 5 is reachable through the direct Anthropic API under the model ID claude-opus-5, through Amazon Bedrock as anthropic.claude-opus-5, and through Google Vertex AI as claude-opus-5. It is available today on Claude Pro, Max, Team, and Enterprise plans, and has replaced Opus 4.8 as the default Opus model and the default model on Claude Max. Training data runs through a cutoff of May 2026. Anthropic's own automated behavioral audit gives Opus 5 the lowest misalignment score the company has recorded for any model, including its own more restricted Mythos 5 variant, and the model cooperates with misuse attempts less than any prior Claude release. The Opus 5 system card also documents a less flattering result from the UK AI Security Institute: given a simulated, standard-but-not-hardened enterprise network, Opus 5 completed the intrusion attack path in 8 of 10 attempts, a sign of how far its offensive cyber capability has advanced alongside its coding scores. Opus 5 fits teams running long autonomous coding or research agents that need million-token recall and are willing to pay Opus-tier pricing for it. It is a poor fit for latency-sensitive interactive products: output speed of roughly 54 tokens per second and time-to-first-token near 2.4 to 3.9 seconds trail lighter models built for chat, where Claude Sonnet 5 or Claude Haiku 4.5 are the better match. Switching from Opus 4.8 also brings breaking changes: thinking: disabled combined with effort xhigh or max now returns a 400 error, and the speed: fast parameter inherited from Opus 4.7 has been removed, so requests using it now error instead of falling back to standard speed.

Pricing

$5 per 1M input tokens, $25 per 1M output tokens, unchanged from Opus 4.8. Fast Mode is $10/$50 per 1M for roughly 2.5x faster output. Batch API is 50% off at $2.50/$12.50. Prompt caching activates from 512 tokens and cuts input cost up to 90%.

Key Features

Pros

Cons

Benchmarks

Frequently Asked Questions

How much does Claude Opus 5 cost per 1M tokens?

Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, unchanged from Opus 4.8. Fast Mode runs at $10/$50 per million for roughly 2.5x faster output, and the Batch API cuts standard pricing 50% to $2.50/$12.50. Prompt caching now activates from 512 tokens and can cut input cost up to 90% on repeat-context workloads.

How does Claude Opus 5 compare on benchmarks vs GPT-5.6?

Claude Opus 5 scores 97.0% on SWE-bench Verified and 30.16% on ARC-AGI-3 at high effort, about 4 times GPT-5.6 Sol Max's result on the same test. GPT-5.6 stays competitive on some general knowledge benchmarks, but Opus 5 leads clearly on agentic coding and abstract reasoning tasks.

Is Claude Opus 5 open source or proprietary?

Claude Opus 5 is fully proprietary. Weights are not published, and access is API-only through Anthropic's own platform, AWS Bedrock, or Google Vertex AI under Anthropic's Commercial Terms of Service.

Does Claude Opus 5 train on user data?

No. Anthropic's API does not train on customer inputs or outputs by default, and zero-data-retention is available for eligible enterprise customers. Consumer Claude.ai usage follows a separate retention policy set in account settings.

Who is Claude Opus 5 best for and who should avoid it?

Teams running long autonomous coding or research agents that need million-token recall get the most from Opus 5's benchmark jump at unchanged pricing. Latency-sensitive chat or voice products should look at Claude Sonnet 5 or Haiku 4.5 instead, since Opus 5's roughly 54 tokens-per-second output trails lighter models built for interactive use.

More AI Models on HokAI

Visit Claude Opus 5 Official Page