Claude Opus 5.5 launched September 22, 2026 with a 128,000 token max output and a June 2026 knowledge cutoff, an upgrade to Claude Opus 5 rather than a new model generation. It suits teams already on Opus-tier agentic workloads who want the latest coding and computer-use gains without switching model families.
Claude Opus 5.5 is Anthropic's large language model, released September 22, 2026, scoring 89.9% on SWE-bench Pro against Claude Opus 5's 79.2%. It ships a 1 million token context window by default and targets long-running agentic coding, computer use, and knowledge work at a lower per-token price than its predecessor.
Where it sits
- $8.00/M$ per 1M tokensBlended price (3:1)Lower is better#53 / 67peer median $1.71/Mvendor price, checked by HokAI
- --tokens/sOutput speedHigher is better-- / 41peer median 90 tok/scited: Artificial Analysis
- --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
- --% correctGPQA DiamondHigher is better-- / 44peer median 88.3%per source, see benchmark scores
Pricier than 79% of the 67 GA models with a published price, and one of 25 that document a zero-data-retention option.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Anthropic · Family: Claude 5
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, pdf, tool-calls · Output: text, tool-calls
About Claude Opus 5.5
Claude Opus 5.5 is Anthropic's Opus-class large language model, released September 22, 2026 as an upgrade to Claude Opus 5 rather than a new model generation. Anthropic's own system card describes it as matching or exceeding Claude Fable 5.1 and Claude Mythos 5.1 on many evaluations, and Anthropic now recommends it as the starting point for most workloads, positioning Claude Fable 5.1 for the harder reasoning and longest agentic runs instead. Architecture details and parameter count are undisclosed, consistent with every proprietary Claude release to date. The API model ID is claude-opus-5-5, a dateless pinned snapshot rather than a preview alias.
On Anthropic's own capability evaluation table, Opus 5.5 scores 89.9% on SWE-bench Pro, ahead of Claude Opus 5's 79.2% and Claude Fable 5.1's 81.2%. It reaches 93.9% on SWE-bench Multilingual and 61.4% on SWE-bench Multimodal, both improvements over Opus 5's 89.5% and 59.4%. On Terminal-Bench 4.0 at xhigh reasoning effort it scores 66.4%, well ahead of Opus 5's 52.3%. On Humanity's Last Exam with tools it reaches 67.7%, versus 63.6% for Opus 5 and 65.6% for Fable 5.1. On GDPval-AA v2.1, a professional-work evaluation Anthropic reports on an Elo-style scale, it posts 1846 against Opus 5's 1708. Artificial Analysis puts its composite Intelligence Index at 58. Notably, the September 22, 2026 system card does not publish SWE-bench Verified, GPQA Diamond, AIME 2025, MMLU-Pro, or ARC-AGI-2 scores for Opus 5.5, benchmarks Anthropic reported for Claude Opus 5 three months earlier; the evaluation suite for this release leans toward agentic and professional-work benchmarks instead.
Opus 5.5 ships with a 1 million token context window, the same default-and-ceiling tier as Claude Opus 5, Claude Fable 5.1, and Claude Sonnet 5. Max output through the synchronous Messages API is 128,000 tokens, extending to 300,000 tokens through the Message Batches API with the output-300k-2026-03-24 beta header. Anthropic's own long-context evaluation, ProgramBench, runs Opus 5.5 across context lengths up to the full 1M-token window.
The model accepts text and image input, including PDF documents through Anthropic's Files API and document support, and returns text output only, with tool-calling built into both directions. Anthropic confirms support for computer use through the new computer_toolset_20260801 toolset, code execution, and server-side web search and web fetch tools. Adaptive thinking is always on; the effort parameter (low, medium, high, xhigh, max) is now the sole control for reasoning depth, and its default on Opus 5.5 is medium, down from high on Claude Opus 5. Anthropic also reports sharper reading of dense charts, diagrams, and screenshots without needing prompt-side vision workarounds built for earlier models.
Standard pricing runs $4 per million input tokens and $20 per million output tokens, a fifth less than Claude Opus 5's $5 and $25. Prompt cache reads cost $0.20 per million tokens, 5% of the base input price, activating from a 512-token minimum. Cache writes cost $5 per million tokens at a 5-minute TTL or $8 per million at a 1-hour TTL. The Batch API cuts standard pricing 50%, to $2 and $10. A research-preview Fast Mode, available on the Claude API only, runs at roughly double the standard price ($8 input, $40 output) with the fast-mode-2026-02-01 beta header for faster output.
Opus 5.5 is reachable through the direct Anthropic API as claude-opus-5-5, through Amazon Bedrock as anthropic.claude-opus-5-5 with separate US-geo and Global routing endpoints, through Claude Platform on AWS and Microsoft Foundry as claude-opus-5-5, and through Google Cloud's Vertex AI under the same ID. Anthropic's Commercial Terms of Service govern access; official SDKs cover Python, TypeScript, Java, Go, C#, PHP, and Ruby.
Anthropic's Responsible Scaling Policy and Frontier Compliance Framework evaluations treat Opus 5.5 as having CB-1 chemical and biological capabilities, relating to synthesis of non-novel weapons, but not CB-2 capabilities for novel weapons; it is deployed with the same expanded biological safeguards applied to Claude Fable 5 and Fable 5.1, plus a new biology safety classifier alongside the existing cybersecurity one. Its AI R&D capability sits at or slightly above Claude Mythos 5.1's without a sustained 2x acceleration in Anthropic's internal development-pace metrics. On cyber evaluations it meets or exceeds Mythos 5.1 and Opus 5, and Anthropic has temporarily widened its jailbreak safety margin while tuning classifier false-positive rates; no critical-severity jailbreak was found. Pre-deployment testing involved METR, Frontier Design, the US Center for AI Standards and Innovation, the UK AI Security Institute, Gray Swan, and Irregular. In raw agentic-safety testing without production safeguards, Opus 5.5 assisted with dual-use and benign security tasks at the highest rate of the models Anthropic evaluated, but also refused malicious requests at the lowest rate.
Opus 5.5 suits teams already running long agentic coding or computer-use workloads on Opus-tier models who want the coding and professional-work gains at a lower per-token price; existing Claude Opus 5 integrations should budget migration time for four breaking API changes, covered below. It is a weaker fit for audio or video agent workloads, since it accepts only text and image input, and for latency-sensitive interactive chat or voice products, where Anthropic's own comparison table rates its latency Moderate against Claude Sonnet 5's Fast and Claude Haiku 4.5's Fastest.
Training data runs through a June 2026 cutoff, both for reliable knowledge and the broader training data range, one month later than Claude Opus 5's May 2026 cutoff. Anthropic commits to keeping Opus 5.5 available on its own platforms no sooner than September 22, 2027. As with other current Claude models, the API does not train on customer inputs or outputs by default, and zero-data-retention is available for eligible enterprise customers.
Four breaking changes separate Opus 5.5 from Claude Opus 5: thinking can no longer be disabled, so requests must drop thinking: disabled and use the effort parameter instead; forced tool_choice (any or a named tool) now returns a 400 error, requiring a switch to auto plus strict tool use; thinking blocks are now bound to the model and conversation that produced them, with stricter enforcement on accounts created after August 31, 2026; and the legacy computer_20251124 computer-use tool is rejected on the Claude API and Google Cloud, though it still works unchanged on Amazon Bedrock. A further non-breaking change worth flagging in production: short text notes the model used to send between tool calls now arrive as thinking blocks rather than text blocks, so streaming interfaces that surfaced them as progress updates need to set a thinking.display value to keep receiving them.
Pricing
$4 per 1M input tokens and $20 per 1M output tokens, 20% below Claude Opus 5's $5/$25. Cache reads are $0.20 per 1M (5% of base input price, from a 512-token minimum). Cache writes are $5/M at a 5-minute TTL or $8/M at a 1-hour TTL. Batch API is 50% off standard at $2/$10. A research-preview Fast Mode (Claude API only) runs $8/$40 for faster output.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.120 | $0.020 | $0.140 |
| Support reply | $0.0080 | $0.0060 | $0.014 |
| One coding agent run | $0.800 | $0.400 | $1.20 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- 1M Token Context Window: Default and maximum context length of 1,000,000 tokens, with Anthropic's ProgramBench evaluation run across the full window.
- Adaptive Thinking, Effort-Controlled: Reasoning depth is set entirely by the effort parameter, from low up through the new xhigh and max tiers, since thinking can no longer be turned off.
- Computer Use Toolset: Agentic desktop control via the computer_toolset_20260801 toolset, reaching 81.8% partial completion on OSWorld 2.0.
- Lower Pricing Across Every Tier: Standard, cached, and Batch pricing all run below Claude Opus 5's rates across every tier, from pay-as-you-go to async batch jobs.
- Task Budgets and Mid-Conversation Tool Definition: Task budgets and a beta inline-tool-definition feature let you add or change tools mid-conversation without resending schemas or breaking the prompt cache.
Pros
- Scores 89.9% on SWE-bench Pro, ahead of Opus 5's 79.2% and Fable 5.1's 81.2%, per Anthropic's own system card.
- Runs cheaper than Claude Opus 5 per token while Anthropic reports output generation over 30% faster.
- 1M token context window ships as the default tier, with support for the Files API, PDF input, and prompt caching for repeat-context workloads.
Cons
- Migrating from Opus 5 costs real engineering time: the thinking-disabled option is gone, forced tool_choice ("any"/"tool") now errors, and both changes require re-testing agent loops built for the previous model.
- No native audio or video input; accepts text, image, and PDF only.
- In Anthropic's own pre-safeguard agentic testing, it refused malicious requests at the lowest rate among the models evaluated, even as it assisted with dual-use security tasks at the highest rate, a tradeoff worth extra review for security-sensitive agent deployments.
Benchmarks
- GDPval-AA v2: 1,846 vendor-reported · 23 Sep 2026 — Real knowledge-work deliverables judged against professionals, run by Artificial Analysis.
- SWE-bench Pro: 89.9% vendor-reported · 23 Sep 2026 — Harder, longer real-repository coding tasks, % solved.
- Humanity's Last Exam: 67.7% vendor-reported · 23 Sep 2026 — Expert-written questions across many fields, % correct.
- AA Intelligence Index: 58 cited: Artificial Analysis · 23 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does Claude Opus 5.5 cost per 1M tokens in 2026?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, 20% below Claude Opus 5's $5/$25. Cache reads cost $0.20 per million tokens, 5% of the base input price, and cache writes cost $5/M at a 5-minute TTL or $8/M at a 1-hour TTL. The Batch API cuts standard pricing 50% to $2/$10, and a research-preview Fast Mode on the Claude API runs $8/$40 for faster output.
How does Claude Opus 5.5 compare on benchmarks to Claude Opus 5 and Claude Fable 5.1?
Claude Opus 5.5 scores 89.9% on SWE-bench Pro in Anthropic's own system card, ahead of Claude Opus 5's 79.2% and Claude Fable 5.1's 81.2%. On Humanity's Last Exam with tools it reaches 67.7%, versus 63.6% for Opus 5 and 65.6% for Fable 5.1. Anthropic states Opus 5.5 matches or exceeds Fable 5.1 on many evaluations, though the system card does not publish SWE-bench Verified, GPQA Diamond, or MMLU-Pro scores for this comparison.
Is Claude Opus 5.5 open source or built on open weights?
No. Claude Opus 5.5 is fully proprietary: Anthropic has not published its weights or architecture details, and access is API-only through the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry under Anthropic's Commercial Terms of Service.
Does Anthropic train Claude Opus 5.5 on customer API data?
No, not by default. Anthropic's API does not train on customer inputs or outputs unless a customer explicitly opts in, and zero-data-retention is available for eligible enterprise customers. Consumer claude.ai usage follows a separate retention policy set in account settings.
Who should use Claude Opus 5.5, and who should look elsewhere?
Teams running long agentic coding or computer-use workloads get the most from Opus 5.5's SWE-bench Pro and Terminal-Bench gains at 20% lower per-token pricing than Opus 5. Teams building audio or video agents should look elsewhere, since Opus 5.5 accepts only text and image input. Latency-sensitive chat or voice products are also a weaker fit, since Anthropic rates its comparative latency "Moderate" against Claude Sonnet 5's "Fast" and Claude Haiku 4.5's "Fastest".