Claude Opus 4.7 is built for engineering teams running long, multi-file coding agents and enterprise teams analyzing huge documents, replacing the previous Opus generation as Anthropic's default heavy-workload model. It even reaches 54.7% on Humanity's Last Exam with tool use, though teams building real-time chat or voice products should pick a faster, cheaper Claude model instead.
Claude Opus 4.7 is Anthropic's flagship generally available large language model, scoring 94.2% on GPQA Diamond graduate-level science reasoning. Released April 16, 2026, its main differentiator is a full 1M-token context window with no long-context surcharge, paired with the first high-resolution vision support in the Claude lineup.
Where it sits
- $10.00/M$ per 1M tokensBlended price (3:1)Lower is better#51 / 64peer median $1.70/Mvendor price, checked by HokAI
- --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
- 87.6%% solvedSWE-bench VerifiedHigher is better#7 / 28peer median 78.3%per source, see benchmark scores
- 94.2%% correctGPQA DiamondHigher is better#5 / 44peer median 88.3%per source, see benchmark scores
Pricier than 84% of the 64 GA models with a published price, in the top third on SWE-bench Verified (rank 7 of 28), and one of 22 that document a zero-data-retention option.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Anthropic · Family: Claude 4
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, pdf, tool-calls · Output: text, tool-calls
About Claude Opus 4.7
Claude Opus 4.7 is Anthropic's most capable generally available model, released on April 16, 2026. Built on a dense Transformer architecture (not MoE) with undisclosed parameter count, it is the fourth major Opus iteration in the Claude 4 generation, following Opus 4.1, 4.5, 4.6, and now 4.7. Anthropic positioned it specifically to close the gap between their generally-available lineup and frontier research previews, with a deliberate focus on agentic coding, long-horizon autonomous tasks, and high-fidelity vision understanding. It replaces Claude Opus 4.6 as Anthropic's recommended starting point for any complex workload.
On benchmarks, Claude Opus 4.7 posts 87.6% on SWE-bench Verified, a 6.8-percentage-point gain over Opus 4.6's 80.8% and ahead of GPT-5.4 on the same leaderboard. SWE-bench Pro, a harder multi-file variant, jumped from 53.4% on Opus 4.6 to 64.3% on Opus 4.7, leapfrogging GPT-5.4 (57.7%) and Gemini 3.1 Pro (54.2%). GPQA Diamond, which tests PhD-level scientific reasoning across physics, chemistry, and biology, is close to saturation: Opus 4.7 scores 94.2%, GPT-5.4 scores 94.4%, and Gemini 3.1 Pro scores 94.3%. For tool use, MCP-Atlas scores 77.3%, ahead of Opus 4.6 at 75.8%, GPT-5.4 at 68.1%, and Gemini 3.1 Pro at 73.9%. Computer use on OSWorld-Verified reached 78.0%, ahead of GPT-5.4 at 75.0%. Humanity's Last Exam without tools sits at 46.9%, and with tools at 54.7%.
The context window is 1,000,000 tokens, the full 1M provided at standard per-token rates with no long-context premium. Maximum output via the synchronous Messages API is 128,000 tokens. On the Batch API with the output-300k-2026-03-24 beta header, Opus 4.7 can produce up to 300,000 output tokens per request. A single request can include up to 600 images or PDF pages. Note: Opus 4.7 uses a new tokenizer that processes text with up to 35% more tokens than Opus 4.6 for equivalent input, so effective costs on existing prompts can increase even though the rate card is unchanged.
Modalities supported as inputs are text and images (PDF documents are parsed as images). Output is text only. No native audio input or output is available. Opus 4.7 is the first Claude model with high-resolution image support, raising maximum image resolution from 1568px / 1.15MP to 2576px / 3.75MP. This is more than a 3x increase in pixel area, which meaningfully improves performance on dense UIs, technical diagrams, chemical structure reading, and computer use screenshots. Image coordinates now map 1:1 to actual pixels, removing the need for scale-factor math in agentic loops. Tool use and function calling are fully supported with native parallel tool calls, client-side tool schemas, and server-side tools including web search, web fetch, and code execution. Computer use is available via the beta API. Task budgets (beta) allow the model to self-regulate token spend across a full agentic loop. The new xhigh effort level sits between high and max, giving finer control over reasoning depth versus speed.
Pricing carries over unchanged from the previous Claude Opus generation. Prompt caching, Batch API processing, and US-only inference each apply their own discount or surcharge on top of the base rate, detailed in the pricing section below.
Claude Opus 4.7 is available via the Anthropic API (model ID: claude-opus-4-7), AWS Bedrock (ID: anthropic.claude-opus-4-7-v1:0), Google Cloud Vertex AI (ID: claude-opus-4-7@20260416), and Microsoft Foundry (via the Azure AI Foundry catalog). Authentication is via API key on Anthropic direct, AWS IAM on Bedrock, GCP IAM on Vertex, and Azure subscription billing on Foundry. Microsoft Foundry regions are restricted to East US2 and Sweden Central. AWS Bedrock supports both global endpoints (dynamic routing) and regional endpoints (geo-guaranteed routing). Cloud providers add a margin on top of Anthropic's base rates for enterprise SLAs and region pinning.
Safety evaluation for Opus 4.7 was published in the system card released April 16, 2026. The model uses Constitutional AI and RLHF alignment. In multi-turn safety evaluations, Opus 4.7 successfully identified escalating harm patterns even when individual prompts appeared benign. In agentic settings, it is better than Opus 4.6 at refusing malicious instructions and resisting prompt injection in Claude Code and computer use contexts. Cyber capabilities are roughly similar to Opus 4.6; an external evaluation by the UK AI Security Institute found it could not complete their full cyber range, unlike the more capable Mythos Preview. Evaluation awareness was detected in about 9% of transcripts, triggered by inconsistencies in mocked tool results. Cybersecurity refusals are new real-time safeguards; teams doing legitimate security work can apply to the Cyber Verification Program.
Training data cutoff is January 2026, meaning the model has reliable knowledge of events through that date. Anthropic does not train on API inputs by default. Enterprise zero-retention plans are available. The model is SOC 2 Type II compliant, HIPAA-eligible, ISO 27001 certified, and GDPR compliant. EU AI Act classification is general-purpose AI with systemic risk obligations. API inputs and outputs are retained for 30 days for abuse monitoring and deleted unless flagged, unless the enterprise zero-retention option is enabled.
Pricing
$5.00 per 1M input tokens, $25.00 per 1M output tokens. Prompt caching: 5-min cache write at $6.25/MTok, 1-hour cache write at $10.00/MTok, cache reads at $0.50/MTok (90% savings). Batch API: 50% discount at $2.50 input / $12.50 output per MTok. US-only inference adds 1.1x multiplier. A new tokenizer may meaningfully increase effective token counts versus the previous generation, so raise max_tokens accordingly.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.150 | $0.025 | $0.175 |
| Support reply | $0.010 | $0.0075 | $0.018 |
| One coding agent run | $1.00 | $0.500 | $1.50 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- High-Resolution Vision (3.75MP): First Claude model to support 2576px / 3.75MP images, a 3.3x increase over previous models. Pixel coordinates map 1:1 to actual image pixels, removing scale-factor math from computer use agents.
- 1M Token Context at Standard Pricing: Full 1 million token context window with no long-context premium. Up to 600 images or PDF pages per request. Reliable knowledge through January 2026.
- Adaptive Thinking with xhigh Effort: New xhigh effort level between high and max for finer reasoning/latency control. Adaptive thinking (replacing extended thinking budgets) is the only reasoning mode and outperforms explicit budget setting in internal evals.
- Task Budgets for Agentic Loops: Beta feature: give Claude an advisory token target across a full agentic session. Model self-regulates via a running countdown, finishing tasks gracefully without runaway token spend.
- Leading Tool-Use Benchmark (MCP-Atlas): MCP-Atlas 77.3%, highest among generally available models. Supports native parallel tool calls, client-side schemas, and server-side tools including web search ($10/1K searches), web fetch (free), and code execution.
- Prompt Caching (90% Savings): Cache system prompts, tool definitions, and document context at $0.50/MTok (10% of base input rate). Stack with Batch API for up to 95% savings vs uncached standard requests.
Pros
- Leads all generally available models on real-world coding benchmarks, translating into fewer failed steps in long autonomous agent sessions.
- Ingests a full codebase or long document corpus in a single call without paying a separate long-context surcharge.
- Tool-use orchestration leads every other generally available model on MCP-Atlas, which matters for agents that chain many tool calls in one session.
Cons
- Time to first token of ~19 seconds and output speed of 42.2 t/s, well below the 59.8 t/s frontier median, ruling it out for latency-sensitive applications.
- New tokenizer silently increases token counts by up to 35%, meaning existing prompts can cost more despite unchanged per-token rates.
- Breaking API changes at launch: temperature/top_p/top_k removed, extended thinking budget syntax removed. Requires migration work for all Opus 4.6 integrations.
Benchmarks
- MMLU: 91.5% vendor-reported · 16 Apr 2026 — General-knowledge exam across 57 subjects, % correct.
- GPQA Diamond: 94.2% vendor-reported · 16 Apr 2026 — PhD-level science questions that are hard to search for, % correct.
- SWE-bench Verified: 87.6% vendor-reported · 16 Apr 2026 — Real GitHub issues fixed end to end, % solved.
- Humanity's Last Exam: 46.9% vendor-reported · 16 Apr 2026 — Expert-written questions across many fields, % correct.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
How much does Claude Opus 4.7 cost in 2026?
Claude Opus 4.7 costs $5.00 per million input tokens and $25.00 per million output tokens. Prompt caching cuts cached reads to a fraction of the base input rate, and the Batch API halves both rates for asynchronous jobs. A new tokenizer processes the same text into more tokens than before, so real bills can rise even though the per-token rate hasn't changed.
Does Claude Opus 4.7 have a free plan?
No, Claude Opus 4.7 has no free tier and no documented trial credit or capped free usage. Anthropic's lower-cost Sonnet tier is the closer fit for teams that want to try Claude-family models without paying Opus rates.
Which tools compete with Claude Opus 4.7 in 2026?
Claude Opus 4.7 competes most directly with GPT-5.4 and Gemini 3.1 Pro among frontier generally available models. For less demanding, latency-sensitive workloads, Anthropic's own lower-cost Sonnet tier is the closer fit within the Claude lineup.
What separates Claude Opus 4.7 from GPT-5.4?
On SWE-bench Pro, a harder multi-file coding benchmark, Opus 4.7 scores 64.3% versus GPT-5.4's 57.7%, and Opus 4.7 also leads on MCP-Atlas tool use. GPT-5.4 stays competitive on general reasoning: GPQA Diamond scores are close to tied. Opus 4.7 also ships a higher-resolution vision pipeline for agentic screenshot work.
How do you get started with Claude Opus 4.7?
Claude Opus 4.7 is available immediately through the Anthropic API, AWS Bedrock, Google Vertex AI, or Microsoft Foundry; direct API access just needs an account and an API key. Teams migrating from the prior Opus generation need to update calls that used the old temperature, top_p, or extended-thinking-budget parameters, since Opus 4.7 rejects them. First-time users should also budget for the new tokenizer's higher token counts when sizing prompts.
Top Alternatives
- GPT-5.4: Pick Opus 4.7 for stronger coding, tool-use, and computer-use benchmarks; pick GPT-5.4 if you need marginally higher scores on general scientific reasoning tests.
- Gemini 3.1 Pro: Pick Opus 4.7 for its stronger SWE-bench Pro and MCP-Atlas tool-use scores; pick Gemini 3.1 Pro if near-identical GPQA Diamond reasoning is enough and you prefer Google's platform.