Sonnet 5 fits engineering teams running high-volume agentic coding and computer-use work who want near-Opus quality at Sonnet pricing. It's the wrong pick for voice or video-first products, since it has no native audio or video I/O, and for configs still pinned to non-default sampling parameters. Max output runs to 128,000 tokens on the synchronous API.
Claude Sonnet 5, Anthropic's mid-tier flagship model released in 2026, sits between Haiku 4.5 and Opus 4.8 in Anthropic's current lineup. It leads on GPQA Diamond at 96.2% and posts 84.7% on ARC-AGI-2, running adaptive hybrid reasoning by default for agentic coding and computer-use automation.
Where it sits
- $6.00/M$ per 1M tokensBlended price (3:1)Lower is better#48 / 64peer median $1.70/Mvendor price, checked by HokAI
- --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
- 82.1%% solvedSWE-bench VerifiedHigher is better#8 / 28peer median 78.3%per source, see benchmark scores
- 96.2%% correctGPQA DiamondHigher is better#1 / 44peer median 88.3%per source, see benchmark scores
Pricier than 78% of the 64 GA models with a published price, in the top third on SWE-bench Verified (rank 8 of 28), and one of 22 that document a zero-data-retention option.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Anthropic · Family: Claude Sonnet 5
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, pdf, tool-calls · Output: text, tool-calls
About Claude Sonnet 5
Claude Sonnet 5 is Anthropic's fifth-generation Sonnet-tier model, released June 30, 2026. Anthropic positions it as a drop-in upgrade for Sonnet 4.6 that closes most of the coding gap to Opus 4.8 while keeping Sonnet's speed and price. It is a dense, hybrid-reasoning Transformer with adaptive thinking on by default; Anthropic has not disclosed a parameter count.
Sonnet 5 uses the tokenizer introduced with Opus 4.7, so identical text now produces roughly 30% more tokens than under Sonnet 4.6's tokenizer. Anthropic's reliable knowledge cutoff for the model is January 2026.
Deployment spans the direct Claude API, AWS Bedrock, Microsoft Foundry, and Google Vertex AI (listed as coming soon at launch), plus day-one access through Claude Code and GitHub Copilot. Anthropic holds SOC 2 Type I and Type II, ISO 27001:2022, and ISO/IEC 42001:2023 certifications, signs HIPAA Business Associate Agreements, and offers Zero Data Retention outside the Batch API.
Safety partners for the June 30, 2026 system card include METR, Apollo Research, and the UK AI Security Institute. Anthropic reports lower misalignment, hallucination, and sycophancy rates than Sonnet 4.6, plus better prompt-injection resistance; the model does not cross the automated AI R&D capability threshold and shows only limited bio/chem uplift.
Pricing
$3 per 1M input tokens and $15 per 1M output tokens at standard pricing, identical to Sonnet 4.6's rate card. An introductory discount of $2/$10 per 1M tokens runs through August 31, 2026, offsetting the new tokenizer's higher per-request token counts. Prompt cache reads bill at 0.1x the standard input rate, and the Batch API is 50% cheaper but excluded from Zero Data Retention, with batch data held up to 29 days.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.090 | $0.015 | $0.105 |
| Support reply | $0.0060 | $0.0045 | $0.010 |
| One coding agent run | $0.600 | $0.300 | $0.900 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- 82.1% SWE-bench Verified: First model to clear the 80% mark on SWE-bench Verified, the benchmark most agentic coding teams weigh first.
- 1M-Token Context Window: 128,000 max output tokens on the synchronous Messages API, extending to 300,000 on the Batch API.
- Computer Use: Scores 81.2% on OSWorld-Verified and 80.4% on Terminal-Bench 2.1, both sharp jumps from Sonnet 4.6's 78.5% and 67.0%.
- Adaptive Thinking: Runs by default on every request unless explicitly disabled, replacing Sonnet 4.6's opt-in extended-thinking mode.
- Prompt Caching: Cache reads are billed well below the standard input rate, cutting costs on repeat-context agentic workloads.
Pros
- Anthropic's own launch results put it ahead of Gemini 3.1 Pro and GPT-5.4 on real-world coding-agent evaluations.
- 40% cheaper than Opus 4.8 on the standard rate card, even before the current introductory discount kicks in.
- Computer-use automation jumped sharply on Anthropic's own tests, with OSWorld-Verified and Terminal-Bench scores both rising over Sonnet 4.6.
Cons
- No native audio or video input/output, so voice and video apps still need a separate model in the stack.
- New tokenizer inflates token counts roughly 30% versus Sonnet 4.6, requiring max_tokens and cost budgets to be re-tuned on migration.
- Sampling params (temperature, top_p, top_k) now error at non-default values, breaking configs carried over directly from Sonnet 4.6.
Benchmarks
- ARC-AGI 2: 84.7% vendor-reported · 30 Jun 2026 — Abstract visual puzzles built to resist memorisation, % solved.
- GPQA Diamond: 96.2% vendor-reported · 30 Jun 2026 — PhD-level science questions that are hard to search for, % correct.
- SWE-bench Verified: 82.1% vendor-reported · 30 Jun 2026 — Real GitHub issues fixed end to end, % solved.
- AA Intelligence Index: 53 cited: Artificial Analysis · 30 Jun 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
How much does Claude Sonnet 5 cost in 2026?
Standard pricing charges $3 for a million input tokens and $15 for a million output tokens, matching Sonnet 4.6's rate card. Through August 31, 2026, Anthropic discounts that to $2/$10 per 1M tokens to cushion the higher counts from the new tokenizer, and cached reads bill at a tenth of the normal input price. Batch requests run 50% cheaper than the synchronous API but give up Zero Data Retention eligibility.
Is Claude Sonnet 5 free to use?
Claude Sonnet 5 has no standalone free tier: it's the default model on claude.ai's Free and Pro consumer plans, subject to standard message-rate limits, but full-rate API access requires a paid API account with no free monthly quota. Teams wanting the model at zero marginal cost only get it through the consumer chat interface, not the developer API.
What are the best alternatives to Claude Sonnet 5?
Inside Anthropic's own lineup, Claude Opus 4.8 suits teams that need the highest coding ceiling and can absorb higher per-token pricing, while Claude Sonnet 4.6 fits teams not yet ready to re-tune token budgets around the new tokenizer. Outside Anthropic, Gemini 3.1 Pro and GPT-5.4 both compete directly on agentic coding and long-context tasks.
Claude Sonnet 5 or GPT-5.4: which should you pick?
Claude Sonnet 5 posts a higher SWE-bench Verified score than GPT-5.4, the first model in either lineup to clear the 80% mark on that benchmark. GPT-5.4 counters with a native Computer Use API built into OpenAI's own stack, while Sonnet 5's computer-use gains come from Anthropic's OSWorld-Verified and Terminal-Bench improvements instead. Teams already standardized on OpenAI's tooling have less reason to switch; teams choosing fresh get a real coding-benchmark edge from Sonnet 5.
How do you get started with Claude Sonnet 5?
Getting started just needs an Anthropic API key and a request to the Messages API with the model field set to claude-sonnet-5. Existing Sonnet 4.6 integrations need any explicit temperature, top_p, or top_k overrides removed, since non-default values now return a 400 error. Claude Code and GitHub Copilot users can switch model versions directly in their settings without touching API code.
Top Alternatives
- Claude Opus 4.8: Pick Opus 4.8 if you need the strongest coding ceiling on SWE-bench Pro; pick Sonnet 5 for near-Opus quality at a lower token price.
- Claude Sonnet 4.6: Pick Sonnet 5 over Sonnet 4.6 for the computer-use and coding-benchmark jump at the same price; stay on 4.6 only if pinned to non-default sampling parameters.
- Gemini 3.1 Pro: Pick Gemini 3.1 Pro for Google Cloud-native deployment; pick Sonnet 5 for the higher SWE-bench Verified score and native computer-use benchmarks.
- GPT-5.4: Pick GPT-5.4 if you're standardized on OpenAI's native Computer Use API; pick Sonnet 5 for the SWE-bench lead and Anthropic's compliance certifications.
HokAI guides covering Claude Sonnet 5
- Claude Sonnet 4.6 vs Claude Sonnet 5: What Migrating Changes: Claude Sonnet 4.6 is active with no retirement date, yet Anthropic already calls it legacy. Here's what changes, and what breaks, when you migrate to Sonnet 5.
- What Qwen Is Actually For, Now It's Priced Below Claude and GPT-5.6: Alibaba priced Qwen3.8-Max at $2/$6 per million tokens, beating Claude and GPT-5.6, but its benchmarks are self-reported and the open-weights license isn't out.
- Best AI for Coding Questions Free in 2026: 8 Real Options, Compared: Which free AI actually answers coding questions well in 2026? DeepSeek, ChatGPT, Claude, Gemini and Qwen, all researched and compared side by side today.
- Our llms.txt Has 708 Links. Google Ignores All of Them.: We fetched eight AI directories' llms.txt on 23 August 2026. Five had none. Ours lists 708 links that match our API exactly, and Google still ignores it.
- Cursor Composer 2.5 vs. Claude Code: Which AI Coding Agent Should You Use in 2026?: Cursor Composer 2.5 vs Claude Code compared on 2026 pricing, benchmarks, and context windows now that Claude Code defaults to Sonnet 5 and Opus 5, not Opus 4.6.