Sonnet 5.5 is Anthropic's faster partner to Opus 5.5, with 1,000,000 tokens of context room and a June 2026 knowledge cutoff. It suits teams running well-scoped coding and office-document work at volume, while complex, open-ended jobs still belong on Opus 5.5.
Claude Sonnet 5.5 is a large language model from Anthropic's Sonnet line, launched on September 28, 2026, and it scores 81.3% on SWE-Bench Pro against Claude Sonnet 5's 63.2%. It takes text and image input, returns text, and runs adaptive thinking across five effort levels for everyday coding, documents, slides, and spreadsheets.
Where it sits
- $4.00/M$ per 1M tokensBlended price (3:1)Lower is better#48 / 72peer median $1.70/Mvendor price, checked by HokAI
- 142 tok/stokens/sOutput speedHigher is better#13 / 44peer median 87 tok/scited: Artificial Analysis
Pricier than 68% of the 72 GA models with a published price, rank 13 of 44 on output speed as cited from Artificial Analysis, and one of 28 that document a zero-data-retention option.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Anthropic · Family: Claude 5
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, pdf, tool-calls · Output: text, tool-calls
About Claude Sonnet 5.5
Released September 28, 2026 as the second model in the Claude 5.5 family, six days after Claude Opus 5.5, Claude Sonnet 5.5 is the latest Sonnet-class model from Anthropic, whose full lineup sits on the Anthropic models page. It upgrades Claude Sonnet 5 rather than starting a new generation. Anthropic positions it as the faster, lower-cost partner to Opus 5.5: strongest at well-scoped everyday tasks, bug fixing, and polished documents, slides, and spreadsheets, while Opus 5.5 stays the choice for complex, open-ended work that needs sustained judgment. Architecture and parameter count are undisclosed. The API model ID is claude-sonnet-5-5, and Anthropic says a Claude Haiku 5.5 will join the family in the coming weeks.
Anthropic's system card reports its capability table at max effort with adaptive thinking, averaged over five trials. On it, Sonnet 5.5 scores 81.3% on SWE-Bench Pro (Sonnet 5: 63.2%, Opus 5.5: 89.9%) and 90.3% on SWE-Bench Multilingual (78.3% and 93.9%). Terminal-Bench 4.0 is the standout: 70.6% against 10.3% for Sonnet 5 and 66.4% for Opus 5.5, though the Opus figure is at xhigh effort and the Sonnet run had safeguards on, with a fallback model answering 1.2% of requests. FrontierCode 1.1 is where it trails: 46.2% at max effort against 54.4% for Opus 5.5 and 49.3% for GPT-6 Sol. On the same table, Anthropic's system card lists Claude Mythos 5.1 at 60.9%, Claude Fable 5.1 at 55.8%, and Claude Opus 5 at 52.3% on Terminal-Bench 4.0, all below Sonnet 5.5. Anthropic notes the Sonnet FrontierCode score is higher at xhigh (52.1%) because max-effort runs more often launched a subagent code-review skill that produced out-of-scope edits, which FrontierCode penalizes.
The Terminal-Bench 4.0 public leaderboard lists GPT-5.6 Sol at 37.3%, and Anthropic notes OpenAI has published no such score for GPT-6 Sol. On the separate Terminal-Bench-Science 0.1 benchmark, Sonnet 5.5 scores 59.9% against 58.7% for Opus 5.5 and 24.7% for Claude Fable 5.
Beyond coding, Anthropic reports Humanity's Last Exam scores of 64.5% when tools are enabled and 56.9% without, 80.1% partial credit on the OSWorld 2.1 computer-use benchmark (43.5% strict pass rate, against 81.8% partial for Opus 5.5), and 1844 on GDPval-AA v2.1, level with Opus 5.5's 1846 and about 400 points above Sonnet 5. Artificial Analysis, an independent evaluator, scores the max-effort variant at 56 on its Intelligence Index with an output speed of about 142 tokens per second, and flags it as verbose: the index run generated 410M output tokens. That verbosity note matters for cost, because Anthropic's own claim of up to 30% lower cost per task versus Sonnet 5 depends on effort level and workload, not on the per-token price alone.
The context window holds 1,000,000 tokens and a single reply can run to 128,000, extendable to 300,000 on the Message Batches API with the output-300k-2026-03-24 beta header. It accepts text and images (including PDFs through the API) and returns text only. Adaptive thinking is on by default; the default effort is high through the Claude API and medium inside the Claude apps and Claude Code. Anthropic's docs list a June 2026 reliable knowledge cutoff, a 512-token minimum for prompt caching, and a 400 error for any non-default temperature, top_p, or top_k. Anthropic sells it directly and through AWS (Bedrock and Claude Platform on AWS), Google Cloud, and Microsoft Foundry, with zero data retention offered, and Anthropic commits to keeping it available no sooner than September 28, 2027.
Effort defaults matter in tools like Claude Code, where medium is the default. On CursorBench 4.0, a benchmark built from real Cursor sessions, it scores 55.5%, second only to Opus 5.5. Early testers quoted in the launch post include Lovable and Zendesk. Customer figures in the launch post, such as Balyasny Asset Management's 121k tokens per answer against 497k for Sonnet 5 on its own 2,441 finance tasks, are vendor-selected testimonials and were not independently reproduced.
On safety, the system card says Sonnet 5.5 is broadly less capable than Opus 5.5 and crosses no new Responsible Scaling Policy thresholds; its chemical and biological capability is estimated at or below Claude Opus 5. It is the first Sonnet model to launch with cyber safeguards, so higher-risk cybersecurity requests visibly fall back to Sonnet 5, and it adds classifiers that block attempts to extract its reasoning. Anthropic's automated behavioral audit (roughly 1,850 scenarios) shows it matching or improving on Sonnet 5 on most alignment measures, and close to Opus 5.5 on how rarely it tries to escape its sandbox, but also with thinking that is more illegible than many earlier models.
Sonnet 5.5 suits teams running high-volume, well-scoped coding, document, and support workloads who want most of Opus-tier output at half the per-token price of Claude Opus 5.5. It is the wrong pick for long architectural work that needs sustained judgment (use Opus 5.5 or Claude Fable 5.1), for audio or video inputs, and for pipelines built on forced tool calls or on switching thinking off. Anthropic lists five breaking changes from Sonnet 5, covered in the quirks below: thinking: disabled is replaced by between_tools, forced tool_choice returns a 400 error, thinking blocks are bound to the model and conversation, the computer_20251124 tool is rejected on the Claude API and Google Cloud, and Opus 4.8, Opus 4.7, and Sonnet 5 are no longer accepted as advisors.
Pricing
$2 per 1M input tokens and $10 per 1M output tokens, the same list price Anthropic states for [Claude Sonnet 5](/hub/models/claude-sonnet-5) and half of [Claude Opus 5.5](/hub/models/claude-opus-5.5)'s $4/$20. Cache reads cost $0.20 per 1M tokens, 5-minute cache writes $2.50, and 1-hour cache writes $4, from a 512-token minimum. The Batch API takes 50% off input and output. Anthropic lists no free API tier for this model.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.060 | $0.010 | $0.070 |
| Support reply | $0.0040 | $0.0030 | $0.0070 |
| One coding agent run | $0.400 | $0.200 | $0.600 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Adaptive thinking with five effort levels: Effort runs low, medium, high, xhigh, and max, with high as the Claude API default and medium in the Claude apps and Claude Code.
- 1M token context and 128K output: Both limits are standard rather than a premium tier, and Message Batches jobs can request 300,000 output tokens under a beta header.
- Computer use through a toolset: Computer use runs through the computer_toolset_20260801 toolset, offered on the Claude API plus Google Cloud, and Anthropic reports an OSWorld 2.1 partial score of 80.1%; Amazon Bedrock still accepts computer_20251124.
- Cyber safeguards with visible fallback: Higher-risk cybersecurity requests fall back to Claude Sonnet 5 with a transparent block, and routine bug finding and fixing is unaffected.
- Compaction and inline tools in beta: The compact-2026-09-04 header returns a signed summary block, and inline-tools-2026-09-15 lets a system message add or change a tool without breaking the prompt cache.
Pros
- Leads Claude Opus 5.5 on Terminal-Bench (70.6% versus 66.4%) and AutomationBench (44.7% versus 42.5%) at half the per-token price.
- Writes clearer prose than Sonnet 5, and Anthropic's testers single out slide templates that need minimal editing.
- Fewer tool calls: testers at Lovable report about a third fewer calls and roughly half the shell runs to finish a task.
Cons
- Ships five breaking API changes from Claude Sonnet 5, so an existing integration needs a migration pass before switching the model ID.
- Cyber requests above routine bug fixing fall back to Claude Sonnet 5, which can change output quality mid-workflow.
- Accepts text and images only, with no native audio or video input.
Benchmarks
- GDPval-AA v2: 1,844 vendor-reported · 29 Sep 2026 — Real knowledge-work deliverables judged against professionals, run by Artificial Analysis.
- SWE-bench Pro: 81.3% vendor-reported · 29 Sep 2026 — Harder, longer real-repository coding tasks, % solved.
- Humanity's Last Exam: 64.5% vendor-reported · 29 Sep 2026 — Expert-written questions across many fields, % correct.
- AA Intelligence Index: 56 cited: Artificial Analysis · 29 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- Output speed: 142 tok/s cited: Artificial Analysis · 29 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What are Claude Sonnet 5.5's pricing plans in 2026?
Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, half of Claude Opus 5.5's $4 and $20. Cache reads cost $0.20 per million, and cache writes cost $2.50 (5-minute) or $4 (1-hour). The Batch API takes 50% off, and because Anthropic says the model needs fewer tokens per task than Sonnet 5, a task can cost less even at identical rates.
Is Claude Sonnet 5.5 better than Claude Opus 5.5 for everyday work?
It depends on the task. Anthropic's tables show Sonnet 5.5 ahead on AutomationBench (44.7% versus 42.5%) and HealthBench Professional (69.2% versus 65.6%), and about two points behind on CursorBench (55.5% versus 57.8%). Anthropic still calls Opus 5.5 clearly stronger on complex, open-ended work that needs sustained judgment, so pick Sonnet for well-scoped jobs and Opus for the hard ones.
Can you download or self-host Claude Sonnet 5.5?
No. It is proprietary, with no published weights or architecture details, and it runs only as a hosted service through Anthropic's own API or its cloud partners. Anthropic's Commercial Terms of Service govern access.
Does Anthropic train on what you send to Claude Sonnet 5.5?
Not for commercial API use: Anthropic's Commercial Terms say it may not train models on Customer Content from its services, and zero data retention is available for this model. The system card does say the training mix can include user data from feedback or bug reports and data users explicitly permitted for training.
Who should pick Claude Sonnet 5.5, and who should skip it?
Pick it for high-volume coding, support, and document workflows where speed and cost per task matter, especially at medium effort. Skip it for audio or video inputs, for pipelines that force a tool call, and for long architectural work, where Claude Opus 5.5 or Claude Fable 5.1 is the better fit. Teams on Sonnet 5 should also budget time for five breaking API changes.