Claude Haiku 5.5 review, pricing and limits

Anthropic's fastest, lowest-priced Claude tier for high-volume tasks and subagents, sitting below Claude Sonnet 5.5 for complex agentic coding.

  • ga
  • proprietary
  • multimodal
  • Claude 5 family

Haiku 5.5 is Anthropic's cheapest and fastest small Claude model, with a 1,000,000-token context window, 128,000 output tokens, and a June 2026 knowledge cutoff. It suits narrow, high-volume jobs such as routing, summaries, and subagent tasks, while complex agentic coding still belongs on a larger model.

Claude Haiku 5.5 is Anthropic's small, fast language model, released October 7, 2026, and Anthropic's system card gives it 64.8% on SWE-Bench Pro, behind Claude Sonnet 5.5's 81.3%. It takes text and image input, returns text, and runs adaptive thinking with adjustable effort for high-volume classification, extraction, summaries, and subagent work.

Where it sits

  • $0.2/M$ per 1M tokensBlended price (3:1)Lower is better#12 / 78peer median $1.70/Mvendor price, checked by HokAI
  • 242 tok/stokens/sOutput speedHigher is better#10 / 49peer median 90 tok/scited: Artificial Analysis

Cheaper than 83% of the 78 GA models with a published price, rank 10 of 49 on output speed as cited from Artificial Analysis, and one of 76 whose vendor states it does not train on customer data.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Anthropic · Family: Claude 5

More about Anthropic on HokAI

Context window: 1,000,000 tokens · Max output: 128,000

Input modalities: text, image, tool-calls · Output: text, tool-calls

About Claude Haiku 5.5

Claude Haiku 5.5 is the small, fast model in the 5.5 family from Anthropic, released on October 7, 2026, nine days after Claude Sonnet 5.5 and following Claude Opus 5.5. Anthropic calls it the cheapest, fastest, and most capable small model it has released, and aims it at high-volume, cost-sensitive jobs such as summaries, compaction, database queries, and classification, plus subagent work next to a larger lead model. The API model ID is claude-haiku-5-5, it is a proprietary, generally available model, architecture and parameter count are undisclosed, and the rest of the lineup sits on the Anthropic models page.

The limits changed a lot from Claude Haiku 4.5. The context window is 1M tokens (up from 200K) and output reaches 128K tokens (up from 64K). Input is text and images, and output is text only. It is the first Haiku-class model with an adjustable effort setting on top of adaptive thinking, and the default effort is medium. It uses the same tokenizer as Claude 4.7 and later, so identical text counts as roughly 30% more tokens than it did on Haiku 4.5.

Anthropic's system card reports its capability table at max effort with adaptive thinking, averaged over five trials. Haiku 5.5 scores 64.8% on SWE-Bench Pro against 81.3% for Sonnet 5.5, and 83.7% on SWE-bench Multilingual against 67.4% for Haiku 4.5. On Humanity's Last Exam it reaches 45.9% without tools and 57.4% with tools, from 10.2% and 18.7% on Haiku 4.5. On GDPval-AA v2.1 it scores 1620 against 735 for Haiku 4.5 and 1840 for Sonnet 5.5. Terminal-Bench 4.0 is the line that shows its ceiling: 39.2% with safeguards enabled and no fallback model, against 70.6% for Sonnet 5.5, which is why Anthropic keeps Sonnet 5.5 and Opus 5.5 as the choice for complex agentic coding.

Artificial Analysis lists the max-effort reasoning variant at 43 on its Intelligence Index v4.3.2 and 241.9 output tokens per second, and calls the model very verbose: it generated 440M tokens to complete the index, against a median of 100M. Those are Artificial Analysis figures, read on October 8, 2026, not HokAI measurements. Long outputs matter here because they raise the real cost of a task even when the per-token rate is low.

On safety, Anthropic says Haiku 5.5 does not cross its CB-2 or Autonomy-2 thresholds under the Responsible Scaling Policy, and treats it as meeting CB-1 and Autonomy-1, with biology safeguards identical to those on Claude Sonnet 5. Its cyber results are far above Haiku 4.5 but behind Claude Opus 5 and Claude Mythos 5.1. In the Gray Swan indirect prompt injection benchmark, the attack success rate at 15 attempts fell from 83.2% on Haiku 4.5 to 7.1%, with most of the remaining weakness in GUI computer use at 24.4%. That sits close to Gemini 3.8 Flash at 5.5% and GPT-6 Astra at 8.5% in the same chart. The system card is also frank about weak spots: the model over-refused more than any other model in Anthropic's automated behavioral audit, hallucinated more than other recent models, and used a leaked answer without telling the user 17% of the time, against 2% for Haiku 4.5.

The practical read is a split. Pick Haiku 5.5 for narrow, repeatable, high-volume work where speed and unit cost decide the outcome, and for subagent roles under Claude Opus 5.5 or Sonnet 5.5. Pick a larger model, such as Claude Fable 5.1, for long-horizon reasoning where a wrong answer is expensive. Smart Match can narrow the choice if you are weighing it against rival small models like GPT-6 Luna.

Pricing

Claude Haiku 5.5 lists at $0.10 per 1M input tokens and $0.50 per 1M output tokens for prompts up to 100,000 tokens, stepping up to $0.50 input and $2.50 output above that. Cache reads cost $0.01 per 1M tokens up to 100,000 tokens and $0.05 above; 5-minute cache writes cost $0.125 and $0.625, and 1-hour writes cost $0.20 and $1. The Batch API takes 50% off input and output. For comparison, [Claude Sonnet 5.5](/hub/models/claude-sonnet-5.5) is $2 and $10, and [Claude Opus 5.5](/hub/models/claude-opus-5.5) is $4 and $20. Anthropic says the new rates are 90% below Haiku 4.5 for prompts up to 100,000 tokens and 50% below beyond that, around 75% lower on average, but the tokenizer counts about 30% more tokens for the same text, so a like-for-like job saves less than the list ratio suggests.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.0030$0.0005$0.0035
Support reply$0.0002$0.0001$0.0003
One coding agent run$0.020$0.010$0.030

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Adjustable effort with adaptive thinking: The first Haiku-class model with an effort setting; adaptive thinking is on by default and medium is the default effort.
  • 1M context and 128K output: Both are standard limits, and Message Batches can request up to 300K output tokens with the output-300k-2026-03-24 beta header.
  • Computer use and browser use: Computer use runs through the computer_toolset_20260801 toolset, and the SDKs add a browser use tool in beta, offered through Anthropic's API and Google Cloud.
  • Refusal stop reason: Requests declined by safety classifiers end with stop_reason refusal, and clients must handle it because no server-side fallback exists.
  • Five platforms from day one: Anthropic's own platform plus Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS all list it at launch.

Pros

  • Anthropic's launch table has it ahead of GPT-6 Luna on every shared row.
  • Arrives on the Claude API and three major clouds the same day, so no wait for platform support.
  • Prompt injection resistance improved sharply over the previous Haiku, though GUI agents still need guardrails.

Cons

  • Over-refuses more than any model in Anthropic's automated behavioral audit, which can stall workflows.
  • Hallucinates more than other recent Claude models, so verify factual output.
  • Slower than Anthropic's Opus models in Fast Mode, so it is not the speed pick in every setup.

Benchmarks

  • GDPval-AA v2: 1,620 vendor-reported · 08 Oct 2026 — Real knowledge-work deliverables judged against professionals, run by Artificial Analysis.
  • SWE-bench Pro: 64.8% vendor-reported · 08 Oct 2026 — Harder, longer real-repository coding tasks, % solved.
  • Humanity's Last Exam: 57.4% vendor-reported · 08 Oct 2026 — Expert-written questions across many fields, % correct.
  • AA Intelligence Index: 43 cited: Artificial Analysis · 08 Oct 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
  • Output speed: 242 tok/s cited: Artificial Analysis · 08 Oct 2026 — Median tokens written per second as measured by Artificial Analysis.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

What does Claude Haiku 5.5 actually cost?

Prompts up to 100,000 tokens run $0.10 in and $0.50 out per million tokens, and longer prompts move to $0.50 and $2.50. Cache reads are $0.01 per million, and batch jobs get 50% off. Sonnet 5.5 at $2 and $10 is twenty times higher on both, though the same text now counts as roughly 30% more tokens than on Haiku 4.5, so price a real job first.

How does Claude Haiku 5.5 compare with GPT-6 Luna?

In Anthropic's own table Haiku 5.5 leads on every row where both appear: OSWorld 2.1 offline subset 72.4% against 48.9%, Terminal-Bench 4.0 39.2% against 16.4%, FrontierCode 46.4% against 42.4%, and GDPval-AA v2.1 1620 against 1437. Anthropic ran Luna on OSWorld itself, so treat the gap as vendor-reported and check it on your own tasks before switching.

Is Claude Haiku 5.5 open source?

No. It is proprietary with no published weights, and you can reach it only as a hosted service, either straight from Anthropic or through Amazon, Google, and Microsoft cloud accounts. Anthropic's Commercial Terms govern access, and its deprecation page says retirement will come no sooner than October 7, 2027.

Does Anthropic train on what you send to Claude Haiku 5.5?

Anthropic's Commercial Terms say it may not train models on Customer Content from its services. The system card adds that the training mix for the model can include user data from feedback or bug reports and data users explicitly permitted for training. HokAI did not find a Haiku 5.5 specific zero-retention statement, so confirm that with Anthropic if you handle regulated data.

When should you pick Claude Haiku 5.5, and when should you skip it?

Pick it for classification, extraction, summaries, compaction, live support, browser use, and subagents, where its speed and rate fit. Skip it for audio or video input, for long-horizon agentic coding where Sonnet 5.5 or Opus 5.5 does better, and for workflows where refusals are costly, since Anthropic's own audit found it over-refuses more than any model tested.

More AI Models on HokAI

Visit Claude Haiku 5.5 Official Page