All AI guides
Analysis12 min read

Claude Haiku 5.5 vs Haiku 4.5: Should You Switch Now?

Claude Haiku 5.5 is Anthropic's small model released on 7 October 2026. It has a 1M-token context window and prices by prompt length: $0.10 input and $0.50 output below 100,000 tokens, then $0.50 and $2.50. Anthropic lists Haiku 4.5 as active, with retirement not sooner than 15 October 2026.

The short version

Haiku 5.5 charges $0.10 input and $0.50 output per million tokens for prompts up to 100,000 tokens, against $1 and $5 for Haiku 4.5. A tokenizer change adds about 30% more tokens, so short prompts save roughly 87% and long ones about 35%. Haiku 4.5 is not scheduled to retire.

Anthropic released Claude Haiku 5.5 on 7 October 2026 at $0.10 per million input tokens, a tenth of what Claude Haiku 4.5 costs.

The catch sits in the pricing footnote. Haiku 5.5 is the only current Claude model priced by prompt length: past 100,000 tokens the rates rise fivefold, and a new tokenizer turns the same text into roughly 30% more tokens. If you run classification, routing or extraction at volume, the saving is real and large. If your prompts are long, it is smaller than the headline, and a prompt that sat at 80,000 tokens on the old model can land on the expensive side of the line.

This guide works through the exact numbers, three cost examples, the API calls that now return errors, what Anthropic's deprecation page does and does not say about the older model's retirement, and who should wait. Every figure was read from Anthropic's own pages on 8 October 2026.

What changed between Haiku 4.5 and Haiku 5.5

Anthropic's announcement calls it "the cheapest, fastest, and most capable small model we've ever released". The numbers behind that sentence come from the Haiku 5.5 overview and the pricing page, and they are best read side by side.

SpecHaiku 4.5Haiku 5.5
Context window200K tokens1M tokens
Max output64K tokens128K tokens
Input, per million$1$0.10, then $0.50
Output, per million$5$0.50, then $2.50
Cache read, per million$0.10$0.01, then $0.05
Batch input / output$0.50 / $2.50$0.05 / $0.25
ThinkingManual budgetAdaptive, with effort

The "then" in each price cell marks the 100,000-token line. Prompts up to that size pay the first number; longer prompts pay the second, on input and output alike. The batch figures in the last price row are the lower tier, and they already include the 50% batch discount Anthropic applies to both models.

Two changes matter beyond price. Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, and its default is medium. It also adds a new computer-use toolset and a browser use tool, and the older model does not support the browser tool.

Anthropic's what's new page lists the 1M-token context window and the 128K output limit as "up from 200k and 64k", and the overview gives a reliable knowledge cutoff of June 2026. The announcement says it pairs well with Claude Opus 5.5 and Claude Sonnet 5.5 as a subagent on coding work.

What it costs: three workloads, worked through

List prices hide the tokenizer. Anthropic's migration guide says the same text produces "approximately 30% more tokens" on the new model, so every prompt and every response grows before a price is applied. The table below applies both effects. Each row assumes 30% more tokens, uses the list rates from the pricing page, and leaves out thinking tokens.

WorkloadHaiku 4.5Haiku 5.5Saving
1M classification calls, 2,000 in / 200 out$3,000$39087%
One 80K-token prompt, 1,000 out$0.085$0.05535%
One 120K-token prompt, 2,000 out$0.130$0.08535%

The first row is the case the model was built for. A million calls at 2,000 input and 200 output tokens each cost $3,000 on the old model. On the new one, the same calls become 2,600 input and 260 output tokens, and the bill falls to $390.

The second and third rows are the long-prompt cases, and they save a third. The reason is arithmetic. Above the line, the new rates are half the old ones ($0.50 against $1 on input), but 30% more tokens eat part of that, leaving about 65 cents of cost for every dollar spent before.

Anthropic puts the blended result at "around 75% less to run". Its footnote says 90% of requests to the old model came in under 100,000 tokens, and that the figure also accounts for the tokenizer. Your mix will differ, so recompute from your own logs.

The 100K line: why 77,000 old-model tokens is the number to remember

Divide 100,000 by 1.3 and you get about 77,000. Any prompt above that size on the old model will count as more than 100,000 tokens on the new one, and it will be billed at the higher tier. The pricing page states the rule plainly: "a prompt of over 100,000 tokens pays higher prices."

The cliff is steep. Take the 80K prompt from the second row. It becomes about 104,000 tokens and is billed at the higher tier, for about $0.055 in total. Had it stayed under the line, the same call would have cost about $0.011. Crossing by four thousand tokens makes the call roughly five times dearer.

That has a practical consequence for retrieval pipelines and agents that stuff context until it is nearly full. A cap at 70,000 old-model tokens, or about 90,000 new ones, keeps a request on the cheap tier. Weigh the quality lost by trimming against the price of crossing, and use caching, which helps too: on both tiers, a cache read costs a tenth of the matching input rate.

Other current models do not share this split. The pricing page says Claude 4.6 and later models, "except Claude Haiku 5.5", include the full 1M window at standard rates, which is why Claude Sonnet 5.5 at $2 and $10 stays flat however long the prompt runs. Claude Sonnet 4.6 and Claude Fable 5 follow the same flat rule.

Who it affects: the winners, and who it hurts

The winners are easy to name. Anyone running short, repetitive calls at volume on Haiku 5.5 gains the most: summaries, compaction, database queries and classification, the jobs Anthropic lists. Teams that run a small model as a subagent under a larger one gain too, because the subagent's share of the bill shrinks sharply.

Computer-use agents are the other clear winners. In Anthropic's own table, Haiku 5.5 scores 72.4% on the OSWorld 2.1 offline subset against 15.7% for Haiku 4.5, and 83.9% for Sonnet 5.5. A model that was close to useless at operating a desktop is now within 12 points of the larger one, at a twentieth of the input price.

The losers are less obvious, and the documentation is candid about them:

  • Long-prompt workloads. They save a third, not nine tenths, and now face a price cliff.
  • Priority Tier customers. The migration guide says "Priority Tier is not supported on Claude Haiku 5.5".
  • Security teams. The announcement says the cyber safeguards "still block penetration testing", and they are stricter than the old model's.
  • Code with 0, prefills or first-block reads. These fail or misbehave, as the next section shows.

If you cannot tell which group your workload falls into, Smart Match can narrow the choice from your constraints, and the model recommender points to a model when one constraint dominates.

What breaks when you change the model ID

Swapping claude-haiku-4-5 for claude-haiku-5-5 is not a one-line change for most codebases. The migration guide lists ten items for every starting model, and Anthropic marks five changes as breaking. A request that sends any of these four returns a 400 error:

  • A manual thinking budget, {"type": "enabled", "budget_tokens": N}.
  • A temperature other than 1, a top_p other than 0.99, or any top_k.
  • A final assistant message used as a prefill.
  • The old computer_20250124 tool on the Claude API or Google Cloud.

The thinking change is the one most teams will hit first. Here is the guide's own before and after:

{
  "model": "claude-haiku-5-5",
  "max_tokens": 16000,
  "thinking": { "type": "adaptive" },
  "output_config": { "effort": "medium" },
  "messages": [{ "role": "user", "content": "..." }]
}

Adaptive thinking is on by default, so a response can open with a thinking block even when you never asked for one. Code that reads content[0] as the answer will break; select blocks by type instead. Thinking tokens also count toward max_tokens, so a small limit can stop the response after the thinking and before any text appears.

Two further items catch teams late. The model runs safety classifiers that can decline a request, and the migration guide says there is no server-side fallback, so your client must handle stop_reason: "refusal". And thinking blocks from the new model work only in the account that produced them, which matters if you replay stored conversations through another account.

Is Haiku 4.5 being retired?

Searchers ask this directly, so here is the exact answer. Anthropic's model deprecations page, read on 8 October 2026, lists claude-haiku-4-5-20251001 as Active, with deprecation marked N/A and a tentative retirement of "Not sooner than October 15, 2026".

That is a floor, not a date. No deprecation has been announced, and the page does not say 15 October is when the model disappears. Be careful with pages that treat it as a deadline.

The same page sets a rule that bounds the timeline: Anthropic provides "at least 60 days' notice before model retirement for publicly released models". A notice posted today would put the earliest retirement in early December. The pattern is visible in the table. Claude Sonnet 4.5 was deprecated on 30 September 2026 and is scheduled to retire on 30 November, 61 days later.

So the honest position is this. You have no deadline today, and you will get at least two months of warning when one arrives. The reason to move is price and capability, not fear of a cutoff. For comparison, Haiku 5.5's own earliest retirement is 7 October 2027.

The case that the saving is smaller than it looks

The best objection to this whole analysis is that price per token is not price per task. Artificial Analysis scores the max-effort reasoning variant at 43 on its Intelligence Index and measures 241.9 output tokens per second, which it calls notably fast. It also calls the model "very verbose": it generated 440M tokens to complete the index, against a median of 100M for comparable models.

Verbosity cuts straight into the saving. Output tokens cost five times what input tokens do, and a model that writes more than four times the median spends the difference there. Artificial Analysis puts the cost of one Index task at $0.21 on average. We could not find a like-for-like figure for the older model on that page, so the comparison to Haiku 4.5 stays unmeasured.

The answer is partial. Those numbers describe the max-effort variant, and the default is medium, which thinks less. The migration guide suggests choosing a lower effort level where the old model ran with thinking off, and says the new model can skip thinking on simple requests. Anthropic's blended 75% figure also already counts the extra tokens used per task.

We found no published cost-per-task comparison against Haiku 4.5 at default settings. Run your own: send a few hundred real requests through both models and compare the invoice, not the rate card. If the output tokens per request rise by more than the savings on input, your saving shrinks, and it can vanish for generation-heavy jobs.

Haiku 5.5 against Sonnet 5.5 and the other cheap models

Cheaper does not mean interchangeable. Anthropic's pricing page positions Claude Fable 5.1, at $10 and $50, for demanding reasoning and long-horizon agentic work. Anthropic's benchmark table puts the new model between its old self and the mid-size model, and the gap to Sonnet is widest on agentic coding.

BenchmarkHaiku 4.5Haiku 5.5Sonnet 5.5
GDPval-AA v2.1 (Elo)73516201840
OSWorld 2.1, offline subset15.7%72.4%83.9%
Humanity's Last Exam, no tools10.2%45.9%56.9%
Terminal-Bench 4.00.0%39.2%70.6%

These scores come from Anthropic's launch chart, so treat them as vendor-reported. The jump over the old model is enormous on every row, but the Terminal-Bench line shows the ceiling. Anthropic says Sonnet 5.5 and Opus 5.5 "remain better choices for complex agentic coding", and describes the small model as best for "narrowly scoped tasks" such as compaction, summarization and subagent work.

On price, the nearest rival is GPT-6 Luna, which lists the same $0.10 input and $0.50 output rates on its HokAI page. Anthropic's chart gives Luna 1437 on GDPval-AA, 48.9% on OSWorld and 16.4% on Terminal-Bench. Those rows favour Haiku 5.5, but they come from the Haiku maker's own chart, so confirm them on your tasks.

Other small models cost more, according to their HokAI pages. Ling 3.1 Flash from InclusionAI lists $0.30 per million input tokens and is still a preview with a free launch trial, and Gemini 3.8 Flash charges $0.75 through the end of 2026. The models leaderboard lets you sort these by price, speed and context, and our Claude model picker covers the rest of the Anthropic lineup. For a wider choice, the LLM selection guide walks through the criteria.

A migration plan for this week

Start with measurement, because the order below decides how much you save:

  1. Count a sample of production prompts with the token-counting endpoint set to claude-haiku-5-5, not with counts taken on the old model.
  2. Split requests into three buckets: under 77,000 old-model tokens, between 77,000 and 100,000, and above that. The middle bucket is where the cliff bites.
  3. Remove temperature, top_p, top_k, prefills and manual thinking budgets, then select content blocks by type.

Then change the code and test the result:

  1. Set effort lower than medium for simple classification and routing, and leave it at the default for anything that reasons.
  2. Replay a sample of production traffic through both models and compare cost per task, accuracy and refusal rate.
  3. Move the cheap bucket first. Keep the old model as a fallback for the rest until the replay shows parity.

Haiku 5.5 is generally available on every major cloud, but model IDs differ. On Amazon Bedrock the new ID is anthropic.claude-haiku-5-5, and teams on Claude in Amazon Bedrock need to change the ID there separately. Teams that call Anthropic directly can compare plans on Anthropic's API service page, and the company profile covers the maker. The Anthropic models directory shows every Claude model side by side.

What to watch next

Two dates matter. If Anthropic posts a deprecation notice for Haiku 4.5 before 15 October, the 60-day rule makes early December the earliest retirement, and the question changes from "should we move" to "by when". If nothing appears, the older model simply stays active, and the case for moving rests on the invoice.

The second thing to watch is whether the pricing split spreads. Haiku 5.5 is the only current model charged by prompt length. If larger models follow, the 77,000-token line becomes a design constraint for every long-context pipeline, and it will be worth capping context before the next launch rather than after the bill.

Frequently asked questions

Is Claude Haiku 4.5 being retired on 15 October 2026?

Not according to Anthropic's deprecations page, read on 8 October 2026. It lists claude-haiku-4-5-20251001 as Active with deprecation N/A and retirement "Not sooner than October 15, 2026", which is a floor and not a date. Anthropic promises at least 60 days' notice for publicly released models.

How much cheaper is Claude Haiku 5.5 than Haiku 4.5?

List prices fall 90% for prompts up to 100,000 tokens ($0.10 against $1 input) and 50% above that. Because the new tokenizer counts about 30% more tokens for the same text, our worked examples save about 87% on short calls and 35% on long prompts. Anthropic puts the blended figure at around 75%.

What is the 100,000-token price line on Haiku 5.5?

Haiku 5.5 is the only current Claude model priced by prompt length. Prompts over 100,000 tokens pay $0.50 input and $2.50 output per million instead of $0.10 and $0.50. A prompt of about 77,000 tokens on Haiku 4.5 grows past the line on the new tokenizer.

What breaks when I switch from Haiku 4.5 to Haiku 5.5?

Manual thinking budgets, non-default temperature, top_p or top_k values, assistant prefill and the old computer_20250124 tool all return 400 errors. Responses can begin with thinking blocks, so read content by type. Priority Tier is not supported, and safety classifiers can return a refusal stop reason.

Should I use Haiku 5.5 or Sonnet 5.5?

Use Haiku 5.5 for narrow, high-volume work such as classification, summaries and subagents. Anthropic says Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding, where Sonnet scores 70.6% on Terminal-Bench 4.0 against 39.2% for Haiku 5.5.

Covered in this guide

  • Claude Haiku 5.5: Anthropic's smallest and fastest 5.5 model, released October 7, 2026, with a 1M context window, built for high-volume work.
  • Claude Sonnet 5.5: Anthropic's fastest Sonnet-class model, launched in September 2026, with a 1M context and output more than 30% faster than Sonnet 5.
  • Anthropic's API service page: Anthropic provides Claude AI via API from $1/MTok input, with Free, Pro ($20/mo), Team, and Enterprise plans covering developers to large organizations.
  • company profile: Anthropic, founded 2021 by 7 ex-OpenAI researchers, builds Claude and was valued near $965B after its May 2026 Series H round.
  • Claude Fable 5: The first generally available Mythos-class Claude model, with 1M context and 95.0% SWE-bench Verified. Released June 2026 by Anthropic.
  • Claude Fable 5.1: Anthropic's second Mythos-class model, released Sept 1, 2026, for long-horizon agentic coding and research at Fable 5's per-token rates.
  • Claude in Amazon Bedrock: Anthropic's Claude models served by AWS in 27 regions, billed per token on your AWS invoice and secured with IAM, KMS keys and CloudTrail logs.
  • Claude Opus 5.5: Anthropic's September 2026 flagship LLM with a 1 million token context window, built for long-running agentic coding and computer use.
  • Claude Sonnet 4.6: Claude Sonnet 4.6 by Anthropic (Feb 2026) scores 79.6% on SWE-bench Verified with a 1M-token context window at $3/$15 per 1M tokens.
  • Gemini 3.8 Flash: Google DeepMind shipped Gemini 3.8 Flash on September 2, 2026, a multimodal Gemini 3 model tuned for long-horizon coding and agentic enterprise work.
  • GPT-6 Luna: GPT-6 Luna (Sept 2026) is OpenAI's lowest-cost GPT-6 model, tuned for high-volume chat, extraction and classification with six adjustable reasoning-effort levels.
  • InclusionAI: Ant Group's open AGI lab behind the free, MIT-licensed Ling, Ring and Ming model families, including the 1T-parameter Ling-1T.
  • Ling 3.1 Flash: InclusionAI's text-only reasoning model for coding agents and long tasks, released on 30 September 2026 as the successor to InclusionAI's July flash model.

Sources

Still deciding?

This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.

Start Smart Match

Related guides

All AI guidesBrowse the AI directory