The Cheapest LLM APIs in 2026, and What Each Cheap Price Leaves Out
On 9 October 2026 the cheapest LLM API per token is Claude Haiku 5.5 for prompts up to 100,000 tokens. GLM-5.3 Flash costs $0.15 in and $0.50 out, and DeepSeek V4.1 Flash costs $0.15 in and $0.60 out off-peak. The six budget models score between 38 and 43 on the Artificial Analysis index.
The short version
Claude Haiku 5.5 is the cheapest LLM API for prompts under 100,000 tokens at $0.10 in and $0.50 out per million, with GLM-5.3 Flash and DeepSeek V4.1 Flash close behind. Each price carries a condition: a length tier, a peak window, a trial or a sale. Test your own prompts before committing.
The cheapest LLM API on 9 October 2026 is not the one with the lowest number on its pricing page. Claude Haiku 5.5 lists at $0.10 per million input tokens, and GLM-5.3 Flash follows at $0.15. Each price has a condition attached. That condition decides what you actually pay, and on a busy workload it can move the bill by a factor of five, which is the gap between a rounding error and a budget line your finance team will ask about.
This guide is for a developer or a small team that sends a few million tokens a day through an API and wants the lowest bill that still gives a usable answer. It compares six budget models from six vendors on price, an independent quality score and the fine print. Prices come from each vendor's pricing page, opened on 9 October 2026 unless a line says otherwise. Quality is the Artificial Analysis Intelligence Index, as recorded on the HokAI model pages.
Two things set this apart from a price list. First, it puts a quality number next to every price, because a model that is 20% cheaper and 10% worse is rarely a saving. Second, it names what each price leaves out: a long-prompt surcharge, a preview label, a peak-hour multiplier, a sale that could end.
The shortlist: which cheap LLM APIs are worth a look in 2026?
Six models cover the low end. Four are generally available. Two, Ling 3.1 Flash and Mistral Large 4, still carry a preview label, which matters for production use.
| Model | Input / output per 1M tokens | AA index | Status |
|---|---|---|---|
| Claude Haiku 5.5 | $0.10 / $0.50 (prompts to 100K) | 43 | Generally available |
| GLM-5.3 Flash | $0.15 / $0.50 | 42 | Generally available |
| Ling 3.1 Flash | $0.30 / $0.90 | 41.1 | Preview |
| Gemini 3.8 Flash | $0.75 / $3.75 | 41 | Generally available |
| DeepSeek V4.1 Flash | $0.15 / $0.60 (off-peak) | 40 | Generally available |
| Mistral Large 4 | $0.68 / $2.09 (sale price) | 38 | Preview |
The scores sit in a tight band, 38 to 43. Tight. That is the first useful finding. Across these six, price varies by a factor of seven on both input and output, while quality varies by about 13%. If your task is one all six handle, such as classification, extraction or short summaries, the quality gap is small. Price should lead.
These are the rates for each vendor's own API. Resellers such as OpenRouter list their own rates, which can run higher or lower, and some run free trials that end with little warning.
How to choose a cheap LLM API before you read the fine print
Start with the job, then the prompt length, then the data rules, and only after those three questions have answers should you open a pricing page and compare numbers. Price is the tiebreaker, not the starting point.
- Short prompts, predictable load. Haiku 5.5 and GLM-5.3 Flash lead on cost and score.
- Batch work you can run at night. DeepSeek off-peak, or Haiku with the Batch API.
- Prompts over 100K tokens. Haiku loses its edge, so compare long-prompt rates.
Three more cases need a different pick.
- Image, audio or video input. Gemini 3.8 Flash is the only one of the six that lists audio input.
- Open weights you can self-host. GLM-5.3 Flash and DeepSeek V4.1 Flash ship under MIT licences.
- Not sure yet. Run your own prompts through two or three models.
The model leaderboard sorts models by price and score. The model recommender filters by a constraint such as budget or open weights. If the choice depends on your stack and the data you handle, Smart Match asks a few questions and returns one to three picks from the whole directory.
What does each cheap price leave out?
A list price is a headline. This table shows the condition behind each one, and the sections after it give the detail.
| Model | The condition behind the price |
|---|---|
| Claude Haiku 5.5 | Prompts over 100K tokens cost five times more |
| GLM-5.3 Flash | None found; a 50% launch discount ended 9 September |
| Ling 3.1 Flash | Trial is free for about two weeks; hosted only |
| Gemini 3.8 Flash | Price doubles on 1 January 2027 |
| DeepSeek V4.1 Flash | Every rate doubles in weekday peak windows |
| Mistral Large 4 | 50% sale with no stated end date; preview |
Claude Haiku 5.5: cheap until the prompt gets long
Anthropic prices Haiku 5.5 by prompt length. A request with 100,000 input tokens or fewer pays the low rate. A request above that pays the high rate.
- Up to 100K tokens: $0.10 in, $0.50 out, per million.
- Over 100K tokens: $0.50 in, $2.50 out, per million.
Anthropic's page says the length counts every input token, including cache reads and cache writes, and that each request is priced on its own. So Haiku 5.5 is the cheapest model here for short and medium prompts. For long ones it is a mid-priced model. Retrieval pipelines that put 150K tokens of context into every call pay the higher rate on every call.
- Cache reads cost $0.01 per million up to the threshold.
- The Batch API takes 50% off both input and output.
- The new tokenizer counts about 30% more tokens for the same text, so a like-for-like saving is smaller than the list ratio suggests.
If you are moving from the older model, the Haiku 5.5 versus Haiku 4.5 comparison walks through the switch.
GLM-5.3 Flash: the cleanest price on the page
Z.ai lists GLM-5.3 Flash at $0.15 per million input tokens and $0.50 for output, with cached input priced far lower. We found no length tiers, peak windows or sale labels on the pricing page. A 50% launch promotion ended on 9 September, so today's number is the standing rate.
The full GLM-5.3 costs $1.40 in and $4.40 out. The Flash version is roughly a tenth of that for a score of 42. The weights are open under an MIT licence, so you can also self-host and pay for compute only.
The catch is the data side. The model card does not set out retention or training terms for the hosted API, and the HokAI record lists China and global resellers as regions.
Ling 3.1 Flash: free now, priced later
Ling 3.1 Flash, from InclusionAI, launched on 30 September. Artificial Analysis records $0.30 in and $0.90 out. On 8 October, three hosts showed a zero price: Novita AI, OpenRouter and Vercel AI Gateway. InclusionAI described the free period as two weeks.
So the honest answer to "what does Ling cost" is "nothing, briefly". After the trial it sits mid-table on price and at 41.1 on quality. That is level with Gemini 3.8 Flash and below GLM and Haiku, which cost less.
Two more notes. The model is verbose, and output tokens are the expensive side, so real bills run above a like-for-like estimate. And the weights are not released. InclusionAI says it plans an open-source release after the trial, with no licence announced.
Gemini 3.8 Flash: the price with a date on it
Google's pricing page lists the paid tier at $0.75 in and $3.75 out per million tokens, thinking tokens included, through 31 December 2026.
- From 1 January 2027 the page says $1.50 in and $7.50 out.
- Cached input is $0.075 now and $0.15 next year, plus an hourly storage fee.
- Batch and flex modes cut the standard rates by half.
- The context window is about a million tokens.
That makes Gemini the most expensive model in this group today. From January it costs 15 times what Haiku 5.5 does on input. What you buy is breadth: it takes text, images, video, audio and PDFs. A free tier exists.
It runs through Google AI Studio, and its data terms differ from the paid tier's, as the data section below explains. If your work is multimodal, the premium may be worth it. If it is text only, it is hard to justify.
DeepSeek V4.1 Flash: cheap, but check the clock
DeepSeek prices V4.1 Flash at $0.15 per million input tokens on a cache miss and $0.60 per million output tokens, in off-peak hours. A cache hit costs a small fraction of that.
Peak hours run 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Every rate doubles then, to $0.30 in and $1.20 out. Weekends and Chinese public holidays are off-peak in full.
For a batch job you can schedule, this is among the cheapest options anywhere. For an interactive product with users in Europe or the US, much of the traffic lands in or near the peak window. The older name deepseek-v4-flash still works but is served by V4.1 Flash at its price, so stored model names keep working.
The weights are open under MIT. The larger DeepSeek V4 Pro is not a budget pick: it costs $0.66 in and $1.98 out off-peak.
Mistral Large 4: a sale on a preview
Mistral bills Mistral Large 4 at $0.68 in and $2.09 out. Its pricing page shows these against crossed-out list prices that are exactly double. That is a 50% sale, and the page states no end date.
The model scores 38, the lowest in this group, and it is a preview. Weights are promised for the end of October, with the licence not yet announced. Our guide to what Mistral Large 4 is for covers the model itself.
On price alone it does not belong in a cheapest-API list. Its predecessor, Mistral Large 3, lists at $0.50 in and $1.50 out. Large 4 earns a place only if you need its image input or its coming open weights, and you accept that the sale could end.
What does a month of real usage cost?
List prices are per million tokens, which hides the shape of a bill. Take one fixed workload, because comparing six vendors on six different mixes of input, output, caching and prompt length is how a team ends up arguing about numbers that were never comparable in the first place: 50 million input tokens and 10 million output tokens a month, every prompt under 100,000 tokens, no caching. That is roughly 2,000 support-ticket summaries a day for a small product.
| Model | Monthly cost | Note |
|---|---|---|
| Claude Haiku 5.5 | $10.00 | About $13 with the larger tokenizer |
| GLM-5.3 Flash | $12.50 | Standing rate |
| DeepSeek V4.1 Flash | $13.50 | $27.00 if all traffic is peak |
| Ling 3.1 Flash | $24.00 | $0 during the trial |
| Mistral Large 4 | $54.90 | $109.80 at list price |
| Gemini 3.8 Flash | $75.00 | $150.00 from 1 January 2027 |
The sums are plain: input tokens times the input rate, plus output tokens times the output rate. For Haiku 5.5 that is 50 times the input rate plus 10 times the output rate. The total is $10.
The top three land within $4 of each other. A gap that small can be reordered by one detail, such as a longer system prompt, a chattier model or a workload that skews toward output rather than input. Tokenizer differences, output verbosity and cache hit rate matter more than the third decimal on a price page.
If you send the same long system prompt on every call, caching changes the picture. A DeepSeek cache hit costs $0.003 per million. A Haiku cache read costs $0.01.
How fast are they?
Price is one cost. Waiting is another. Both count. If a user sits watching a chat reply stream, speed is part of the product.
| Model | Output speed (tokens per second) | Note |
|---|---|---|
| Claude Haiku 5.5 | 241.9 | Max-effort reasoning variant |
| Gemini 3.8 Flash | 237 | Artificial Analysis, High setting |
| DeepSeek V4.1 Flash | 218.7 | Off-peak figure not separated |
| Ling 3.1 Flash | 211.4 | Preview, may change |
| Mistral Large 4 | 116.1 | Preview |
| GLM-5.3 Flash | 52.5 | Cheapest tier, slowest output |
These are Artificial Analysis measurements as recorded on the HokAI model pages. Treat them as a guide, not a promise. Speed varies with load, region and reasoning settings.
The spread is large. Haiku 5.5 streams about 4.6 times faster than GLM-5.3 Flash while costing less per output token. For a batch job that nobody watches, GLM's slow speed does not matter. For a live chat, it does. A cheaper model that makes users wait can cost you more in lost sessions than it saves in tokens, and nothing on a pricing page will warn you about that until a customer complains in a support ticket three weeks after launch.
Who should skip the cheapest option?
Three groups should not chase the lowest price.
- Teams whose bill is dominated by wrong answers. If a bad output triggers five minutes of human review, a model that is 10% worse can cost more than it saves.
- Anyone who needs a stable price for a year. Gemini's price is scheduled to double, Mistral's is a sale and Ling's is a trial.
- Regulated or confidential workloads. Cheap tiers often come with thin data terms.
For the first group, the step up is Claude Sonnet 5.5, at $2 in and $10 out. Above it sits Claude Opus 5.5. Our guide to which Claude model to use sets out where each tier pays off.
Only Haiku, GLM and DeepSeek carry prices without a dated change on the page, and DeepSeek reserves the right to adjust its rates at any time.
What do the cheap tiers do with your data?
This is the part price pages do not show, and it is where the lowest bill can turn into the highest risk.
- Anthropic. Its Commercial Terms say it may not train models on customer content from its services.
- Google. On the paid Gemini API and Vertex AI, prompts and responses are not used to train models. They are kept briefly for abuse detection, safety and legal reasons.
- Mistral. Its help centre says free-mode data may be used for training with an opt-out. Pay-as-you-go customers can opt out at any time.
- DeepSeek. The API terms state no retention period or training policy. The privacy policy covers the consumer chat app, not the API.
- Z.ai. The model card gives no retention or training terms for the hosted API.
- InclusionAI. Nothing is published for Ling 3.1 Flash. Retention depends on the host you use.
The free Google AI Studio tier has different terms from the paid tier, so read them before sending anything private.
Readers often ask for the cheapest provider in Europe. We cannot give a confirmed answer. Mistral is the one European company in this group, but we found no statement on its pages that this preview API keeps data in Europe, so we make no claim. If European data residency is a requirement, get it from the vendor in writing before you build.
When is a cheap model the wrong call?
The strongest objection to this guide is that all six scores are close, so the choice barely matters. Pick on price and move on. For narrow tasks there is some truth in that.
But the Artificial Analysis index is an average across many tests. A model that scores 42 overall can score far lower on your task. It can also fail in ways an average hides, such as ignoring a formatting rule or inventing a field.
The fix is cheap. Really cheap. Take 50 real prompts from your own traffic, run them through two or three candidates and compare the outputs by hand. It takes an afternoon. It is the only comparison that reflects your data, your tone requirements, your output format and the odd edge cases your users actually send, which no public benchmark covers and no vendor page will describe honestly.
Score each answer pass or fail against a short checklist you write before you look. Count the failures. Then multiply the failure rate by what a failure costs you, and add that figure to the token bill. Now compare. Often the cheapest model on paper is not the cheapest model in practice. The guide to choosing the right LLM lists what to test for.
What would change this ranking?
Four dated events could reorder the table within three months.
- Ling's trial ends. Paid rates near the recorded ones keep it mid-table. A rise drops it out.
- Gemini's 1 January change. Unless Google revises it, Gemini becomes the most expensive model here by a wide margin.
- Mistral's sale ends or weights ship. Open weights would let you self-host and ignore the API price, if the licence allows commercial use.
- Your prompts pass 100K tokens. Haiku's tier changes, so re-run the sums.
Treat the tables as a snapshot dated 9 October 2026. Prices move. Check each vendor's page again before you sign anything, because a price that was true on the morning this guide was written can be a discount, a trial or a mistake by the time your first invoice arrives, and browse the full model directory or the other HokAI guides for newer comparisons.
Frequently asked questions
What is the cheapest LLM API right now? For prompts up to 100,000 tokens, Claude Haiku 5.5. GLM-5.3 Flash and DeepSeek V4.1 Flash off-peak are close behind.
Is a cheap LLM API good enough for coding? The index is an average, not a coding test. Run your own coding prompts before you commit.
Are free LLM APIs worth using? They suit prototypes. Ling 3.1 Flash was free on several hosts for about two weeks, then paid. Free tiers often carry different data terms from paid ones.
Is a Chinese-hosted API safe for company data? That depends on your rules, not on the vendor's nationality. DeepSeek and Z.ai publish thin API data terms, so get written answers before sending confidential data.
Frequently asked questions
What is the cheapest LLM API right now?
For prompts up to 100,000 tokens it is Claude Haiku 5.5, priced at $0.10 per million input tokens. GLM-5.3 Flash and DeepSeek V4.1 Flash off-peak are close behind. Longer prompts change the order.
Is a cheap LLM API good enough for coding?
The Artificial Analysis index is an average, not a coding test. Run 50 of your own coding prompts through two or three candidates and compare by hand before you commit.
Are free LLM APIs worth using?
They suit prototypes. Ling 3.1 Flash was free on several hosts on 8 October for about two weeks, then paid. Free tiers often carry different data terms from paid ones.
Which cheap LLM APIs will get more expensive?
Gemini 3.8 Flash is scheduled to double on 1 January 2027, to $1.50 in and $7.50 out. Ling 3.1 Flash ends its free trial, and Mistral Large 4 is on a sale with no stated end date.
Does a cheaper model always cost less per month?
No. Tokenizers, output verbosity, caching and long-prompt tiers change the bill. On a fixed 50M input and 10M output workload, the top three models land within $4 of each other.
Covered in this guide
- Claude Haiku 5.5: Anthropic's smallest and fastest 5.5 model, released October 7, 2026, with a 1M context window, built for high-volume work.
- Anthropic: Anthropic is a Public Benefit Corporation started by 7 ex-OpenAI researchers. It builds the Claude model family and counts Amazon and Google among its largest backers.
- Claude Opus 5.5: Anthropic's September 2026 flagship LLM with a 1 million token context window, built for long-running agentic coding and computer use.
- Claude Sonnet 5.5: Anthropic's fastest Sonnet-class model, launched in September 2026, with a 1M context and output more than 30% faster than Sonnet 5.
- DeepSeek: DeepSeek is a Chinese AI research company developing frontier language and reasoning models, including DeepSeek V4-Pro. Founded by High-Flyer hedge fund CEO Liang Wenfeng in 2023, the company is known for achieving GPT-level performance at dramatically lower compute and API cost.
- DeepSeek V4 Pro: DeepSeek V4 Pro: 1.6T-param open-source MoE (April 2026), 80.6% SWE-bench Verified with 1M token context under MIT license.
- DeepSeek V4.1 Flash: DeepSeek-V4.1-Flash is a 552B-parameter, MIT-licensed multimodal model with native vision, released in September 2026.
- Gemini 3.8 Flash: Google DeepMind shipped Gemini 3.8 Flash on September 2, 2026, a multimodal Gemini 3 model tuned for long-horizon coding and agentic enterprise work.
- GLM-5.3: GLM-5.3 arrived in August 2026 as Z.ai's coding- and cybersecurity-focused update, post-trained on the same base as its predecessor.
- GLM-5.3 Flash: Open-weight, natively multimodal MoE model from Z.ai with a 1M-token context window and MIT license, released in August 2026.
- InclusionAI: Ant Group's open AGI lab behind the free, MIT-licensed Ling, Ring and Ming model families, including the 1T-parameter Ling-1T.
- Ling 3.1 Flash: InclusionAI's text-only reasoning model for coding agents and long tasks, released on 30 September 2026 as the successor to InclusionAI's July flash model.
- Mistral: Mistral AI, founded in April 2023 in Paris by three ex-Meta researchers, builds Mistral, Mixtral, and Le Chat and raised $1.47B including $830M debt (Mar 2026).
- Mistral Large 3: Mistral Large 3 (Dec 2025) is a 675B MoE model (41B active) with 256K context and Apache 2.0 open weights.
- Mistral Large 4: Mistral Large 4 is Mistral AI's largest open-weight multimodal MoE model, launched 6 Oct 2026 in public preview with image input.
- OpenRouter: Single API endpoint for 300+ AI models from OpenAI, Anthropic, Google, and other providers, with one bill and no vendor lock-in.
- Z.ai: Z.ai (formerly Zhipu AI) is a Beijing lab behind the GLM model family, founded 2019 and listed on the Hong Kong Stock Exchange (2513) since January 2026.
Sources
Still deciding?
This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart MatchRelated guides
- The AI Tool Ecosystem in 2026: Buy the Meter, Not the CategoryBuyer's guideHow to pick, across a category
- AIML API vs OpenRouter: Which AI Gateway Should You Use in 2026?ComparisonHead-to-head, with a verdict
- Best AI Chatbots in 2026: Pick by the Job, Not the LeaderboardUpdatedRechecked against current sources
- Best AI Coding Assistants in 2026: Pick the Job, Not the BrandBuyer's guideHow to pick, across a category
- Best AI Companies in 2026: Who Is Actually LeadingBuyer's guideHow to pick, across a category
- Best AI for Writing a Business Plan in 2026: Tested Picks and a VerdictBuyer's guideHow to pick, across a category