What Mistral Large 4 Is Actually For
Mistral Large 4 is a 1.05 trillion parameter mixture of experts model from Mistral AI, released 6 October 2026 as a public preview. It activates 52 billion parameters per token, accepts text and images, and has no announced licence. Mistral promises downloadable weights by the end of October 2026.
The short version
Mistral Large 4 is a hosted preview with promised open weights, aimed at cyber work, EU data residency and image tasks. It trails cheaper closed models on general scores, and its $0.68 price is a sale. Use it for those niches only.
Mistral Large 4 launched on 6 October 2026 as a public preview, priced at half its own list price, with open weights promised for the end of the month.
The model is a 1.05 trillion parameter mixture of experts that activates 52 billion parameters per token. Mistral AI built it for cybersecurity, agents and image work. It says it trained Mistral Large 4 on 3,800 Nvidia Grace Blackwell GPUs in its own European datacenters. For a security lead, an EU procurement team or anyone planning to self-host, the question is narrow: is the preview worth building on before the weights and the licence arrive?
The short answer: yes for three jobs, no for the rest. This guide separates what Mistral reports from what an outside party measured, and it flags the two facts that are still open: the weights date and the licence.
What changed: a preview, a sale price and a promise
Three separate things landed on 6 October. They are easy to blur, because the launch post, the pricing page and the benchmark site each tell a slightly different part of the story and none of them says so.
First, a hosted preview. Mistral's launch post says you can try the preview API on Mistral Studio today. The model page labels it Public Preview, version 26.10, so it is not a generally available release with a stable version string.
Second, a price. Mistral's pricing page lists Large 4 at $0.68 per million input tokens. Output costs $2.09. Both numbers carry a "Sale price" label beside a struck-through original, and the originals are exactly double.
Third, a promise. Mistral says the weights will arrive by the end of October, along with more detail on the architecture and post-training. The launch post does not name a licence. Until the weights ship, nobody can download, audit or self-host the model.
That makes Large 4 an open-weight model in intent and a hosted API in practice. Artificial Analysis, the independent benchmarking site, still labels it proprietary today. That label should flip once the files are public, and it is accurate until then.
How the benchmarks read, and who ran them
Almost every number in the launch post comes from Mistral, and the post presents most of them as charts. Treat them as vendor claims until a third party repeats them. Here is what Mistral reports for the coding and agent work:
- 61.7% on DeepSWE v1.1, a software engineering benchmark.
- 59.4% on SWE-Atlas-QnA, which tests questions about real repositories.
- 28.3% on Terminal-Bench 4, which tests multi-step command line tasks.
- 49.8% on its own combined Coding Agent Index, which Mistral says beats DeepSeek V4 Pro 0813 and Qwen3.8 Max.
- 59.9% on AutomationBench, covering 657 business workflows across Gmail, Slack and Salesforce.
Mistral also ran a blind human rating of coding output with Surge AI. Large 4 scored 3.74 out of 5 and placed second of five models. Claude Opus 5 took first at 4.22. Kimi K3 scored 3.59 and GLM-5.3 scored 3.60. Five models is a thin field, so read the order, not the decimals.
The independent view is narrower. It is also cleaner. Artificial Analysis scores Large 4 at 38 on its Intelligence Index v4.3.2, against a median of 26 for comparable models. It measures 116.1 output tokens per second.
It also records that the model produced 200 million tokens to finish the index, which is about 2.5 times the median of 81 million and the kind of figure that quietly inflates a bill, a latency budget and a test suite all at once. The site's summary calls the model well priced, fast and very verbose.
One more caution on context. Mistral's model card lists 1M tokens. Artificial Analysis lists 524,000 for the live preview endpoint. Plan around the smaller number until Mistral reconciles the two.
What Large 4 costs: the $0.68 price is a sale price
The cost question has a trap in it. Two sources quote two different prices. Both are accurate.
Artificial Analysis shows $1.36 for input and $4.18 for output. That is the original price. Mistral's own docs show half of each, labelled as a sale. Cached input follows the same halving. The pages I read give no end date. Mistral has not said whether the lower price becomes permanent when the weights ship.
That matters for budgeting. A prototype priced at the sale rate can double its bill overnight if the label comes off. Build your cost model on the original rates and treat the lower ones as a bonus.
The table below puts Large 4 next to three rivals on a blended price, which weights input and output three to one. HokAI calculated the blended figures from each vendor's listed rates. The index scores come from Artificial Analysis as recorded on the linked HokAI model pages.
| Model | Blended $ per M tokens | AA index | Weights |
|---|---|---|---|
| Claude Haiku 5.5 | 0.20 | 43 | Closed |
| Ling 3.1 Flash | 0.45 | 41.1 | Closed |
| Large 4 at sale price | 1.03 | 38 | Promised |
| Large 4 at list price | 2.07 | 38 | Promised |
| Kimi K3 | 2.31 | 44 | Open |
Read it plainly. Cheap wins. Large 4 scores below all three rivals on the general index and costs more than two of them even at the sale price. Anthropic's pricing page charges Haiku 5.5 by prompt length, with the low rate applying only to short prompts of 100,000 tokens or fewer. Under that limit, the cheap closed models win on cost per point of general intelligence. Above it, Haiku 5.5 costs about the same as Large 4 at the sale price.
So price is not the reason to pick Large 4. If you only want the cheapest capable API, start with the LLM chooser and the Haiku 5.5 comparison. The model leaderboard shows current prices and scores side by side.
Who it affects: security teams, EU buyers and self-hosters
Mistral aimed Large 4 at three groups, and each has a real reason to look.
Security teams. This is the sharpest pitch. Mistral reports 82% on CyberGym-E2E, a test that asks a model to reproduce a real vulnerability and then patch it. It reports 93% on Cybench, a set of 40 security competition challenges. Mistral says several closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the CyberGym test because they refuse the task.
That is Mistral's claim. It is not an independent finding. The logic is still worth taking seriously. A defender who needs to prove a flaw is real can hit a refusal wall on a closed model. An open-weight model that runs under your own policies removes that wall. Mistral also says it measured a higher refusal rate than any other open model on harmful cyber prompts, so the model is not unguarded.
EU buyers. Mistral says it trained and serves the preview in its own European datacenters. It describes a European deployment that it runs end to end, independently of other digital service providers and under European law. Mistral says more than 160 languages appear in the training data, including every official EU language. For a public body or a regulated firm, that residency story is a feature the US labs do not offer in the same form.
Self-hosters. If the weights ship, a 1.05 trillion parameter model is a serious download. Only about 5% of parameters activate per token, which sets compute cost. The full set still has to sit in memory. That puts it out of reach of a single workstation. Teams that run models on their own hardware, and that already know what a multi-GPU server costs to rent or to buy, should read the local AI guide first and expect a multi-GPU server.
Image work is a fourth case. It is smaller. Mistral reports 42% on the Dense 200 visual grounding test against 41% for GPT-6 Astra, using a separate 1.6 billion parameter vision encoder. A one point gap on a vendor chart is a hint, not a ranking.
Why Large 4 isn't for most teams today
Most readers should not switch. Outside the three groups above, the case against is stronger than the case for.
The general index score is the first problem. A 38 trails Haiku 5.5 at 43, Ling 3.1 Flash at 41.1, Kimi K3 at 44 and DeepSeek V4 Pro 0813 at 53. Those last two rows come from the HokAI records for each model, which carry Artificial Analysis scores. A model that trails on the general index but leads on cyber is a specialist. Specialists make poor defaults.
Verbosity is second. It is expensive. Producing 2.5 times the median output tokens means your bill and your latency both grow beyond what the per-token price suggests. Artificial Analysis puts the cost of running its index at $1.13 per task at list price. Test your own prompts before trusting a price table.
Preview status is third. Mistral calls the release a preview and says the reinforcement learning run behind it is still in flight. Behaviour can change under a fixed model name. That breaks tests. A production workload needs a version that does not move.
The fourth is the licence. Mistral lists Mistral Large 3 under Apache 2.0, at a flat $0.50 per million input tokens. A team that needs open weights this month can run Large 3 today with a licence everyone understands. Large 4 offers no such clarity yet.
There is also no system card. Anyone who needs a safety document for an internal review has nothing to attach, and that alone can end a procurement conversation before it starts.
What to do: use the preview, wait for the weights, or skip
The right move depends on your job. Here it is, job by job.
- Security work. Trial the preview API on non-sensitive code. Compare refusals with your current closed model.
- EU procurement. Get the data-processing terms in writing. A launch post is not a compliance file.
- Self-hosting. Wait. Plan hardware after the weights and licence are public. Ollama is the usual first stop for small local models.
- Everyone else. Stay put. The small closed models in the table give more score per dollar. The Claude model guide covers Anthropic's tiers.
If you are unsure which bucket you are in, Smart Match asks about your stack and constraints and narrows the field. The Mistral AI listing covers the company's other products, and OpenRouter is one way to test several models behind one key, once a router lists the model.
A practical switching note: the API supports function calling, structured outputs, document question answering, batching and an agents endpoint. If you already call another provider through an OpenAI-style interface, a test swap is cheap. The cost shows up later, when you rewrite evaluations for a model that talks more than your current one.
The data and licence questions you cannot answer yet
Two sets of facts are open, and you should write them down as risks rather than assumptions.
On licence, Mistral has said nothing in the launch post. Large 3 uses Apache 2.0, but one model's licence is not a promise about the next. Some labs release large models under custom terms with usage limits. Do not plan a commercial product on the assumption that Large 4 matches Large 3.
On data, the launch post describes the European deployment but does not set out retention, training use or logging for the preview. Mistral also says vetted cyber partners and state authorities test the same model with reduced moderation. That is a statement about who else has access to a different configuration, not about your data. Ask for the terms before you send anything sensitive.
Mistral did publish safety results. It says the model resists 93.3% of attacks on Lakera's B3 agent security benchmark, and it reports a score of 1.691 out of 2 on the KORA benchmark. Those are Mistral's runs on public test sets. They suggest care went into the release. They do not replace an audit of your own use.
What would change this verdict
Four events would move the advice from "wait" to "adopt".
- A permissive licence. An Apache-style licence would make Large 4 a leading open option for cyber work.
- Outside proof. Artificial Analysis runs a cyber index. A repeat of the CyberGym and Cybench figures would back the specialist claim.
- A permanent low price. The gap to Ling 3.1 Flash stays open, but verbosity would hurt less.
- A better final model. Mistral says its training run is still going. The 38 could move.
Two events would confirm the cautious reading. The first is a delay past 31 October with no new date. The second is a licence that bars commercial use or caps deployment size.
Check back after the end of October. The weights date is the one fact that settles most of this.
Frequently asked questions
When do the Mistral Large 4 weights arrive?
Mistral says it will release the weights by the end of October 2026. The launch post gives no exact day. It also names no licence, so check Mistral's announcement before you plan a self-hosted deployment.
How much does Mistral Large 4 cost?
Mistral's docs list $0.68 per million input tokens and $2.09 per million output tokens, labelled as a sale price. The original prices are $1.36 and $4.18, which is what Artificial Analysis shows. Mistral has not published an end date for the sale.
Is Mistral Large 4 better than Claude Haiku 5.5?
Not on the general index. Artificial Analysis scores Large 4 at 38 and, per HokAI's model record, Haiku 5.5 at 43. Large 4 also costs more per token. Its case rests on cyber tasks, open weights and EU hosting.
Can I run Mistral Large 4 on my own hardware?
Not yet, because the weights are unreleased. Once they ship, the model has 1.05 trillion parameters that all need memory even though only 52 billion activate per token. Expect a multi-GPU server rather than a workstation.
Is Mistral Large 4 good for security work?
Mistral reports 82% on CyberGym-E2E and 93% on Cybench, and says closed models often refuse the CyberGym task. Those are Mistral's figures. Test the preview on non-sensitive samples before relying on them.
Covered in this guide
- Mistral AI: Mistral AI, founded in April 2023 in Paris by three ex-Meta researchers, builds Mistral, Mixtral, and Le Chat and raised $1.47B including $830M debt (Mar 2026).
- Mistral Large 4: Mistral Large 4 is Mistral AI's largest open-weight multimodal MoE model, launched 6 Oct 2026 in public preview with image input.
- Claude Opus 5: Anthropic is a Public Benefit Corporation started by 7 ex-OpenAI researchers. It builds the Claude model family and counts Amazon and Google among its largest backers.
- Claude Haiku 5.5: Anthropic's smallest and fastest 5.5 model, released October 7, 2026, with a 1M context window, built for high-volume work.
- Claude Opus 5.5: Anthropic's September 2026 flagship LLM with a 1 million token context window, built for long-running agentic coding and computer use.
- DeepSeek: DeepSeek is a Chinese AI research company developing frontier language and reasoning models, including DeepSeek V4-Pro. Founded by High-Flyer hedge fund CEO Liang Wenfeng in 2023, the company is known for achieving GPT-level performance at dramatically lower compute and API cost.
- DeepSeek V4 Pro 0813: DeepSeek-V4-Pro-0813 is DeepSeek's flagship 1.6T-parameter mixture-of-experts model, generally available since August 2026 with MIT-licensed open weights.
- GLM-5.3: GLM-5.3 arrived in August 2026 as Z.ai's coding- and cybersecurity-focused update, post-trained on the same base as its predecessor.
- GPT-6 Astra: OpenAI's flagship model, launched September 2026 as the first ever rated at the Preparedness Framework's Critical cybersecurity level.
- Kimi K3: 2.8T-parameter open-weight MoE model from Moonshot AI (July 2026) with a 1M-token context window and 93.5% GPQA Diamond, the top open score.
- Ling 3.1 Flash: InclusionAI's text-only reasoning model for coding agents and long tasks, released on 30 September 2026 as the successor to InclusionAI's July flash model.
- Mistral AI listing: Mistral AI is a Paris-based AI lab valued at $13.7B, building open-weight foundation models such as Mistral Large 3 alongside the Vibe consumer assistant.
- Mistral Large 3: Mistral Large 3 (Dec 2025) is a 675B MoE model (41B active) with 256K context and Apache 2.0 open weights.
- Moonshot AI: Beijing AI lab founded in 2023 by Yang Zhilin, maker of Kimi and the open-weight Kimi K2 models, valued at $20B in 2026.
- Ollama: MIT-licensed runtime that runs open-weight LLMs on macOS, Windows and Linux with a local API, plus an optional paid cloud for larger models.
- OpenAI: OpenAI builds ChatGPT, the GPT-6 family (Astra, Sol, Luna), Codex and the OpenAI API. It reports 1.2B weekly ChatGPT users and an $852B valuation after its March 2026 round.
- OpenRouter: Single API endpoint for 300+ AI models from OpenAI, Anthropic, Google, and other providers, with one bill and no vendor lock-in.
- Qwen3.8 Max: Qwen3.8-Max is Alibaba's flagship MoE model, launched Aug 2026, with a 1M-token context, native multimodal input, and a text-only open-weight checkpoint released weeks later.
Sources
Still deciding?
This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart MatchRelated guides
- What Happened When 7 AI Agents Got Real Bank Accounts and No SupervisionAnalysisWhat changed and who it affects
- The AI Tool Ecosystem in 2026: Buy the Meter, Not the CategoryBuyer's guideHow to pick, across a category
- AIML API vs OpenRouter: Which AI Gateway Should You Use in 2026?ComparisonHead-to-head, with a verdict
- Best Agentic AI in 2026: Pick the Job, Not the GeneralistBuyer's guideHow to pick, across a category
- Best AI Assistant Apps for Android in 2026, Now That Google Assistant Is DeadBuyer's guideHow to pick, across a category
- Best AI Chatbots in 2026: Pick by the Job, Not the LeaderboardUpdatedRechecked against current sources