Mercury 2.5 targets latency-bound work such as voice agents, search pipelines and coding subagents, where one user action fans out into dozens of model calls. It offers 4 reasoning-effort levels and a 65,536-token output cap, but it performs like a small cost-optimized model, so keep a stronger model for the hardest steps.
Mercury 2.5 is a diffusion language model from Inception, released on 8 September 2026, that Artificial Analysis scores 12.3 on its Intelligence Index. It refines many tokens in parallel instead of writing them one at a time, which Inception reports at 1,107 tokens per second on standard NVIDIA GPUs, and it accepts text only.
Where it sits
- $0.068/M$ per 1M tokensBlended price (3:1)Lower is better#4 / 76peer median $1.86/Mvendor price, checked by HokAI
- 711 tok/stokens/sOutput speedHigher is better#2 / 48peer median 87 tok/scited: Artificial Analysis
Cheaper than 96% of the 76 GA models with a published price, rank 2 of 48 on output speed as cited from Artificial Analysis, and one of 5 whose vendor states it may train on customer data.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Inception · Family: Mercury
Context window: 260,000 tokens · Max output: 65,536
Input modalities: text, tool-calls · Output: text, tool-calls
About Mercury 2.5
Mercury 2.5 is the flagship of the Mercury family from Inception and the successor to Mercury 2, which launched on 24 February 2026. Inception released it on 8 September 2026 and describes it as the largest diffusion language model trained so far. A diffusion LLM (dLLM) is a Transformer trained to denoise text, so it drafts and revises blocks of tokens in parallel instead of predicting them left to right. Weights are closed, and access is through an API.
Inception reports a 40% intelligence gain over Mercury 2 and says quality is comparable to cost-optimized models such as GPT-5.6 Luna (at low effort), Gemini 3.5 Flash-Lite and Claude Haiku 4.5. Independent data is more modest. Artificial Analysis rates the model below average in intelligence, and OpenRouter's benchmark tab, which relays Artificial Analysis data, shows 11.8% on Humanity's Last Exam, 38.5% on SciCode, 71.7% on a long-context reasoning test and 22.7% accuracy on AA-Omniscience. The same tab lists 0.0% on GDPval-AA and CritPt, which is a warning against using it for long professional-task agents or research-level physics.
Speed is the reason to pick it. Inception reports 1,107 tokens per second on widely available NVIDIA GPUs. Artificial Analysis measured 711.2 on its test, and on 3 October 2026 OpenRouter showed a live median of 261 tokens per second with a median response time of 1.82 seconds. The spread comes from prompt length, reasoning effort and load, so time your own prompts before committing.
Access is through the Inception API, OpenRouter and Baseten, and the API follows the OpenAI chat format, so frameworks such as LangChain connect with a new base URL. Inception names Augment Code, which moved context compaction, model routing and tool search to Mercury and reports an 82% latency cut on compaction.
Choose another model when you need image or audio input, reproducible seeded sampling, or the hardest multi-step reasoning. If open weights matter, DiffusionGemma is the nearest diffusion alternative on HokAI. On HokAI it sits in the proprietary, text-input and 128K to 400K context groups, next to the other Inception models.
Pricing
Mercury 2.5 lists at $0.20 per million input tokens and $0.75 per million output tokens, with cached input at $0.02. An 80% launch discount, with no end date published, currently bills $0.04, $0.15 and $0.004. HokAI stores the discounted rates because the API charges them today, so expect the stored figures to move to list when the discount ends. New accounts get a one-time credit of 100 million free tokens, which is a trial rather than a free tier. HokAI's records list [Gemini 3.5 Flash-Lite](/hub/models/gemini-3.5-flash-lite) at $0.30 input and $2.50 output and [GPT-5.6 Luna](/hub/models/gpt-5.6-luna) at $1 input and $6 output.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.0012 | $0.0001 | $0.0013 |
| Support reply | $0.0001 | $0.0000 | $0.0001 |
| One coding agent run | $0.0080 | $0.0030 | $0.011 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Diffusion decoding: Refines blocks of text in parallel instead of one token at a time, and the API can stream blocks or show the denoising effect.
- Four reasoning levels: reasoning_effort accepts instant, low, medium (the default) and high, so one model serves both reflex turns and multi-step planning.
- 260K context, 65,536-token output: Holds large retrieved context in one call and can return long structured results.
- Parallel tool calls and JSON schemas: OpenAI-compatible tool calling returns several calls in one turn, and schema-aligned output constrains replies to typed JSON.
- Automatic prefix caching: Cached input bills at one tenth of the normal input rate with no request parameter, though no cache lifetime is published.
Pros
- Built for fan-out workloads: Augment Code reports compaction dropping from about 150 seconds to 27 after moving to Mercury.
- Cheap enough to call on every step of an agent loop, well under the list prices of the small models Inception compares itself with.
- Drop-in for OpenAI-style clients, so adoption means a new base URL and model name.
Cons
- Independent scores sit well below the small models Inception benchmarks against, so hard reasoning belongs on a stronger model.
- Live throughput is far below the headline figure, and the gap changes with prompt length and load.
- Default terms allow training on your prompts until you opt out in settings.
Benchmarks
- Humanity's Last Exam: 11.8% independent · 03 Oct 2026 — Expert-written questions across many fields, % correct.
- AA Intelligence Index: 12.3 cited: Artificial Analysis · 03 Oct 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- Output speed: 711 tok/s cited: Artificial Analysis · 03 Oct 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does Mercury 2.5 cost per million tokens?
Billing is per token: $0.04 input and $0.15 output per million at the 80% launch discount, rising to $0.20 and $0.75 at list, with cached input at one tenth of the input rate. An agent loop of 100 calls, each sending 5,000 tokens and returning 500, costs about $0.03 at launch pricing and $0.14 at list. Gemini 3.5 Flash-Lite lists at $0.30 input in HokAI's record, 1.5 times the Mercury list rate.
How does Mercury 2.5 compare with Gemini 3.5 Flash-Lite on quality?
Inception says quality is comparable to cost-optimized models, but the independent numbers are lower. Artificial Analysis gives Mercury 2.5 an index of 12.3, and HokAI's records show Gemini 3.5 Flash-Lite at 23 and GPT-5.6 Luna at 45, though those were logged on different dates and reasoning settings, so the gap may be overstated. Flash-Lite also takes images, audio, video and PDFs, while Mercury is the faster text-only option.
Is Mercury 2.5 open source or proprietary?
It is proprietary. Inception has not released weights, and you reach the model through its API, OpenRouter or Baseten, so you cannot self-host it. The terms let you use outputs for any purpose within their limits. DiffusionGemma, a diffusion model on HokAI under Apache 2.0, is the open alternative.
Does Mercury 2.5 train on your prompts?
By default the terms of use allow Inception to train on what you submit, and you opt out by switching off 'Improve the model for everyone' in the platform's user settings. The enterprise page promises no training on customer data, configurable retention and private networking, and zero-retention arrangements go through support. No SOC 2 or ISO 27001 certificate was found on 3 October 2026.
Who should use Mercury 2.5, and who should skip it?
Use it for chained, latency-bound calls: voice turns, query rewriting, reranking, context compaction and tool search. Skip it for image or audio input, repeatable seeded output, and hard reasoning, where Gemini 3.5 Flash-Lite or GPT-5.6 Luna give more headroom. Its 3,000-request-per-minute pay-as-you-go limit is also worth checking against high fan-out traffic.
Top Alternatives
- GPT-5.6 Luna: Pick Mercury 2.5 for the lowest latency and cost per call on text; pick GPT-5.6 Luna when you need image input or stronger reasoning.
- Gemini 3.5 Flash-Lite: Pick Mercury 2.5 for diffusion speed on text agents; pick Gemini 3.5 Flash-Lite for its 1M-token context and audio, video and PDF input.
- DiffusionGemma 26B-A4B: Pick Mercury 2.5 for a managed reasoning API with tool calling; pick DiffusionGemma when you want open Apache 2.0 weights you can self-host.