MiMo-V2.6-Pro-RL ships a 1M-token context window and a 128K-token max output behind a 5-layer speculative decoder built for long agentic sessions. It suits teams that want an MIT-licensed, self-hostable alternative to closed frontier models, not teams needing low-latency small-prompt chat or native audio and image output.
MiMo-V2.6-Pro-RL is Xiaomi's flagship open-weight reasoning model: a 1.02-trillion-parameter Mixture-of-Experts system with 42 billion active parameters, MIT licensed since its September 2026 debut. Independent benchmarking gives it the top intelligence score of any open-weight model at 46.32, and it reads text, image, video, and audio across a 1 million token context window.
Where it sits
- $0.544/M$ per 1M tokensBlended price (3:1)Lower is better#20 / 71peer median $1.69/Mvendor price, checked by HokAI
- 49 tok/stokens/sOutput speedHigher is better#36 / 43peer median 85 tok/scited: Artificial Analysis
Cheaper than 70% of the 71 GA models with a published price, and rank 36 of 43 on output speed as cited from Artificial Analysis.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Xiaomi Corporation · Family: MiMo V2.6
More about Xiaomi Corporation on HokAI
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, video, audio · Output: text
About MiMo-V2.6-Pro-RL
MiMo-V2.6-Pro-RL is the flagship checkpoint of Xiaomi's MiMo-V2.6 series, released open-weight on September 21, 2026 under the MIT license. Development was led by Luo Fuli, who joined Xiaomi's MiMo team in late 2025 after previously working at DeepSeek. It is a sparse Mixture-of-Experts transformer with 1.02 trillion total parameters and 42 billion active per token, routed across 384 experts (8 per token) over 70 layers. Xiaomi positions it as the successor to MiMo-V2.5-Pro, built for long-horizon agentic work: coding, cybersecurity, computer-use, and research.
MiMo-V2.6-Pro-RL scores 46.32 on the Artificial Analysis Intelligence Index, the highest of any open-weight model at launch, ahead of Kimi K3 (44) and GLM-5.3 (45), though Xiaomi says leading closed-source models score around 53. On agentic evals Xiaomi reports, it scores 89.9 on Terminal-Bench 2.1 and 82.0 on OSWorld-Verified, close to Claude Opus 5 (89.1, 83.4) and GPT-5.6 Sol (88.8, 83.0). It trails those same models on cybersecurity exploit evals: 47.9 vs Claude Opus 5's 70.0 on ExploitBench.
The model ships a 1 million token context window and a 128,000 token max output, confirmed on Xiaomi's own API docs. A 5-layer speculative multi-token-prediction decoder predicts 7 tokens per forward pass to keep long sessions fast despite the trillion-parameter backbone. Artificial Analysis independently measured 48.6 output tokens per second and a 4.01-second time-to-first-answer-token, both toward the slow end of its open-weight peer group.
Input is natively omnimodal: text, image, video, and audio, via a 681M-parameter vision encoder and a two-part audio stack (a 308M AudioTokenizer plus a 127M audio patch encoder). Output is text only; there is no native audio or image generation. The model supports tool calling, streaming, structured output, web search, and context caching.
Xiaomi's API prices MiMo-V2.6-Pro-RL at $0.435 per 1M input tokens on a cache miss, $0.0036 per 1M on a cache hit (about a 99% discount), and $0.87 per 1M output tokens. Artificial Analysis computes a blended rate of $0.18 per 1M tokens, below the class median. A separate Token Plan subscription and a batch API for non-real-time workloads are also available.
Weights are published on Hugging Face and ModelScope under MIT, with vLLM and SGLang recipes; the full checkpoint needs 8 GPUs (vLLM) or 2 nodes of 16 GPUs (SGLang). It is not yet on Bedrock, Vertex, or Azure; access is via Xiaomi's own API, AI Studio, MiMo Desktop, and OpenRouter, with only 3 API providers live at listing time.
Xiaomi's Trust Center publishes training-data summaries for the wider MiMo-V2 family and MiMo-V2.5, but not MiMo-V2.6 specifically. Its technical report describes an "Aligned RL" stage where the model rewrites its own misaligned turns into grounded responses, plus adversarial screening and verifier cross-checks against reward hacking during RL.
MiMo-V2.6-Pro-RL suits teams wanting an open-weight model competitive with closed frontier models on agentic and coding benchmarks, needing the full 1M-token context, able to self-host on multi-GPU infrastructure or accept an early-stage API. It is weaker for exploit-development work, native audio/image output, low-latency small-prompt use, or enterprise compliance needs.
Training used one asynchronous GRPO reinforcement-learning run mixing coding, agent, visual, and cybersecurity tasks in the same batches, at 1,568 prompts times 16 rollouts per step. Xiaomi reports the Pro run completed 30 steps in under six days using about 750,000 trajectories at a reported cost of $2.62 million (a further $850,000 covered the smaller MiMo-V2.6-Flash-RL checkpoint trained alongside it).
MiMo-V2.6-Pro-RL beats its predecessor MiMo-V2.5-Pro by a wide margin on every benchmark reported side by side, for example 89.9 vs 65.2 on Terminal-Bench 2.1. Xiaomi has already previewed MiMo-V3, reported to adopt a new "HySparse2" architecture to cut long-context serving costs, as the next model in the line. A smaller research sibling, MiMo-V2.6-Distill-Qwen-9B, was released alongside it for self-hosting on far less hardware.
Pricing
Verified 2026-09-28 on Xiaomi's own pricing page: fresh input tokens run $0.435 per 1M, previously-cached prompt prefixes drop to $0.0036 per 1M (roughly a 99% saving), and generated output runs $0.87 per 1M. Token Plan subscriptions and a separate batch API exist for high-volume or non-real-time work; those rates were not independently verified for this listing.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.013 | $0.0009 | $0.014 |
| Support reply | $0.0009 | $0.0003 | $0.0011 |
| One coding agent run | $0.087 | $0.017 | $0.104 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Sparse Mixture-of-Experts backbone: 384 routed experts with 8 activated per token across 70 transformer layers; only the first layer uses dense global attention, the rest interleave sliding-window and global attention with MoE feed-forward blocks.
- 1M-token context window: 128K max output tokens, paired with a 5-layer speculative decoder that drafts several tokens per step so long agentic sessions stay responsive despite the model's size.
- Native omnimodal input: A 681M-parameter vision encoder and a two-stage audio encoder (308M + 127M parameters) let the model read text, image, video, and audio in one call; output remains text only.
- Groupwise agentic RL training: Trained with a single mixed reinforcement-learning run (GRPO) across coding, agent, visual, and cybersecurity tasks together, using an automated grader that ranks passing rollouts by quality rather than a simple pass/fail signal.
- Prompt caching: Repeated prompt prefixes bill at a discounted cache-hit rate instead of the full input price; see the pricing FAQ for the exact figures.
Pros
- Highest Artificial Analysis Intelligence Index score of any open-weight model at launch (46.32), ahead of Kimi K3 and GLM-5.3.
- MIT license with no commercial-use restrictions, plus published vLLM and SGLang self-hosting recipes.
- Beats the closed-source Claude Opus 5 comparison model on Terminal-Bench 2.1 (89.9 vs 89.1) despite being free to self-host.
Cons
- Output speed (48.6 tokens/second) and latency (4.01s time-to-first-token) trail the median open-weight model in its size class.
- Not yet available on AWS Bedrock, Google Vertex AI, or Azure; only 3 API providers total were live at the time of this listing.
- No published system card, red-team partner disclosure, or MiMo-V2.6-specific training-data summary as of this listing.
Benchmarks
- GDPval-AA v2: 1,673 vendor-reported · 28 Sep 2026 — Real knowledge-work deliverables judged against professionals, run by Artificial Analysis.
- OSWorld Verified: 82% vendor-reported · 28 Sep 2026 — Tasks completed by operating a real desktop, % solved.
- Terminal-Bench 2.1: 89.9% vendor-reported · 28 Sep 2026 — Multi-step tasks completed in a real command line, % solved.
- AA Intelligence Index: 46.3 cited: Artificial Analysis · 28 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- AA blended price: $0.18/M cited: Artificial Analysis · 28 Sep 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
- Output speed: 49 tok/s cited: Artificial Analysis · 28 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does MiMo-V2.6-Pro-RL cost to run?
On Xiaomi's own API, verified 2026-09-28, fresh input tokens run $0.435 per 1 million, cached prompt prefixes drop to $0.0036 per 1 million (roughly 99% cheaper), and generated output runs $0.87 per 1 million. That works out to an $0.18 blended rate under Artificial Analysis's methodology, below the median for models in its class. High-volume or non-real-time users can also subscribe to a Token Plan or use the batch API, though those rates were not independently verified here.
MiMo-V2.6-Pro-RL vs Claude Opus 5: which scores higher?
It depends on the task: MiMo-V2.6-Pro-RL scores 89.9 on Terminal-Bench 2.1 against Claude Opus 5's 89.1, essentially a tie in Xiaomi's own reported comparison. Claude Opus 5 leads on OSWorld-Verified (83.4 vs 82.0) and by a wide margin on cybersecurity exploit evals such as ExploitBench (70.0 vs 47.9). MiMo-V2.6-Pro-RL's 46.32 tops every open-weight model on the Artificial Analysis Intelligence Index, though Xiaomi puts the best closed-source models around 53 on the same scale.
Are MiMo-V2.6-Pro-RL's weights publicly available?
Yes. MiMo-V2.6-Pro-RL is released under the MIT license, which permits commercial use, and the weights are published on Hugging Face and ModelScope. Self-hosting the full 1.02-trillion-parameter checkpoint needs a multi-GPU setup: Xiaomi's own recipes use either 8 GPUs with vLLM or two nodes of 16 GPUs each with SGLang.
Does MiMo-V2.6-Pro-RL train on user data?
Xiaomi has not published a MiMo-V2.6-specific data-retention or training-on-inputs policy as of this listing; its Trust Center hosts training-data summaries for the broader MiMo-V2 family and for MiMo-V2.5, but not yet for V2.6. Teams with strict data requirements can self-host the MIT-licensed weights instead of using the hosted API, which removes the question entirely.
When should you choose MiMo-V2.6-Pro-RL over a rival model?
Choose it for long-horizon coding or computer-use agents that benefit from a 1M-token context and can tolerate self-hosting or an early-stage hosted API, especially where an MIT license and no per-seat cost matter. Avoid it for offensive-security or exploit-development work, where it trails named closed-source models by a wide margin, and for latency-sensitive small-prompt chat, where its 48.6 tokens/second output speed is below the class median.