by Xiaomi Corporation

MiMo-V2.6-Flash-RL review, pricing and limits

The efficiency tier of Xiaomi's MiMo-V2.6 series: near-Pro-RL agentic performance at roughly a third of the price, built for high-frequency and large-scale workflows.

  • ga
  • open source
  • multimodal
  • MiMo V2.6 family
checked

Flash-RL pairs a 1M-token context window with 128K max output and 59.2 tokens/second throughput, faster than its Pro-RL sibling. Choose it for high-volume, cost-sensitive agentic work; choose Pro-RL instead when the intelligence gap matters or when confirmed video and audio input is required.

MiMo-V2.6-Flash-RL is the efficiency-tier model in Xiaomi's MiMo-V2.6 series: a 309-billion-parameter Mixture-of-Experts system with 15 billion active parameters, MIT licensed since its September 2026 debut. Its Artificial Analysis Intelligence Index score of 38 sits well above the open-weight median, and it reads text and image input across a 1 million token context window at roughly a third of Pro-RL's per-token price.

Where it sits

  • $0.175/M$ per 1M tokensBlended price (3:1)Lower is better#9 / 71peer median $1.69/Mvendor price, checked by HokAI
  • 59 tok/stokens/sOutput speedHigher is better#30 / 43peer median 85 tok/scited: Artificial Analysis

Cheaper than 87% of the 71 GA models with a published price, and rank 30 of 43 on output speed as cited from Artificial Analysis.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Xiaomi Corporation · Family: MiMo V2.6

More about Xiaomi Corporation on HokAI

Context window: 1,000,000 tokens · Max output: 128,000

Input modalities: text, image · Output: text

About MiMo-V2.6-Flash-RL

MiMo-V2.6-Flash-RL is the efficiency-tier checkpoint of Xiaomi's MiMo-V2.6 series, released open-weight on September 21, 2026 under the MIT license, led by the same Luo Fuli-headed MiMo team as MiMo-V2.6-Pro-RL. It is a smaller sparse Mixture-of-Experts transformer: 309 billion total parameters with 15 billion active per token, routed across 256 experts (8 activated) over 48 layers. Xiaomi built it as the balance point of the two RL-trained V2.6 checkpoints, for high-frequency calls and large-scale workflows where trillion-parameter Pro-RL would cost too much.

Flash-RL scores 38 on the Artificial Analysis Intelligence Index, well above the open-weight median (18) though behind Pro-RL's 46.32. On agentic evals Xiaomi reports directly, it scores 87.6 on Terminal-Bench 2.1 and 80.8 on OSWorld-Verified, both close to Pro-RL's 89.9 and 82.0 and to the closed-source Claude Opus 5 (89.1, 83.4) named in the shared technical report. It trails badly on cybersecurity exploit evals: 25.3 on ExploitBench against Claude Opus 5's 70.0.

The model ships the same 1 million token context window as Pro-RL, with a 128,000 token max output and an identical 5-layer speculative decoder design. Artificial Analysis measured 59.2 output tokens per second, faster than Pro-RL's 48.6 but still under the open-weight class median of 85.7 t/s, plus a 3.80-second time-to-first-answer-token. Flash-RL was also the more verbose of the two on the Intelligence Index, generating 240 million output tokens against Pro-RL's 140 million.

Xiaomi's own product page and this model's Hugging Face card both list Flash-RL's input modality as text, image, video, and audio, matching Pro-RL. Artificial Analysis, which tests models directly rather than citing vendor claims, found only text and image input actually working through Flash-RL's API as of this listing. Output remains text only in every source.

Xiaomi's API prices Flash-RL far below Pro-RL: $0.14 per 1M input tokens on a cache miss, $0.0028 per 1M on a cache hit (about a 98% discount), and $0.28 per 1M output tokens. Artificial Analysis computes a blended rate of $0.06 per 1M tokens, a third of Pro-RL's $0.18.

Weights are published on Hugging Face and ModelScope under MIT. Self-hosting recipes are lighter than Pro-RL's: SGLang with tensor-parallel 8 and data-parallel 2, or vLLM with tensor-parallel-size 4, roughly half the GPU count Pro-RL calls for. Distribution is narrower than Pro-RL: only 1 API provider was live on Artificial Analysis's tracker at listing time, versus 3 for Pro-RL; access otherwise runs through the same first-party channels (Xiaomi's own API, AI Studio, MiMo Desktop, OpenRouter).

As with Pro-RL, Xiaomi's Trust Center had not published a MiMo-V2.6-specific training-data summary or system card at listing time. The shared technical report describes the same "Aligned RL" self-correction stage and adversarial/verifier hardening used to train both checkpoints together.

Flash-RL suits high-frequency, cost-sensitive workloads needing agentic and coding competence close to Pro-RL's, and teams that can accept its narrower API distribution or self-host on roughly half the GPU budget. It is weaker for tasks needing Pro-RL's top intelligence score, offensive-security work, or workflows relying on the vendor's claimed video/audio input, which Artificial Analysis could not confirm.

Flash-RL trained alongside Pro-RL in the same reinforcement-learning campaign: Xiaomi reports its run also completed 30 GRPO steps in under six days, sharing the roughly 750,000-trajectory batch structure, at a reported cost of $850,000, well below Pro-RL's $2.62 million.

There is no MiMo-V2.5-Flash predecessor in Xiaomi's comparison table, so Flash-RL's closest lineage comparison is MiMo-V2.5-Pro, which it beats on every shared benchmark, for example 87.6 vs 65.2 on Terminal-Bench 2.1. Xiaomi has already previewed MiMo-V3, reported to adopt a new "HySparse2" architecture to cut long-context serving costs, as the next model in the line. A smaller research sibling, MiMo-V2.6-Distill-Qwen-9B, was released alongside it for self-hosting on a single GPU.

Pricing

Verified 2026-09-28 on Xiaomi's own pricing page: fresh input runs $0.14 per 1M tokens, cache-hit input drops to $0.0028 per 1M (about 98% cheaper), and output runs $0.28 per 1M tokens, roughly a fifth of Pro-RL's rate across the board. A separate Token Plan subscription and batch API exist for high-volume or non-real-time work; those rates were not independently verified for this listing.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.0042$0.0003$0.0045
Support reply$0.0003$0.0001$0.0004
One coding agent run$0.028$0.0056$0.034

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Sparse Mixture-of-Experts backbone: 256 routed experts with 8 activated per token across 48 transformer layers and a 4096-dimension hidden size, about half the layer count of Pro-RL's 70-layer backbone.
  • 1M-token context window: 128K max output tokens on the same 1M-token context as Pro-RL, paired with a 5-layer speculative decoder to keep long agentic sessions responsive.
  • Vision-capable input: Reads text and image input; Xiaomi's own pages also list video and audio, though Artificial Analysis's independent testing did not confirm those two modalities working via the API.
  • Groupwise agentic RL training: Trained in the same mixed reinforcement-learning run as Pro-RL (GRPO across coding, agent, visual, and cybersecurity tasks), using an automated grader that ranks passing rollouts by quality rather than a simple pass/fail signal.
  • Lighter self-hosting footprint: Xiaomi's reference recipes need 4 GPUs (vLLM) or 8 GPUs (SGLang), roughly half of what the full Pro-RL checkpoint requires.

Pros

  • Costs about a third of Pro-RL's blended per-token rate while staying within a few points of it on Terminal-Bench 2.1 and OSWorld-Verified.
  • MIT license with no commercial-use restrictions, plus a lighter self-hosting footprint than Pro-RL (4-8 GPUs vs 8-16).
  • Faster output than its larger sibling: 59.2 tokens/second against Pro-RL's 48.6, per Artificial Analysis.

Cons

  • Vendor markets video and audio input identical to Pro-RL, but Artificial Analysis's independent tests only confirmed text and image working via the API.
  • Hosted-API distribution is thinner than Pro-RL's right now: Artificial Analysis tracked just 1 provider for Flash-RL versus 3 for its sibling.
  • Generated 240 million output tokens on the Intelligence Index, more verbose than the larger Pro-RL's 140 million.

Benchmarks

  • OSWorld Verified: 80.8% vendor-reported · 28 Sep 2026 — Tasks completed by operating a real desktop, % solved.
  • Terminal-Bench 2.1: 87.6% vendor-reported · 28 Sep 2026 — Multi-step tasks completed in a real command line, % solved.
  • AA Intelligence Index: 38 cited: Artificial Analysis · 28 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
  • AA blended price: $0.06/M cited: Artificial Analysis · 28 Sep 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
  • Output speed: 59 tok/s cited: Artificial Analysis · 28 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

How much do you pay to use MiMo-V2.6-Flash-RL?

On Xiaomi's own API, verified 2026-09-28, fresh input tokens run $0.14 per 1 million, cached prompt prefixes drop to $0.0028 per 1 million (roughly 98% cheaper), and generated output runs $0.28 per 1 million. That works out to a $0.06 blended rate under Artificial Analysis's methodology, about a third of what Pro-RL costs. High-volume or non-real-time users can also subscribe to a Token Plan or use the batch API, though those rates were not independently verified here.

How does MiMo-V2.6-Flash-RL compare on benchmarks vs MiMo-V2.6-Pro-RL?

On the Artificial Analysis Intelligence Index, Flash-RL's 38 trails Pro-RL's 46.32, but the gap narrows sharply on agentic evals: 87.6 vs 89.9 on Terminal-Bench 2.1, despite Flash-RL running roughly a third of the active parameters. It also costs about a third as much and outputs faster (59.2 vs 48.6 tokens/second), making it the better pick for high-volume workloads where the small intelligence gap does not matter.

Can you self-host MiMo-V2.6-Flash-RL?

Yes. It ships under the MIT license, which permits commercial use, with weights published on Hugging Face and ModelScope. Xiaomi's reference setup calls for 4 GPUs under vLLM or 8 under SGLang, about half the hardware Pro-RL's full checkpoint needs.

Does MiMo-V2.6-Flash-RL retain or train on API inputs?

Xiaomi has not published a MiMo-V2.6-specific data-retention or training-on-inputs policy as of this listing; its Trust Center hosts summaries for the broader MiMo-V2 family and MiMo-V2.5, but not yet V2.6. Teams with strict data requirements can self-host the MIT-licensed weights instead of using the hosted API.

Who should skip MiMo-V2.6-Flash-RL?

Teams that specifically need the top-end intelligence score in the family should use Pro-RL instead, since Flash-RL trails it by 8 points on the Artificial Analysis Intelligence Index. Teams building a workflow around video or audio input should also be cautious: Xiaomi markets full omnimodal support, but Artificial Analysis's independent testing only confirmed text and image working through the API.

More AI Models on HokAI

Visit MiMo-V2.6-Flash-RL Official Page