by Xiaomi Corporation

MiMo-V2.6-Distill-Qwen-9B review, pricing and limits

A small, self-hostable research checkpoint in Xiaomi's MiMo-V2.6 family: an MIT-licensed SFT distillation of Qwen3.5-9B for agentic coding, cybersecurity, and visual tasks.

  • ga
  • open source
  • multimodal
  • MiMo V2.6 family
checked

This 9B checkpoint trains on 77.4 billion tokens split across code, cybersecurity, general agent tasks, and visual coding, with 20 community finetunes already published. It suits researchers and self-hosters wanting a small, MIT-licensed agentic base, not teams needing a hosted API or flagship-level intelligence.

MiMo-V2.6-Distill-Qwen-9B is a 9-billion-parameter dense model Xiaomi released by supervised fine-tuning Qwen3.5-9B on MiMo-generated agentic data. Released September 21, 2026 under the MIT license, it raises SWE-bench Verified from its base model's 60.0 to 61.1 and ships purely as open weights, with no hosted API.

Where it sits

  • 61.1%% solvedSWE-bench VerifiedHigher is better#25 / 30peer median 77.6%per source, see benchmark scores

In the bottom third on SWE-bench Verified (rank 25 of 30), and one of 70 whose vendor states it does not train on customer data.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Xiaomi Corporation · Family: MiMo V2.6

More about Xiaomi Corporation on HokAI

Input modalities: text, image · Output: text

About MiMo-V2.6-Distill-Qwen-9B

MiMo-V2.6-Distill-Qwen-9B is a 9-billion-parameter dense model Xiaomi's MiMo team released alongside the flagship MiMo-V2.6-Pro-RL and MiMo-V2.6-Flash-RL checkpoints on September 21, 2026, under the MIT license. Unlike its two RL-trained siblings, it is a supervised fine-tune of Alibaba's Qwen3.5-9B base model on MiMo-generated data, not a reinforcement-learning checkpoint in its own right. Xiaomi frames it as a starting point for open research into agentic reinforcement learning, covering coding, general-purpose agent tasks, visual coding, and cybersecurity.

Xiaomi's own evaluation table compares the SFT checkpoint against its Qwen3.5-9B base rather than competing 9B models. On SWE-bench Verified it improves from 60.0 to 61.1 (avg@3), and on SWE-bench Pro from 32.0 to 44.6. On Terminal-Bench 2.1 it rises from 27.0 to 37.1 (avg@1). Larger jumps appear on Xiaomi's own internal evaluation sets, which are not independently verifiable: an internal coding set from 19.5 to 51.6, and an internal general-agent set from 28.5 to 62.2.

As a Qwen3.5-9B derivative, the model inherits that architecture's context handling and tokenizer rather than Pro-RL's own sparse-MoE design; Xiaomi's model card does not publish a specific context window for this checkpoint the way it does for the two hosted flagships. It ships as BF16 safetensors and supports the same explicit thinking on/off toggle used across the MiMo-V2.6 family.

Input is text and image, listed under Hugging Face's image-text-to-text task and inherited from the Qwen3.5-9B base's own vision support. There is no evidence of video or audio input: unlike Pro-RL and Flash-RL, its Hugging Face tags do not list video-understanding or audio capability. Output is text only.

The supervised fine-tuning mixture totals 77.4 billion tokens, of which 27.2 billion are loss-bearing. By domain: code is 23.2 billion tokens (29.9%), cybersecurity 11.0 billion (14.2%), general agent tasks 22.0 billion (28.5%), and visual coding 21.2 billion (27.4%). This is distillation data Xiaomi's own larger MiMo models generated, not raw pretraining text.

There is no hosted API for this checkpoint; Xiaomi's pay-as-you-go product pages list only MiMo-V2.6-Pro and MiMo-V2.6-Flash. It is distributed purely as open weights on Hugging Face, with a documented SGLang quickstart, and broad downstream community activity: 20 finetunes, 7 merges, and 54 quantizations were already published by others at listing time, alongside 9,994 downloads in the preceding month.

As with the two flagship checkpoints, Xiaomi has not published a MiMo-V2.6-specific training-data summary or system card covering this distilled model. Because it trains on Xiaomi's own MiMo-generated data rather than raw web text, its behavior is shaped by whatever the teacher models produced, which is not independently documented.

MiMo-V2.6-Distill-Qwen-9B suits researchers and hobbyists who want a small, self-hostable, MIT-licensed agentic model to experiment with or fine-tune further, and teams building on Qwen3.5's ecosystem who want MiMo-style agentic behavior without the trillion-parameter Pro-RL checkpoint. It is a weaker choice for anyone needing a supported hosted API, a long context window, or capability close to the much larger RL-trained siblings.

The checkpoint sits inside Xiaomi's 5-item MiMo-V2.6 Hugging Face collection alongside Pro-RL and Flash-RL, and Xiaomi frames it as a base for others to run their own agentic reinforcement learning on top of, rather than a finished product. As with the rest of the family, Xiaomi has already previewed MiMo-V3 as the next generation in the line.

Pricing

No hosted API exists for this checkpoint as of this listing; Xiaomi's pay-as-you-go pricing pages list only MiMo-V2.6-Pro and MiMo-V2.6-Flash. Running MiMo-V2.6-Distill-Qwen-9B means self-hosting the open weights, so the real cost is whatever compute you provide, not a per-token vendor rate.

Key Features

  • 9B dense SFT of Qwen3.5-9B: Supervised fine-tune of Alibaba's Qwen3.5-9B on 77.4 billion tokens of MiMo-generated agentic data, released as an open research starting point rather than a flagship.
  • Covers four agentic domains: The SFT mixture splits across code (29.9%), general agent tasks (28.5%), visual coding (27.4%), and cybersecurity (14.2%).
  • Thinking on/off toggle: Supports the same enable_thinking chat-template switch as the rest of the MiMo-V2.6 family, separating reasoning_content from the final answer.
  • Vision-capable, text output: Handles image-text-to-text input, inherited from the Qwen3.5-9B base, but Xiaomi has not documented video or audio input for this checkpoint the way it has for Pro-RL and Flash-RL.
  • Active open-source community: A downstream community of finetunes, merges, and quantizations has already formed around the checkpoint on Hugging Face, unusually fast for a research-framed release.

Pros

  • Small enough (9B dense) to self-host on a single modern GPU, unlike the 1.02T-parameter Pro-RL.
  • MIT license with no commercial-use restrictions, plus an active downstream community already building finetunes and quantizations on Hugging Face.
  • Improves meaningfully over its base model on every benchmark Xiaomi reports, for example 37.1 vs 27.0 on Terminal-Bench 2.1.

Cons

  • No hosted API; Xiaomi's pay-as-you-go pricing pages cover only MiMo-V2.6-Pro and MiMo-V2.6-Flash.
  • A supervised fine-tune rather than a checkpoint trained with reinforcement learning, so its ceiling is well below the family's two flagships.
  • No published context window for this specific checkpoint, unlike the two hosted flagships.

Benchmarks

  • SWE-bench Pro: 44.6% vendor-reported · 28 Sep 2026 — Harder, longer real-repository coding tasks, % solved.
  • SWE-bench Verified: 61.1% vendor-reported · 28 Sep 2026 — Real GitHub issues fixed end to end, % solved.
  • Terminal-Bench 2.1: 37.1% vendor-reported · 28 Sep 2026 — Multi-step tasks completed in a real command line, % solved.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

Is MiMo-V2.6-Distill-Qwen-9B free to use?

Yes. It ships under the MIT license, which permits commercial use, and there is no hosted API or per-token charge: you download the open weights from Hugging Face and run them on your own infrastructure. The only cost is whatever compute you provide.

What separates MiMo-V2.6-Distill-Qwen-9B from its Qwen3.5-9B base model?

Fine-tuning on MiMo-generated agentic data lifts SWE-bench Verified from 60.0 to 61.1, SWE-bench Pro from 32.0 to 44.6, and Terminal-Bench 2.1 from 27.0 to 37.1, per Xiaomi's own reported figures. Qwen3.5-9B itself remains available separately from Alibaba as an untuned base.

Is MiMo-V2.6-Distill-Qwen-9B open source or proprietary?

It is open source under the MIT license, and the weights are published on Hugging Face. At 9 billion dense parameters, hardware requirements are far below the 1.02-trillion-parameter Pro-RL checkpoint in the same family, which needs an 8-16 GPU cluster to serve.

What happens to your data when you use MiMo-V2.6-Distill-Qwen-9B?

Nothing leaves your own infrastructure, because there is no hosted API: you download the weights and run inference yourself, so Xiaomi never sees prompts or outputs. That is a different posture from the hosted Pro-RL and Flash-RL checkpoints, which do have a live API.

What is MiMo-V2.6-Distill-Qwen-9B not good at?

It is a supervised fine-tune, not an RL-trained checkpoint, so it should not be expected to approach Pro-RL or Flash-RL on agentic benchmarks or the Artificial Analysis Intelligence Index. It also has no documented context window, no hosted API, and no confirmed video or audio input, unlike its two larger siblings.

More AI Models on HokAI

Visit MiMo-V2.6-Distill-Qwen-9B Official Page