MiniMax released M3.1-Flash-Preview on September 27, 2026, as a coding-focused model inside its MiniMax Code agent tool, letting callers dial reasoning depth across 5 distinct effort settings. It targets everyday bug fixes and feature work but stays gated to MiniMax's own subscription plan rather than a standalone metered API key.
MiniMax-M3.1-Flash-Preview is a coding- and agent-focused language model MiniMax launched on September 27, 2026, with a 1,000,000-token context window and five tunable reasoning-effort levels from low to max. It handles multimodal text, image, and video input but ships without published benchmarks or per-token pricing.
Provider: MiniMax · Family: MiniMax M3
Context window: 1,000,000 tokens
Input modalities: text, image, video, tool-calls · Output: text, tool-calls
About MiniMax-M3.1-Flash-Preview
MiniMax-M3.1-Flash-Preview is a text-generation model from MiniMax, the Shanghai AI lab listed on the Hong Kong Stock Exchange (0100.HK). MiniMax added it to its API documentation and MiniMax Code agent tool on September 27, 2026, positioning it inside the M-series lineage alongside MiniMax-M3 and MiniMax-M2.7. Unlike those siblings, M3.1-Flash-Preview launched without a dedicated marketing page, model card, or benchmark report; MiniMax's own release-notes changelog had no entry for it as of this writing, and independent tech press (Pandaily, AlphaSignal, Startup Fortune) described the rollout as quiet, with one outlet noting it shipped "without a price tag."
The model is built for agentic reasoning, tool use, coding, and long-context tasks, per MiniMax's official API documentation. It supports a context window of up to 1,000,000 tokens for long documents, codebases, and multi-step agent sessions, and accepts multimodal chat input across text, images, and video. MiniMax has not disclosed a parameter count, active-expert count, or architecture family (dense vs. Mixture-of-Experts) specifically for this variant, in contrast to sibling MiniMax-M3, documented separately as a 229.9B-parameter Mixture-of-Experts model with MiniMax Sparse Attention.
The defining mechanical feature is tunable reasoning depth. The effort parameter (output_config.effort on the Anthropic-compatible endpoint, reasoning_effort on the OpenAI-compatible endpoint) accepts low, medium, high, xhigh, and max, with higher levels producing more thorough reasoning at higher latency and token cost. Omitting the parameter defaults to max, the deepest and slowest setting. Reasoning cannot be switched off: sending thinking:{type:"disabled"} or effort:"none" returns an HTTP 400 error stating the model requires adaptive thinking.
No official benchmark suite has been published for M3.1-Flash-Preview. It does not appear on Artificial Analysis's index, and no responsible SWE-bench, GPQA, or Terminal-Bench figure exists for it from MiniMax itself, unlike published rivals such as DeepSeek V4 and DeepSeek V4 Flash. Scattered third-party community testing exists, but none of it is independently verified to a standard this page can responsibly repeat, so no benchmark scores are listed here; treat any vendor or blog claim of coding-quality gains as provisional until MiniMax releases formal evaluation numbers.
Access runs through two paths, and neither is a conventional metered API key at launch. MiniMax's Token Plan is a flat monthly subscription, tiered by usage volume (exact tier pricing is in the pricing FAQ below), covering M3.1-Flash-Preview alongside the rest of the M-series and MiniMax's image, speech, and video models, billed against 5-hour rolling and weekly usage quotas rather than per-token spend. The model is also usable directly inside MiniMax Code, MiniMax's own coding agent product. The model is entirely absent from MiniMax's public pay-as-you-go pricing table, which lists exact per-token rates for MiniMax-M3, M2.7, M2.7-highspeed, and the legacy M2.x line but nothing for M3.1-Flash-Preview. MiniMax's Token Plan documentation notes that once a subscription quota is exhausted, usage can fail over to a standard pay-as-you-go API key, implying a metered price exists internally, but no such rate has been published for this model as of this writing.
MiniMax aims the model squarely at everyday engineering work, bug fixes, refactors, and end-to-end feature implementation, rather than research-grade reasoning benchmarks, and documents integration guides for reaching it from several popular coding-agent tools, including Claude Code and Cursor (the full list is under Key Features below). Two request formats are documented, letting teams already wired to a major model SDK point their base URL at MiniMax and swap in the model name rather than adopt a new client library.
MiniMax has not published a system card, named red-teaming partners, or a responsible scaling policy for the M-series generally, and no model-specific safety disclosure accompanies this preview. It also has not published SOC 2, ISO 27001, or HIPAA attestations for its platform as a whole. How MiniMax handles API inputs for training is covered in the FAQ below.
Overall, M3.1-Flash-Preview reads as MiniMax testing a faster, subscription-gated coding model ahead of a fuller rollout: the underlying capability, a huge context window and tunable reasoning, is real and documented, but the commercial and evaluation groundwork other frontier launches ship with, pricing, benchmarks, and a system card, has not caught up yet. The FAQ below covers exactly who it fits today.
Pricing
No metered pay-as-you-go rate is published for this model; it sits outside MiniMax's per-token pricing table entirely, unlike sibling models MiniMax-M3 and M2.7 (verified 2026-09-28). The pricing FAQ below covers the Token Plan subscription tiers and MiniMax Code access that currently gate it.
Key Features
- 1M-Token Context Window: Handles up to 1,000,000 tokens of input for long codebases, documents, and multi-step agent sessions, per MiniMax's official API documentation.
- Five-Level Tunable Reasoning Effort: The effort parameter accepts low, medium, high, xhigh, and max, trading thinking depth for latency and token cost; omitting it defaults to the deepest max setting.
- Multimodal Chat Input: Accepts text, image, and video input in the same request for content understanding alongside coding tasks.
- Anthropic- and OpenAI-Compatible Endpoints: Reachable through an Anthropic-compatible endpoint (api.minimax.io/anthropic) and an OpenAI-compatible endpoint (api.minimax.io/v1), so existing SDK integrations can point at it with a base URL and model-name change.
- Built for Coding Agent Tools: Ships with integration guides for Claude Code, Cursor, Codex, TRAE, Hermes Agent, and OpenClaw under the Token Plan, aimed at everyday bug fixes, refactors, and feature implementation.
Pros
- Context window large enough for full codebases and multi-step agent sessions, matching frontier-class long-context models.
- Five explicit reasoning-effort levels (low to max) give fine control over latency and cost per request.
- Drop-in compatible with both Anthropic and OpenAI SDK request formats, easing adoption for teams already using Claude Code or GPT-based coding tools.
Cons
- No published per-token pay-as-you-go price at launch; access is gated to a paid monthly subscription or MiniMax Code, not a standalone metered API key.
- No official benchmark suite, model card, or system card released alongside launch.
- Thinking cannot be disabled, and omitting the effort parameter silently defaults to the slowest, most expensive max setting.
Frequently Asked Questions
How much does MiniMax-M3.1-Flash-Preview cost per 1M tokens?
MiniMax has not listed a per-token rate for this model anywhere in its published API pricing, unlike MiniMax-M3, M2.7, and the legacy M2.x line, which all carry exact input and output prices. Use instead requires a Token Plan subscription, Plus at $22/month, Max at $55/month, or Ultra at $132/month, each metered on 5-hour rolling and weekly quotas, or free access inside the MiniMax Code agent tool.
How does MiniMax-M3.1-Flash-Preview compare on benchmarks to rival coding models?
MiniMax has not released an official benchmark suite or model card for M3.1-Flash-Preview, so no verified head-to-head score exists yet against rivals like DeepSeek V4 Flash or Gemini 3 Flash Lite Preview, both of which shipped with published evaluation tables. Treat any vendor or third-party comparison as provisional until MiniMax publishes formal numbers.
Is MiniMax-M3.1-Flash-Preview open source or proprietary?
It is proprietary and API-only: MiniMax has not released model weights for this preview, unlike its sibling MiniMax-M3, which shipped under an open-weights modified-MIT license on Hugging Face. Access requires a Token Plan Subscription Key or use inside MiniMax Code, with no Hugging Face or GitHub weight release announced.
Does MiniMax-M3.1-Flash-Preview train on user data?
MiniMax's platform-wide policy states that customer API inputs are not stored or used for model training by default unless a customer explicitly opts in, per its Privacy Policy and Terms of Service. No separate retention or training disclosure specific to M3.1-Flash-Preview has been published.
Who is MiniMax-M3.1-Flash-Preview best for, and who should avoid it?
It suits teams already inside MiniMax Code or willing to pay for a Token Plan subscription who want a fast, coding-focused model with a 1M-token context and tunable reasoning depth for everyday bug fixes and feature work. Teams needing a standalone metered API key, verified benchmark scores, or open weights for self-hosting should choose a sibling model like MiniMax-M3 instead.