MAI-Image-2.5 suits teams already working inside Microsoft 365 who want in-document image editing without installing a separate app, especially for work needing accurate in-image text, where Microsoft reports a 107-point Arena gain over the prior MAI-Image-2. Skip it if you need the single top-ranked Arena text-to-image model or a self-hosted, open-weight option.
MAI-Image-2.5 is Microsoft AI's diffusion model for text-to-image generation and localized image editing, released June 2, 2026. It holds Arena Elo scores near 1,269 for text-to-image and 1,401 for image editing, and ships as a standard SKU plus a faster MAI-Image-2.5-Flash variant.
Where it sits
- $15.50/M$ per 1M tokensBlended price (3:1)Lower is better#60 / 64peer median $1.70/Mvendor price, checked by HokAI
- --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
- --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
- --% correctGPQA DiamondHigher is better-- / 44peer median 88.3%per source, see benchmark scores
Pricier than 94% of the 64 GA models with a published price, and one of 65 whose vendor states it does not train on customer data.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Microsoft AI · Family: MAI-Image
More about Microsoft AI on HokAI
Context window: 32,000 tokens
Input modalities: text, image · Output: image
About MAI-Image-2.5
MAI-Image-2.5 is a diffusion-based text-to-image and image-editing model built by Microsoft AI, the in-house model group led by Mustafa Suleyman that Microsoft formed in November 2025 to reduce its reliance on OpenAI. Microsoft announced the model on June 2, 2026 at Build 2026, alongside a batch of seven new MAI models covering voice, transcription, coding and reasoning. MAI-Image-2.5 is the fifth release in Microsoft's image lineage, following MAI-Image-1 (October 2025), MAI-Image-2 (March 2026) and the cost-optimized MAI-Image-2-Efficient, and it ships in two SKUs: the full MAI-Image-2.5 and a faster, cheaper MAI-Image-2.5-Flash variant.
On the Arena leaderboard, the industry's crowd-voted comparison of image models, MAI-Image-2.5 placed third globally for text-to-image generation and second for image editing, trailing only GPT-Image-2 on the editing leaderboard. Microsoft has not published FID scores, CLIP scores, or head-to-head numbers against FLUX, Stable Diffusion 3 or Imagen 3/4, so those comparisons cannot be made with confidence.
The model accepts up to 32,000 tokens of text context per request and outputs images across seven aspect ratios. MAI-Image-2.5-Flash is Microsoft's first model in the line to handle both text-to-image and image-to-image workloads in one SKU, supporting localized edits such as replacing a single object, updating in-image text or removing motion blur without touching the rest of the frame. Third-party analysis estimates 10 billion to 50 billion non-embedding parameters, but Microsoft has not disclosed an official parameter count, so that figure should be read as an estimate rather than a confirmed spec.
Access runs primarily through Microsoft Foundry (formerly Azure AI Foundry) as the developer API, with additional surfaces via the MAI Playground, PowerPoint's built-in image tools, and third-party routers including OpenRouter and Fireworks AI. Pricing is token-based, with the Flash SKU priced well below the standard tier for high-volume batch work. Enterprise customers can also reserve provisioned throughput (PTU) capacity, though Microsoft has not published PTU rates publicly.
Microsoft's published model cards for the MAI-Image line describe a two-phase safety evaluation (pre-mitigation and post-mitigation) with prompt- and output-level filtering against violent or gory content, sexual content and nudity, depictions of real public figures, and reproduction of trademarked or protected material. The model cards recommend human review before using generated images in identity, medical, legal, financial or news contexts, rather than deploying the model autonomously in those workflows. Microsoft Ireland Operations Limited is listed as the EU regulatory contact on the model card. Whether MAI-Image-2.5's API output embeds C2PA content-provenance metadata specifically, versus Microsoft's broader February 2026 rollout of that same provenance standard across its productivity apps, is not spelled out in the public model card and should be treated as unconfirmed.
MAI-Image-2.5 works well for teams already inside the Microsoft ecosystem: PowerPoint and OneDrive users doing in-document image edits, and Azure customers who want commercial imagery and packaging mockups with reliable in-image text. It is a weaker fit for solo creators or hobbyists who want a one-click consumer app, since the primary access path runs through Azure Foundry's developer tooling rather than a standalone consumer interface comparable to Midjourney or DALL-E 3's ChatGPT integration; Bing Image Creator and Copilot rollout for the 2.5 generation was not confirmed as live at launch.
Microsoft has not published a training data cutoff or a detailed description of the training corpus for MAI-Image-2.5 beyond stating the model was trained with curated data selection and evaluation feedback from creative professionals, with content filtering mitigations applied before training. No architecture paper or technical report accompanies the release, and no GitHub repository exists for the MAI-Image line, consistent with Microsoft's closed-weights approach across the MAI family.
The release fits a clear strategic pattern: Microsoft restructured its OpenAI partnership in October 2025 to gain the contractual right to pursue frontier model development independently, then stood up the Microsoft AI Superintelligence team the following month. MAI-Image-2.5's Foundry pricing, positioned below licensing GPT-Image-2 at scale, makes it the image-generation leg of that in-house stack alongside MAI-Voice-2 for speech and MAI-Thinking-1 for reasoning.
Pricing
Microsoft Foundry bills MAI-Image-2.5 per token: $5 per 1M text-input tokens, $8 per 1M image-input tokens, $47 per 1M image-output tokens. The Flash SKU is cheaper: $1.75/1.75/$19.50 per 1M for text-input/image-input/image-output, about 41% below MAI-Image-2. Provisioned throughput (PTU) reservations exist for enterprise workloads but rates are not publicly disclosed.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.150 | $0.047 | $0.197 |
| Support reply | $0.010 | $0.014 | $0.024 |
| One coding agent run | $1.00 | $0.940 | $1.94 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Localized image editing: Replaces a single object, updates in-image text, or removes motion blur without altering the rest of the frame, ranked #2 on Arena's editing leaderboard.
- Improved text rendering: In-image typography scores 107 Arena points higher than MAI-Image-2, the largest single-category gain in the 2.5 release.
- Two-SKU lineup: The Flash SKU generates roughly 22% faster than the standard SKU, trading a small quality difference for speed and a lower per-token price.
- Seven aspect ratios: Natively supports 1:1, 4:3, 3:4, 16:9, 9:16, 3:2 and 2:3 output at up to 1024x1024 pixels.
- Native Microsoft 365 integration: Built into PowerPoint's image tools, with OneDrive integration rolling out, so no separate app is needed for in-document edits.
Pros
- Ranks #2 globally on Arena's image-editing leaderboard, trailing only GPT-Image-2.
- The Flash SKU meaningfully undercuts the standard SKU's per-token cost, making high-volume batch editing more affordable for cost-sensitive teams.
- Native PowerPoint access removes the friction of a separate app or account for existing Office users.
Cons
- Ranks #3, not #1, on Arena's text-to-image leaderboard, behind GPT-Image-2 and the leaderboard's top entry.
- No official parameter count or training data cutoff has been disclosed by Microsoft.
- Primary access is Azure/Microsoft Foundry developer tooling rather than a standalone consumer app.
Benchmarks
- LMArena Elo: 1269 vendor-reported · 02 Jun 2026 — Rating from blind human votes on which answer is better.
- LMArena rank: #3 vendor-reported · 02 Jun 2026 — Position on the blind human-preference leaderboard; #1 is best.
- Lmarena Elo Image Edit: 1401 vendor-reported · 02 Jun 2026
- Lmarena Rank Image Edit: #2 vendor-reported · 02 Jun 2026
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What are MAI-Image-2.5's pricing plans in 2026?
On Microsoft Foundry, standard MAI-Image-2.5 runs $5.00 per 1M text-input tokens, $8.00 per 1M image-input tokens and $47.00 per 1M image-output tokens. Choosing the Flash SKU instead drops those rates to $1.75, $1.75 and $19.50 per 1M respectively, about 41% cheaper on output. There is no flat per-image price; enterprise buyers can also negotiate provisioned throughput (PTU) capacity, though Microsoft keeps those rates private.
Is MAI-Image-2.5 free to use?
No, MAI-Image-2.5 has no free tier. Every request bills through Microsoft Foundry's per-token pricing on either the standard or Flash SKU, with no self-serve trial credit disclosed. PowerPoint users generate images inside their existing Office subscription, but Microsoft has not stated whether that usage is metered separately or included at no extra cost.
What are the best alternatives to MAI-Image-2.5?
GPT Image 2 is the closest rival on Arena, holding the top spot for image editing where MAI-Image-2.5 places second. Gemini 3 Pro Image (also marketed as Nano Banana Pro) is worth checking for native 4K output and Search-grounded factual accuracy. FLUX.2 from Black Forest Labs leads on photorealism and exact color matching when brand-accurate rendering matters more than in-image text.
MAI-Image-2.5 or GPT Image 2: which should you pick?
GPT Image 2 holds the #1 spot on Arena's image-editing leaderboard, with MAI-Image-2.5 close behind at #2, and GPT Image 2 also outranks MAI-Image-2.5 for text-to-image. Go with MAI-Image-2.5 when a PowerPoint-based workflow and lower per-token editing costs outweigh topping the leaderboard; go with GPT Image 2 when the strongest all-around Arena ranking matters more than platform convenience.
What does it take to start using MAI-Image-2.5?
Developers sign up for Microsoft Foundry (formerly Azure AI Foundry) and call the API with an Azure API key or Entra ID credential. Non-developers can try the model through the MAI Playground at playground.microsoft.ai, or generate images directly inside PowerPoint with no separate setup. Third-party routers OpenRouter and Fireworks AI also carry the model for teams already using those platforms.
Top Alternatives
- GPT Image 2: Pick GPT Image 2 for the top Arena image-editing rank and a stronger text-to-image placement; pick MAI-Image-2.5 for native PowerPoint access and lower per-token pricing.
- Gemini 3 Pro Image: Pick Gemini 3 Pro Image for native 4K output and Search-grounded factual accuracy; pick MAI-Image-2.5 for its Office-embedded workflow and Flash-tier pricing.
- FLUX.2: Pick FLUX.2 for photorealism and exact hex color matching; pick MAI-Image-2.5 for its #2 Arena editing rank and built-in PowerPoint access.