Qwen3.8-Max launched GA on August 3, 2026 with 2.4 trillion total parameters (about 95 billion active) and a 1-million-token context window across a single flat pricing tier. It targets teams building large-context, multimodal agentic workloads on Alibaba Cloud who can tolerate its still-unpublished safety and training model card.
Qwen3.8-Max is Alibaba Cloud's 2.4-trillion-parameter Mixture-of-Experts model, launched GA on August 3, 2026, with a 1-million-token context window and native text, image and video input in a single call. It scores 93.0 on PaperBench, ahead of GPT-5.6 Sol, Fable 5 and Opus 4.8.
Provider: Alibaba Cloud · Family: Qwen3.8
More about Alibaba Cloud on HokAI
Context window: 1,000,000 tokens · Max output: 131,072
Input modalities: text, image, video · Output: text
About Qwen3.8-Max
Qwen3.8-Max is Alibaba Cloud's flagship large language model, built by the Qwen team and released as a preview on July 19, 2026 at the World AI Conference in Shanghai, then taken to general availability on August 3, 2026. It is a sparse Mixture-of-Experts model with 2.4 trillion total parameters and roughly 95 billion active parameters per token, built on the Qwen3.5 architecture with a hybrid attention mechanism. The model is designed as Alibaba's answer to closed frontier labs like OpenAI and Anthropic, aiming to compete on both reasoning benchmarks and raw context capacity while undercutting them on price. On benchmarks, Alibaba published a launch table covering Terminal-Bench 2.1, SWE-bench Pro, PaperBench, IFBench and Humanity's Last Exam, rather than the more commonly cited SWE-bench Verified or GPQA Diamond. Qwen3.8-Max scores 93.0 on PaperBench, ahead of GPT-5.6 Sol (90.5), Claude Fable 5 (88.8) and Claude Opus 4.8 (80.3). On IFBench it posts 82.8 against GPT-5.6 Sol's 72.7. It trails the frontier leaders on harder reasoning and coding evaluations: 43.6 on Humanity's Last Exam versus Fable 5's 53.3, and 67.7 on SWE-bench Pro versus Fable 5's 80.0. On Terminal-Bench 2.1 it scores 86.6, ahead of both Opus 4.8 and Fable 5 (84.6 each) but behind GPT-5.6 Sol's 88.8. Arena.AI's blind pairwise human-preference leaderboard placed it second globally on multimodal tasks, behind only Claude Fable 5. The model ships with a 1-million-token context window: up to 991,800 input tokens in non-thinking mode, 983,610 in thinking mode, plus 131,072 output tokens. It is a native multimodal foundation, accepting text, image and video input in a single call, though output is text-only at launch. Preserved thinking mode is enabled by default; multi-turn clients must return the complete, unmodified reasoning_content history, and that preserved reasoning counts toward input-token billing. The model supports function calling, structured output, tool use and built-in tools with three cache modes for cost control. Pricing is a single flat tier covering the entire context window: $2.00 per 1M input tokens and $6.00 per 1M output tokens through Alibaba Cloud Model Studio, undercutting most frontier-tier competitors including Alibaba's own prior flagship. New Model Studio activations get a 1M-token free quota in the Singapore (international) region. The model is reachable via Alibaba Cloud Model Studio APIs, QwenWork (Alibaba's enterprise workplace agent platform), and OpenRouter; there is no confirmed listing yet on AWS Bedrock, Google Vertex AI, Together AI or Fireworks. Alibaba has not published a training or safety model card for Qwen3.8-Max as of the GA launch: no activated-parameter breakdown, no reproducible evaluation configuration, no disclosed training data cutoff, and no red-team partner list. Open weights were promised to ship "next week" at the August 3, 2026 announcement, alongside a smaller Qwen3.8-27B checkpoint, but no Hugging Face model card existed at launch and no license terms have been named, so the model is proprietary and API-only for now. Coverage from Tech Times also flags that QwenWork's enterprise workflow integration routes data through Alibaba Cloud's China-linked infrastructure, which raises data-residency questions for regulated Western enterprises that should be checked directly in Model Studio's region settings. Qwen3.8-Max fits teams that need very large context windows and native multimodal input at low per-token cost, particularly ones already building on Alibaba Cloud, or workloads (like PaperBench-style research-paper tasks and instruction-following) where it already leads. It is a weaker choice for teams that need the strongest possible agentic coding or graduate-level reasoning scores, or that require a published safety/training card before production deployment. Teams evaluating it against GPT-5.6 Sol, Claude Fable 5 or Claude Opus 4.8 should weigh its price and context advantage against its gaps on Humanity's Last Exam and SWE-bench Pro.
Pricing
$2.00 per 1M input tokens and $6.00 per 1M output tokens, one flat tier covering the full 1M-token context window (0 to 1M tokens, no tiered step-up). Preserved reasoning_content in thinking mode counts as input tokens and is billed accordingly. New Model Studio activations get a 1M-token free quota in the Singapore region.
Key Features
- 1M-Token Context Window: Up to 991.8K input tokens in non-thinking mode (983.6K in thinking mode) plus 131K output tokens, in one flat pricing tier across the whole window.
- Native Text, Image and Video Input: A single multimodal call accepts text, image and video, ranked second globally on Arena.AI's multimodal leaderboard behind only Fable 5.
- Sparse MoE with Hybrid Attention: 2.4 trillion total parameters with roughly 95 billion active per token, built on the Qwen3.5 architecture.
- Preserved Thinking Mode: Visible reasoning_content carried across multi-turn calls by default, giving inspectable chain-of-thought at the cost of extra billed input tokens.
- Built-In Tool Calling and Caching: Native function calling, structured output, built-in tools, and three cache modes for controlling repeat-context cost.
Pros
- Leads PaperBench at 93.0, ahead of GPT-5.6 Sol, Fable 5 and Opus 4.8.
- 1-million-token context window in a single flat $2/$6 per-1M-token pricing tier.
- Native text, image and video input, second globally on Arena.AI's multimodal leaderboard behind only Fable 5.
Cons
- No published safety or training model card as of the 2026-08-03 GA launch.
- Trails Fable 5 on Humanity's Last Exam (43.6 vs 53.3) and SWE-bench Pro (67.7 vs 80.0).
- Open weights promised at launch but not yet shipped; no confirmed listing on Bedrock, Vertex, Together or Fireworks.
Benchmarks
- ifbench: 82.8
- paperbench: 93
- lmarena rank: 2
- swe bench pro: 67.7
- terminal bench 2 1: 86.6
- humanitys last exam: 43.6
Frequently Asked Questions
How much does Qwen3.8-Max cost per 1M tokens?
Qwen3.8-Max costs $2.00 per 1M input tokens and $6.00 per 1M output tokens on Alibaba Cloud Model Studio, a single flat tier that covers the entire 1M-token context window. That undercuts GPT-5.6 Sol and Claude Opus 4.8 on list price. New Model Studio activations also get a 1M-token free quota in the Singapore region.
How does Qwen3.8-Max compare on benchmarks vs GPT-5.6 Sol?
Qwen3.8-Max leads GPT-5.6 Sol on PaperBench (93.0 vs 90.5) and IFBench (82.8 vs 72.7), but trails on Terminal-Bench 2.1 (86.6 vs 88.8). Alibaba's published table doesn't include a head-to-head on SWE-bench Verified or GPQA Diamond, so coding and PhD-level reasoning parity is unverified.
Is Qwen3.8-Max open source or proprietary?
Qwen3.8-Max is proprietary and API-only at its August 2026 GA launch. Alibaba promised open weights would ship the following week, alongside a smaller Qwen3.8-27B checkpoint, but no Hugging Face model card or license terms existed at launch, so self-hosting isn't yet possible.
Does Qwen3.8-Max train on user data?
Alibaba has not published a specific data-training or retention policy for Qwen3.8-Max. Confirm current retention and training opt-out settings directly in the Alibaba Cloud Model Studio console before sending sensitive or regulated data.
Who is Qwen3.8-Max best for and who should avoid it?
It suits teams that need a 1M-token context window and native multimodal input at low cost, especially on Alibaba Cloud. Teams needing the strongest agentic coding or graduate-level reasoning scores, or a published safety card before deployment, should look at GPT-5.6 Sol or Claude Fable 5 instead.