Qwen3.8-Max launched GA on August 3, 2026 as a large-scale Mixture-of-Experts model with a 1-million-token context window across a single flat pricing tier. Alibaba later open-weighted a smaller, text-only variant under a restricted custom license, though the hosted multimodal API remains proprietary and still lacks a published safety model card.
Alibaba Cloud's Qwen3.8-Max combines a 1-million-token context window with multimodal input, accepting images and video alongside text in a single API call, at a flat $2/$6 per-million-token rate that undercuts most frontier rivals. It launched to general availability on August 3, 2026, still without a published safety or training model card.
Where it sits
- $3.00/M$ per 1M tokensBlended price (3:1)Lower is better#37 / 64peer median $1.70/Mvendor price, checked by HokAI
- 41 tok/stokens/sOutput speedHigher is better#34 / 39peer median 90 tok/scited: Artificial Analysis
- --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
- 92.6%% correctGPQA DiamondHigher is better#11 / 44peer median 88.3%per source, see benchmark scores
Priced around the middle of the 64 GA models with a published price (rank 37), and in the top third on GPQA Diamond (rank 11 of 44).
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Alibaba Cloud · Family: Qwen3.8
More about Alibaba Cloud on HokAI
Context window: 1,000,000 tokens · Max output: 131,072
Input modalities: text, image, video · Output: text
About Qwen3.8-Max
Qwen3.8-Max is Alibaba Cloud's flagship large language model, built by the Qwen team and released as a preview on July 19, 2026 at the World AI Conference in Shanghai, then taken to general availability on August 3, 2026. It is a sparse Mixture-of-Experts model with 2.4 trillion total parameters and roughly 95 billion active parameters per token, built on the Qwen3.5 architecture with a hybrid attention mechanism. The model is designed as Alibaba's answer to closed frontier labs like OpenAI and Anthropic, aiming to compete on both reasoning benchmarks and raw context capacity while undercutting them on price.
On benchmarks, Alibaba published a launch table covering Terminal-Bench 2.1, SWE-bench Pro, PaperBench, IFBench, GPQA Diamond and Humanity's Last Exam, rather than the more commonly cited SWE-bench Verified. Qwen3.8-Max scores 93.0 on PaperBench, ahead of GPT-5.6 Sol (90.5), Claude Fable 5 (88.8) and Claude Opus 4.8 (80.3), and 92.6 on GPQA Diamond, level with Claude Fable 5 and just behind GPT-5.6 Sol. On IFBench it posts 82.8 against GPT-5.6 Sol's 72.7. It trails the frontier leaders on harder reasoning and coding evaluations: 43.6 on Humanity's Last Exam versus Fable 5's 53.3, and 67.7 on SWE-bench Pro versus Fable 5's 80.0. On Terminal-Bench 2.1 it scores 86.6, ahead of both Opus 4.8 and Fable 5 (84.6 each) but behind GPT-5.6 Sol's 88.8. Arena.AI's blind pairwise human-preference leaderboard placed it second globally on multimodal tasks, behind only Claude Fable 5.
The model ships with a 1-million-token context window: up to 991,800 input tokens in non-thinking mode, 983,610 in thinking mode, plus 131,072 output tokens. It is a native multimodal foundation, accepting text, image and video input in a single call, though output is text-only at launch. Preserved thinking mode is enabled by default; multi-turn clients must return the complete, unmodified reasoning_content history, and that preserved reasoning counts toward input-token billing. The model supports function calling, structured output, tool use and built-in tools with three cache modes for cost control.
Pricing is a single flat tier covering the entire context window, through Alibaba Cloud Model Studio, undercutting most frontier-tier competitors including Alibaba's own prior flagship (see the pricing table for the exact per-token rates and free-quota terms). The model is reachable via Alibaba Cloud Model Studio APIs, QwenWork (Alibaba's enterprise workplace agent platform), and OpenRouter; there is no confirmed listing yet on AWS Bedrock, Google Vertex AI, Together AI or Fireworks.
Alibaba still has not published a training or safety model card for Qwen3.8-Max: no activated-parameter breakdown, no reproducible evaluation configuration, no disclosed training data cutoff, and no red-team partner list, as of this review. Alibaba did follow through on the open-weights promise made at the August 3, 2026 GA launch: on August 13, 2026 it released Qwen3.8-2.4T-A95B on Hugging Face under a custom "qwen3.8-max" license, alongside the smaller Qwen3.8-27B under Apache 2.0. The open-weight Max-class checkpoint is text-only with a 262,144-token native context, well short of the hosted API's 1-million-token multimodal window, so self-hosters should not expect parity with the Model Studio product. Coverage from Tech Times also flags that QwenWork's enterprise workflow integration routes data through Alibaba Cloud's China-linked infrastructure, which raises data-residency questions for regulated Western enterprises that should be checked directly in Model Studio's region settings.
Qwen3.8-Max fits teams that need very large context windows and native multimodal input at low per-token cost, particularly ones already building on Alibaba Cloud, or workloads (like PaperBench-style research-paper tasks and instruction-following) where it already leads. It is a weaker choice for teams that need the strongest possible agentic coding or graduate-level reasoning scores, or that require a published safety/training card before production deployment. Teams evaluating it against GPT-5.6 Sol, Claude Fable 5 or Claude Opus 4.8 should weigh its price and context advantage against its gaps on Humanity's Last Exam and SWE-bench Pro.
Pricing
$2.00 per 1M input tokens and $6.00 per 1M output tokens, one flat tier covering the full 1M-token context window (0 to 1M tokens, no tiered step-up). Preserved reasoning_content in thinking mode counts as input tokens and is billed accordingly. New Model Studio activations get a 1M-token free quota in the Singapore region.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.060 | $0.0060 | $0.066 |
| Support reply | $0.0040 | $0.0018 | $0.0058 |
| One coding agent run | $0.400 | $0.120 | $0.520 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- 1M-Token Context Window: Up to 991.8K input tokens in non-thinking mode (983.6K in thinking mode) plus 131K output tokens, in one flat pricing tier across the whole window.
- Native Text, Image and Video Input: A single multimodal call accepts text, image and video, ranked second globally on Arena.AI's multimodal leaderboard behind only Fable 5.
- Sparse MoE with Hybrid Attention: 2.4 trillion total parameters with roughly 95 billion active per token, built on the Qwen3.5 architecture.
- Preserved Thinking Mode: Visible reasoning_content carried across multi-turn calls by default, giving inspectable chain-of-thought at the cost of extra billed input tokens.
- Built-In Tool Calling and Caching: Native function calling, structured output, built-in tools, and three cache modes for controlling repeat-context cost.
Pros
- Leads PaperBench outright, ahead of Claude Fable 5 and Claude Opus 4.8.
- 1-million-token context window in a single flat $2/$6 per-1M-token pricing tier.
- Multimodal input (text, image, video) in one call is a genuine edge over text-only rivals at this price point.
Cons
- No published safety or training model card as of the 2026-08-03 GA launch.
- Trails Fable 5 on Humanity's Last Exam (43.6 vs 53.3) and SWE-bench Pro (67.7 vs 80.0).
- Open-weight checkpoint (Aug 13, 2026) is text-only with a smaller context window than the hosted API; still no confirmed listing on Bedrock, Vertex, Together or Fireworks.
Benchmarks
- IFBench: 82.8% independent · 14 Sep 2026 — How precisely the model follows detailed instructions, % passing.
- Paperbench: 93 independent · 14 Sep 2026
- GPQA Diamond: 92.6% independent · 14 Sep 2026 — PhD-level science questions that are hard to search for, % correct.
- LMArena rank: #2 independent · 14 Sep 2026 — Position on the blind human-preference leaderboard; #1 is best.
- SWE-bench Pro: 67.7% independent · 14 Sep 2026 — Harder, longer real-repository coding tasks, % solved.
- Terminal-Bench 2.1: 86.6% independent · 14 Sep 2026 — Multi-step tasks completed in a real command line, % solved.
- Humanity's Last Exam: 43.6% independent · 14 Sep 2026 — Expert-written questions across many fields, % correct.
- AA Intelligence Index: 40 cited: Artificial Analysis · 14 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- AA blended price: $1.18/M cited: Artificial Analysis · 14 Sep 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
- Output speed: 41 tok/s cited: Artificial Analysis · 14 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does Qwen3.8-Max actually cost?
Qwen3.8-Max is priced at a single flat rate: $2.00 per million input tokens, $6.00 per million output tokens, covering the entire 1M-token context window with no tiered step-up. That undercuts GPT-5.6 Sol and most other frontier-tier rivals on list price. New Model Studio activations also get a limited free quota before billing kicks in.
Can you use Qwen3.8-Max without paying?
Not indefinitely: Qwen3.8-Max has no permanent free tier, but new Alibaba Cloud Model Studio activations get a one-time 1-million-token free quota in the Singapore (international) region. After that quota is used, billing switches to the standard flat per-token rate described above.
What are Qwen3.8-Max's closest competitors?
GPT-5.6 Sol and Claude Fable 5 are the closest frontier-tier rivals Alibaba benchmarked Qwen3.8-Max against directly. For a cheaper large-context alternative, DeepSeek's V4.1-Flash targets a similar low-cost, high-context niche built around undercutting Western frontier pricing.
How does Qwen3.8-Max compare to GPT-5.6 Sol in 2026?
Qwen3.8-Max leads GPT-5.6 Sol on PaperBench (93.0 vs 90.5) and IFBench (82.8 vs 72.7), but trails on Terminal-Bench 2.1 (86.6 vs 88.8) and GPQA Diamond. Alibaba's published table doesn't include a head-to-head on SWE-bench Verified, so coding parity there is unverified.
Are Qwen3.8-Max's weights open source?
Partially. Alibaba shipped an open-weight checkpoint on Hugging Face in mid-August 2026 under a restricted custom license that has not been confirmed as OSI-approved. That checkpoint is text-only with a smaller native context than the hosted API's full multimodal window, so it is not a drop-in replacement for Model Studio. The hosted API itself remains a closed, proprietary service.
HokAI guides covering Qwen3.8-Max
- What Happened When 7 AI Agents Got Real Bank Accounts and No Supervision: Bottleneck Labs gave 7 AI models real money and 72 unsupervised hours. Zero revenue, $12,431 in fake invoices sent, and what it actually means for agent safety.