Alibaba open-weighted this checkpoint in August 2026, shortly after Qwen3.8-Max's API debut, extending context to roughly 1 million tokens with FP8, GGUF and NVFP4 builds available the same day. It replaces the prior generation as Alibaba's flagship reasoning and agentic-coding model for teams that need self-hosting or multi-cloud deployment instead of a single vendor API.
Qwen3.8-2.4T-A95B is Alibaba Cloud's mixture-of-experts flagship, activating 95 billion of its 2.4 trillion parameters per token. It is the first Qwen-Max-class model Alibaba has released as open weights, aimed at agentic coding and long-horizon reasoning workloads that previously required a closed API.
Where it sits
- $3.00/M$ per 1M tokensBlended price (3:1)Lower is better#37 / 64peer median $1.70/Mvendor price, checked by HokAI
- 38 tok/stokens/sOutput speedHigher is better#36 / 39peer median 90 tok/scited: Artificial Analysis
- --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
- 92.6%% correctGPQA DiamondHigher is better#11 / 44peer median 88.3%per source, see benchmark scores
Priced around the middle of the 64 GA models with a published price (rank 37), and in the top third on GPQA Diamond (rank 11 of 44).
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Alibaba Cloud · Family: Qwen3.8
More about Alibaba Cloud on HokAI
Context window: 1,010,000 tokens · Max output: 131,072
Input modalities: text, tool-calls · Output: text, tool-calls
About Qwen3.8-2.4T-A95B
Qwen3.8-2.4T-A95B is Alibaba Cloud's largest foundation model release to date, open-weighted on Hugging Face on August 12, 2026. It is the checkpoint behind the QwenCloud-hosted Qwen3.8-Max product, GA'd on August 3, 2026. The architecture is a fine-grained mixture-of-experts transformer: 2.4 trillion total parameters, 95 billion active per token, across 512 experts (10 routed plus 1 shared per token) and 92 layers interleaving linear and full attention in a 3:1 pattern. It succeeds Qwen3.7-Max, announced at the Apsara Summit on May 20, 2026, and is the first Qwen-Max-class flagship Alibaba has open-weighted rather than kept API-only.
Alibaba reports 92.6% on GPQA Diamond, 67.7% on SWE-bench Pro, 86.6% on Terminal-Bench 2.1, 86.1% on OSWorld-Verified, 93.0% on PaperBench and 82.8% on IFBench; no SWE-bench Verified score has been published for this checkpoint by Alibaba, Artificial Analysis or any independent tracker as of September 2026. On the independently tracked DeepSWE 1.1 coding benchmark it scores 56.6, trailing several frontier agentic-coding competitors despite the strong GPQA result. Artificial Analysis puts its Intelligence Index at a comparatively modest 40 and its blended price at $1.18 per 1M tokens (7:2:1 cache/input/output ratio), an independent cross-check on Alibaba's own figures. AIME 2025, MMLU-Pro and ARC-AGI 2 remain unpublished, and independent verification is limited given the model's age.
Native context is 262,144 tokens, extensible to roughly 1,010,000, which is where the "1 million token context" figure in most coverage comes from. Recommended max output is 131,072 tokens for a final answer, though up to 262,144 of that budget can go to visible reasoning traces first, since thinking mode is mandatory and cannot be disabled.
Despite the "Max" branding, Qwen3.8-2.4T-A95B is text-only: its Hugging Face model card confirms multimodal input is not supported. It accepts text and tool-call input and returns text and tool-calls, with native OpenAI-compatible function calling, structured output and parallel tool calls. This mirrors Qwen3.7-Max, whose text-only "Max" flagship was often confused with the multimodal "Plus" variant; several writeups already describe Qwen3.8-Max as accepting image and video input, which this checkpoint does not. It ships configurable reasoning: low, high and xhigh effort levels per request.
QwenCloud's International (Singapore) endpoint undercuts Qwen3.7-Max's prior per-1M-token rate despite the larger size, with a steep discount on cached input (exact figures in the pricing FAQ below). The Mainland China endpoint runs an estimated 60-70% cheaper, and OpenRouter's routed price runs lower still. A 1M-in/200K-out coding-agent run costs about $3.20; a 1,000-turn support chat averaging 2K in/500 out per turn runs about $7.00.
Open weights ship on Hugging Face as safetensors, FP8, GGUF (Unsloth) and NVFP4 (RadixArk) builds. Day-0 serving came from SGLang, vLLM, NVIDIA NIM/Dynamo and Fireworks AI, alongside the official QwenCloud API and Alibaba Model Studio. NVIDIA reports over 4,000 tokens/sec per GPU and 350 tokens/sec per user on its GB300 NVL72 reference system, a figure specific to that deployment rather than any hosted API's guaranteed rate. No AWS Bedrock, Google Vertex or Azure availability is confirmed yet.
Alibaba has not published a safety system card, training data cutoff, or named red-teaming partners for this model. Its custom license, not Apache 2.0, requires products over 100 million monthly active users to display the model name in their UI, and "Model as a Service" or "AI Work Assistant" businesses over $50 million in trailing 12-month revenue to sign a separate commercial license before commercial use.
It suits teams building self-hosted or multi-cloud agentic coding and long-document pipelines who want frontier-scale reasoning without single-vendor lock-in, and who have the multi-GPU infrastructure a 2.4T-parameter checkpoint needs even at FP8 or NVFP4. Cost-sensitive teams often weigh it against DeepSeek V4, Kimi K3 and Gemini 3.7 Flash for the same long-context reasoning workloads. Teams needing native image or video understanding should use Qwen's separate VL line, or OpenAI's multimodal models, instead; teams needing a published system card for compliance may prefer GPT-5.6 Sol or Anthropic's Claude Opus 5, which document red-teaming partners this model lacks. Browse the rest of the large language model category for open- and closed-weight alternatives.
Screenshots

Pricing
QwenCloud's International (Singapore) endpoint prices Qwen3.8-2.4T-A95B at $2.00 per 1M input tokens and $6.00 per 1M output tokens, with cached input at $0.25 per 1M (an 8x discount). The Mainland China (Beijing) endpoint runs an estimated 60-70% cheaper for the same model. OpenRouter's routed pricing has floated slightly lower, around $1.80/$5.40 per 1M, depending on which upstream host serves the request. No separate batch-API discount has been published for this model.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.060 | $0.0060 | $0.066 |
| Support reply | $0.0040 | $0.0018 | $0.0058 |
| One coding agent run | $0.400 | $0.120 | $0.520 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Configurable Reasoning Effort: Switches between low, high and xhigh reasoning modes per request, trading inference cost for depth on hard reasoning and coding tasks.
- 1M-Class Context Window: Native 262,144-token context extends to roughly 1,010,000 tokens, enough for large codebases or multi-session agent transcripts in one call.
- Open Weights in Four Formats: Released on Hugging Face as full safetensors, FP8, GGUF (via Unsloth) and NVFP4 (via RadixArk) checkpoints for self-hosted deployment.
- Fine-Grained MoE Routing: 512 total experts with 10 routed plus 1 shared expert active per token, keeping inference cost tied to the 95B active count rather than the full 2.4T parameters.
- Day-0 Inference Stack Support: Four independent serving stacks added same-day support rather than waiting on community ports, with NVIDIA's reference hardware reporting over 4,000 tokens/sec per GPU.
Pros
- First Qwen-Max-class model Alibaba has open-weighted, priced below its own predecessor's rate on the official API (see the pricing FAQ for the exact numbers).
- Scores among the highest reported GPQA Diamond results in Alibaba's Qwen3.8 series to date, ahead of its immediate predecessor on graduate-level reasoning.
- Ships same-day in FP8, GGUF and NVFP4 quantized formats plus Fireworks/SGLang/vLLM/NVIDIA NIM serving support, so self-hosting doesn't require waiting on the community to catch up.
Cons
- DeepSWE coding score of 56.6 lags the field's current top performers, so the strong GPQA Diamond result doesn't carry through to every coding benchmark.
- Text-only despite the 'Max' branding: no image, audio or video input, unlike what several third-party writeups imply.
- No published training data cutoff date or safety system card as of August 2026, which complicates enterprise compliance review.
- Custom license requires a separate commercial agreement once trailing-year revenue crosses a set threshold for hosting or AI coding/office-assistant products, unlike a standard permissive open-source release.
Benchmarks
- IFBench: 82.8% vendor-reported · 20 Sep 2026 — How precisely the model follows detailed instructions, % passing.
- GPQA Diamond: 92.6% vendor-reported · 20 Sep 2026 — PhD-level science questions that are hard to search for, % correct.
- SWE-bench Pro: 67.7% vendor-reported · 20 Sep 2026 — Harder, longer real-repository coding tasks, % solved.
- OSWorld Verified: 86.1% vendor-reported · 20 Sep 2026 — Tasks completed by operating a real desktop, % solved.
- Terminal-Bench 2.1: 86.6% vendor-reported · 20 Sep 2026 — Multi-step tasks completed in a real command line, % solved.
- AA Intelligence Index: 40 cited: Artificial Analysis · 20 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- AA blended price: $1.18/M cited: Artificial Analysis · 20 Sep 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
- Output speed: 38 tok/s cited: Artificial Analysis · 20 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does Qwen3.8-2.4T-A95B cost to run via the API?
QwenCloud's international (Singapore) endpoint bills this model at .00 for every 1M input tokens processed and .00 per 1M output tokens generated, with an 8x discount on cached input at /usr/bin/bash.25 per 1M. That undercuts the prior-generation flagship's higher per-token rate, and OpenRouter routes the same model through multiple upstream hosts at a lower blended rate, close to .80/.40 per 1M. See the price notes above for the Mainland China endpoint's regional discount.
How does Qwen3.8-2.4T-A95B perform on independent and third-party benchmarks?
Alibaba's own model card reports 67.7% on SWE-bench Pro for this checkpoint, matching the figure covered in the full review above. On the independently tracked DeepSWE coding benchmark, though, it scores well below its graduate-level reasoning result, and Artificial Analysis's Intelligence Index likewise lands on the lower end, showing the model's reasoning strength doesn't carry through to every coding or composite benchmark equally. AIME 2025, MMLU-Pro and SWE-bench Verified scores remain unpublished as of this writing.
Is Qwen3.8-2.4T-A95B open source or proprietary?
This checkpoint ships as open weights on Hugging Face in safetensors, FP8, GGUF and NVFP4 formats: the first Qwen-Max-class release Alibaba has open-weighted. The license is custom to Qwen, not a standard permissive open-source license, and layers in two size-based obligations once a product scales: a UI-attribution requirement at high user counts, and a separate commercial agreement once trailing-year revenue from model-hosting or AI coding/office-assistant products crosses a set threshold. Read the LICENSE file on the Hugging Face repo for the exact figures.
Does Qwen3.8-2.4T-A95B train on user data?
Alibaba has not published a training data cutoff date or a system card for Qwen3.8-2.4T-A95B as of August 2026. For hosted use through QwenCloud, Alibaba's general terms of service govern data handling; no zero-retention or enterprise data-processing agreement specific to this model has been publicly documented yet.
Which teams should choose Qwen3.8-2.4T-A95B, and who should look elsewhere?
Qwen3.8-2.4T-A95B suits teams running self-hosted or multi-cloud agentic coding pipelines who want frontier-scale reasoning without single-vendor API lock-in. Teams needing native image or video input should look elsewhere: the checkpoint is text-only despite the 'Max' branding. Teams needing a published safety system card for compliance review may prefer [Claude Opus 5](/hub/models/claude-opus-5) or [GPT-5.6 Sol](/hub/models/gpt-5.6-sol), which both document red-teaming partners.
HokAI guides covering Qwen3.8-2.4T-A95B
- Best AI Models You Can Run Locally for Coding in 2026: Qwen3.8-27B beat Gemma 4 31B 12 to 6 on real coding tasks in a 24GB-GPU test. See which open-weight models actually fit your GPU, license and RAM in 2026.
- The Quirks Field: 373 Sentences No Vendor Would Publish: HokAI publishes 373 quirk entries across 102 of 103 model pages, each caveat checked today against the vendor's own model card, license, or terms page.