Qwen3.8-Max is Alibaba's flagship large language model, a 2.4 trillion parameter mixture-of-experts system released August 3, 2026 with a 1 million token context window and text, image and video input. It targets coding and agentic workloads at $2 per million input tokens and $6 per million output tokens, undercutting Claude and GPT-5.6 on price.
Summary: Alibaba priced its new Qwen3.8-Max model at $2 input and $6 output per million tokens, cheaper than Claude Sonnet 5's promotional rate and OpenAI's GPT-5.6 Terra, and promised open weights this week. The benchmark numbers behind that pitch are the vendor's own, and the license for those weights still hasn't been named.
Alibaba priced its newest flagship model, Qwen3.8-Max, at $2 per million input tokens and $6 per million output tokens on August 3, 2026, undercutting every major US rival on a like for like basis.
The timing is not an accident. The cut lands the same week the company promised to publish open weights for the model, a first at flagship scale. For a startup routing a new agent feature through a model API and watching per-token cost line by line, that combination reads like a free upgrade: frontier-class scores at a fraction of the going rate, with a self-hosted option due days later. The numbers behind that pitch, though, come from one place: the vendor itself.
Qwen3.8-Max is a 2.4 trillion parameter mixture-of-experts model, the largest Qwen release to date, with a 1 million token context window and native text, image and video input, according to MarkTechPost's coverage of the release. Cached input runs $0.25 per million tokens, a tier most competitors do not publish at all.
At $2 and $6 per million tokens, it costs less on output than Claude Sonnet 5's introductory rate of $2 and $10, a promotion that expires August 31, 2026 and reverts to $3 and $15 afterward, according to Digital Applied's pricing tracker. It also beats OpenAI's GPT-5.6 Terra, cut 20% to $2 and $12 on July 30, and sits well below GPT-5.6 Sol's $5 and $30 or Claude Opus 5's $5 and $25, according to Forbes' reporting on the launch.
Alibaba Cloud frames this as a computing strategy, not a model-quality flex. AI products already make up 30% of its external revenue, with triple-digit quarterly growth the division expects to push past 50% within a year, Forbes reported. A cheap, widely used model is how a cloud provider sells the underlying compute. Selling access below cost for a while is a familiar trade if it locks in developers early, the same logic that has driven price wars in databases and search before it.
Teams that route model calls through OpenRouter get the clearest win, at least on paper. A new flagship model slots in as a lower-cost option the moment it appears on the marketplace, no vendor negotiation required, and no rewritten integration beyond swapping a model string.
Together AI and other neutral hosts stand to gain too. If the open weights land under a usable license, they can serve the model at their own margin instead of routing traffic through a single vendor's cloud, the same way they already host open checkpoints from several labs.
Solo developers and small teams building agent features on tight budgets are the other clear winner, at least on the hosted API. A four person startup burning through $2,000 a month in tokens on a non-critical agent loop has a concrete reason to test the new model this week, not next quarter. Swapping one API endpoint for another is a low-risk experiment when the downside is a worse answer, not a broken pipeline.
DeepSeek feels this first. Its V4 Flash coding model built its entire pitch on being the cheap, good enough alternative to frontier models. Qwen3.8-Max now claims similar territory at a much larger scale, with multimodal input the smaller model does not offer, narrowing the gap that made V4 Flash the obvious budget pick.
Anthropic and OpenAI absorb pressure from a different direction. Claude Sonnet 5's discount pricing was already scheduled to expire August 31. A cheaper, credibly-positioned option gives price-sensitive teams a live alternative to sit with while that clock runs down, rather than a reason to simply accept the coming 50% increase without shopping around first.
Here is the part the pricing story leaves out. Forbes reported that the benchmark scores circulating online, including Terminal-Bench, GPQA Diamond and OSWorld-Verified numbers other outlets have repeated without qualification, came from internal testing, and that no independent evaluation existed at launch. By the vendor's own account, on the LMArena leaderboards, the new model ranked second in the Vision Arena and fifth in the Text Arena, trailing the Claude family rather than beating it outright.
That does not make the price cut fake. It makes the frontier-for-a-quarter-of-the-cost framing premature. A team switching workloads on a single scorecard is trusting the seller's own numbers, the same position buyers were in with every prior model launch before outside evaluators caught up and occasionally told a different story.
The open-weights promise is the harder test, and it is unresolved as of this writing. Alibaba said on August 3 that weights for Qwen3.8-Max and a smaller Qwen3.8-27B would reach Hugging Face and ModelScope during the week of August 10, but as of publication neither repository is live and no license has been named, according to Digital Applied's license checklist.
History offers a hint, not a promise. Qwen3.5 and Qwen3.6-27B shipped under Apache 2.0, which permits commercial self-hosting without restriction. Some larger past Qwen releases shipped under the more restrictive Tongyi Qianwen terms instead, which caps free commercial use below a fixed number of monthly active users.
If the license lands as Apache 2.0 before August 17, treat Qwen3.8-Max as a real self-hosting option for coding and agent workloads, not just a cheaper API call. If it lands restricted, or the repository stays empty past that date, the only thing that shipped this month is a good price on a hosted API. Still worth using. Just not the open default it was pitched as.
Qwen3.8-Max is Alibaba's flagship large language model, built for coding, agentic workflows and multimodal tasks that mix text, images and video. It is priced at $2 per million input tokens and $6 per million output tokens, which Alibaba positions as frontier-class performance at a fraction of Claude's or GPT-5.6's cost. The benchmark scores behind that positioning come from Alibaba's own testing rather than an independent lab.
At $2 input and $6 output per million tokens, Qwen3.8-Max undercuts Claude Sonnet 5's introductory rate of $2 and $10, which reverts to $3 and $15 on August 31, 2026, and OpenAI's GPT-5.6 Terra at $2 and $12. It is also far cheaper than the flagship tiers: Claude Opus 5 at $5 and $25, and GPT-5.6 Sol at $5 and $30.
Not as of publication. Alibaba said on August 3, 2026 that it would publish weights for Qwen3.8-Max and a smaller Qwen3.8-27B on Hugging Face and ModelScope during the week of August 10, but as of this writing neither model has appeared and no license has been named.
Alibaba has not said. Its two most recent open releases, Qwen3.5 and Qwen3.6-27B, shipped under the permissive Apache 2.0 license, but that is prior behavior, not a commitment for this release, and some larger past Qwen models shipped under the more restrictive Tongyi Qianwen license instead.
The clearest pressure lands on DeepSeek, whose V4 Flash coding model built its pitch on being the cheap alternative to frontier models, a position Qwen3.8-Max now contests directly at larger scale. Anthropic and OpenAI feel it too: Claude Sonnet 5's discount pricing was already set to expire August 31, and Qwen gives cost-sensitive teams a reason not to simply accept that increase.
All AI guides · Browse the AI directory
Still deciding? Get matched.
Smart Match checks every listing in the directory against how you work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart Match