Qwen3.7-Max review, pricing and verdict

Alibaba Cloud's text-first agent flagship combining 1M context and 92.4 GPQA Diamond at half the cost of most Western frontier models.

  • ga
  • proprietary
  • chat
  • Qwen3.7 family
checked

Qwen3.7-Max fits teams building long-horizon coding agents and scientific reasoning pipelines who want frontier-level reasoning without frontier pricing. It scores 60.6 on SWE-Pro, ahead of many proprietary rivals on agentic coding. Skip it for multimodal work (use Qwen3.7-Plus instead) or air-gapped deployments, since there are no open weights.

Qwen3.7-Max is Alibaba Cloud's proprietary agent and reasoning model, scoring 92.4 on GPQA Diamond as of its May 2026 release. It is a text-only Mixture-of-Experts model built for long-horizon autonomous coding and tool-use workflows, with the largest context window in the Qwen3.x lineup and no open-weight release.

Where it sits

  • $3.75/M$ per 1M tokensBlended price (3:1)Lower is better#43 / 64peer median $1.70/Mvendor price, checked by HokAI
  • 113 tok/stokens/sOutput speedHigher is better#14 / 39peer median 90 tok/scited: Artificial Analysis
  • --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
  • 92.4%% correctGPQA DiamondHigher is better#13 / 44peer median 88.3%per source, see benchmark scores

Pricier than 67% of the 64 GA models with a published price, in the top third on GPQA Diamond (rank 13 of 44), and one of 65 whose vendor states it does not train on customer data.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Alibaba Cloud · Family: Qwen3.7

More about Alibaba Cloud on HokAI

Context window: 1,000,000 tokens · Max output: 65,536

Input modalities: text, tool-calls · Output: text, tool-calls

About Qwen3.7-Max

Qwen3.7-Max is Alibaba Cloud's proprietary flagship reasoning and agent model, released May 19, 2026, one day ahead of its official announcement at the Alibaba Cloud Summit in Hangzhou. Built by the Qwen research team under Alibaba Group, the model sits at the top of the Qwen3.7 series, which also includes the multimodal Qwen3.7-Plus released June 3, 2026. The architecture is Mixture-of-Experts with an undisclosed but estimated parameter count exceeding 1 trillion, following the MoE lineage of its predecessor Qwen3-235B-A22B. It was designed specifically as an "Agent Frontier" model for long-horizon autonomous execution, multi-step reasoning, and tool-use workflows rather than general-purpose chat.

On reasoning benchmarks, Qwen3.7-Max scores 92.4 on GPQA Diamond, 97.1 on HMMT 2026 Feb, and 90 on IMOAnswerBench, placing it among the strongest reasoning models in mid-2026. On agentic coding it achieves 60.6 on SWE-Pro and 78.3 on SWE-Multilingual, beating Qwen3.6-Max and matching or exceeding several proprietary competitors at a fraction of the cost. Artificial Analysis ranks it at 56.6 on its composite Intelligence Index with output speed of 113.1 tokens per second. LM Arena placed the Qwen3.7-Max-Preview at approximately Elo 1,475, ranked 13th overall and 7th for math as of May 2026. Compared to GPT-5.5 and Claude Opus 4.7, Qwen3.7-Max wins on cost efficiency per reasoning task while trailing on general assistant tasks and modality coverage.

Qwen3.7-Max ships with a 1,000,000-token context window and a maximum output of 65,536 tokens per request. Alibaba Cloud DashScope lists a maximum input of 991,800 tokens after internal formatting overhead. This is a fourfold increase over Qwen3-Max's 262,144-token limit, enabling processing of large codebases, lengthy legal documents, or multi-session agent transcripts in a single call. Prompt caching applies to repeated system prompt content at a steep discount versus standard input pricing, making repeat-context agent loops significantly cheaper.

Qwen3.7-Max is a text-only model: it accepts text and structured tool call inputs and produces text and tool call outputs, but does not support vision, audio, video, or image inputs. For multimodal workflows, Alibaba released Qwen3.7-Plus on June 3, 2026, which adds image and video understanding, deep reasoning, self-programming, tool invocation, verification testing, and autonomous iteration. Qwen3.7-Max supports native function calling with OpenAI-compatible tool schemas, structured JSON output, parallel tool calls, and an extended thinking mode for deeper chain-of-thought. The model targets sustained autonomous execution across hundreds or thousands of steps, including code writing, debugging, and office workflow automation.

Exact per-token pricing and the current launch promotion are detailed in the pricing section below. Summarizing a 200K-token document costs roughly $0.50 input plus $0.08 output at standard rates; a daily coding agent processing 1M tokens in and 100K tokens out costs about $3.25; a 1,000-turn customer support pipeline at 3K in and 800 out per turn runs roughly $8.10 per day. Compared to Claude Opus 4.7 or GPT-5 at higher output rates, Qwen3.7-Max is among the most affordable frontier-tier options for output-heavy agentic workloads.

Qwen3.7-Max is available via Alibaba Cloud Model Studio (DashScope API), OpenRouter (qwen/qwen3.7-max), Together AI, Fireworks AI, and ModelScope, all live from May 19, 2026. Together AI and Fireworks provide US-region hosting for teams with data residency requirements outside China. The model ID on DashScope is qwen3.7-max and authentication requires a DashScope API key obtained through Alibaba Cloud account creation. There are no open weights, no self-hosting option, and no fine-tuning support for Qwen3.7-Max.

Qwen3.7-Max uses RLHF and instruction-tuning alignment, with a hallucination rate of 22.9% on TruthfulQA-style evaluations. The model abstains on more than 50% of questions it previously attempted in prior versions, reflecting a conservative posture on ambiguous inputs. Alibaba has not published a standalone system card for Qwen3.7-Max, though the official Qwen3.7 blog post at qwen.ai provides technical details and safety notes. The model refuses CSAM, weapons manufacturing instructions, and malware generation by default.

Qwen3.7-Max is best for teams building text-based long-horizon autonomous agents, coding agents processing large codebases, and scientific reasoning pipelines where top-tier GPQA Diamond performance and competition math scores matter. Its SWE-Multilingual score of 78.3 makes it a strong pick for multilingual software engineering. Teams building multimodal applications requiring vision or audio should use Qwen3.7-Plus instead. Organizations with strict air-gapped or on-premise requirements cannot use Qwen3.7-Max (no open weights); Qwen3-32B under Apache 2.0 is the best open alternative. Teams prioritizing human preference alignment over raw reasoning scores may prefer Claude Opus 4.7 or GPT-5.5, which rank higher on LM Arena overall.

Pricing

$2.50 per 1M input, $7.50 per 1M output, $0.25 per 1M cached input (90% discount). A 50% launch promotion is available at $1.25/$3.75 per 1M through June 22, 2026. Max output is 65,536 tokens per request.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.075$0.0075$0.082
Support reply$0.0050$0.0022$0.0072
One coding agent run$0.500$0.150$0.650

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • 1M-Token Context Window: Handles up to 1,000,000 input tokens per request, a fourfold increase over the prior Qwen3-Max, enabling full-codebase and long-document analysis without chunking.
  • Extended Thinking Mode: Generates visible chain-of-thought before the final answer, boosting accuracy on hard reasoning and competition-level math tasks like HMMT 2026 Feb (97.1).
  • Native Tool Use and Function Calling: Supports parallel function calls with OpenAI-compatible tool schemas and structured JSON output, built for sustained agentic execution across hundreds of tool-call turns.
  • Prompt Caching at $0.25/M: Cached system prompt reads cost 90% less than standard input tokens, making repeat-context agent loops and multi-turn sessions significantly cheaper.
  • Multi-Provider Access: Available on Alibaba Cloud DashScope, OpenRouter, Together AI, and Fireworks AI from day one, covering both Asian and US-region hosting options.

Pros

  • Ranks among the strongest reasoning models available under $10/M output, per its GPQA Diamond and HMMT 2026 benchmark scores.
  • A 1M-token context window, the largest in the Qwen3.x series, suits large codebase and document workflows without chunking.
  • 113.1 tokens per second output speed (Artificial Analysis) is above average for reasoning-tier models in the same price band.

Cons

  • Text-only input with no vision or audio support, requiring a model switch to Qwen3.7-Plus for any multimodal task.
  • Closed weights with no self-hosting or fine-tuning, ruling out air-gapped deployments and custom adaptation.
  • 22.9% hallucination rate on TruthfulQA-style evaluations, higher than competitors specifically optimized for factual grounding.

Benchmarks

  • LMArena Elo: 1475 independent · 20 May 2026 — Rating from blind human votes on which answer is better.
  • GPQA Diamond: 92.4% vendor-reported · 19 May 2026 — PhD-level science questions that are hard to search for, % correct.
  • LMArena rank: #13 independent · 20 May 2026 — Position on the blind human-preference leaderboard; #1 is best.
  • AA Intelligence Index: 56.6 cited: Artificial Analysis · 18 Jun 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
  • AA blended price: $3.63/M cited: Artificial Analysis · 18 Jun 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
  • Output speed: 113 tok/s cited: Artificial Analysis · 18 Jun 2026 — Median tokens written per second as measured by Artificial Analysis.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

How much do you pay for Qwen3.7-Max?

Qwen3.7-Max's standard Alibaba Cloud DashScope rate is $2.50 for input and $7.50 for output, each per 1M tokens, with a steep discount on cached input reads. A limited-time launch promotion cuts both rates roughly in half through late June 2026. The model is also available through OpenRouter, Together AI, and Fireworks AI at comparable rates.

Is Qwen3.7-Max free to use?

No, Qwen3.7-Max has no free tier: every request through Alibaba Cloud DashScope, OpenRouter, Together AI, or Fireworks AI is billed per token from the first call. Teams wanting to test it without commitment should budget for a small trial run using the published cost examples, since there is no free allowance to fall back on.

What are Qwen3.7-Max's closest competitors?

Within the Qwen family, Qwen3.7-Plus is the direct upgrade path for teams that need image and video understanding rather than text-only agentic work. Qwen3-32B, released under an Apache 2.0 license, is the closest open-weight alternative for teams that need self-hosting or air-gapped deployment. Outside Alibaba's lineup, Claude Opus 4.7 and GPT-5.5 are the main proprietary competitors for teams prioritizing general assistant quality and human-preference alignment over raw reasoning-per-dollar.

Qwen3.7-Max or Claude Opus 4.7: which should you pick?

Qwen3.7-Max leads Claude Opus 4.7 on GPQA Diamond and other hard-reasoning benchmarks, and it is priced well below Opus's output rate on Alibaba Cloud DashScope. Opus still leads on general assistant polish, tool-orchestration reliability, and LM Arena human-preference ranking, and it supports vision input that Qwen3.7-Max does not. Teams optimizing purely for reasoning-per-dollar on text tasks should lean toward Qwen3.7-Max; teams wanting a single multimodal assistant should stay with Opus.

What does it take to start using Qwen3.7-Max?

Sign up for an Alibaba Cloud account and generate a DashScope API key, then call the qwen3.7-max model ID with an OpenAI-compatible request for text and tool-call inputs. Teams outside China can instead reach the model through OpenRouter, Together AI, or Fireworks AI, which skips the DashScope account-creation step and adds US-region hosting. There is no open-weight download and no self-hosting path, so every integration goes through one of these four API providers.

Top Alternatives

  • Claude Opus 4.7: Pick Qwen3.7-Max if you want frontier-level reasoning at a fraction of Opus's output pricing; pick Claude Opus 4.7 if you need native vision input and broader modality support.
  • GPT-5.5: Pick Qwen3.7-Max for cheaper agentic-coding output pricing than GPT-5.5; pick GPT-5.5 for native audio, video, and image understanding Qwen3.7-Max lacks.
  • Qwen3.7-Plus: Pick Qwen3.7-Max for pure text and agentic coding workloads; pick Qwen3.7-Plus if you need vision and video input in the same Qwen3.7 family.

More AI Models on HokAI

Visit Qwen3.7-Max Official Page