Qwen3.7-Plus review, pricing and limits

The multimodal, budget-tier agent model in the Qwen3.7 lineup, built for screen-reading and vision-driven automation at a sixth of Max's price.

  • ga
  • proprietary
  • multimodal
  • Qwen3.7 family
checked

Qwen3.7-Plus is built for teams that need an agent to see, not just read text: screen-reading automation, visual QA pipelines, and multimodal document review at a fraction of Qwen3.7-Max's cost. Its 1M-token context window holds an entire screen-recording transcript or document set in one call, but skip it for latency-critical loops or hard reasoning benchmarks.

Qwen3.7-Plus scores 39 on the Artificial Analysis Intelligence Index. Alibaba Cloud's multimodal agent model, released June 1, 2026, reads images and video as native input alongside text, unlike its text-only sibling Qwen3.7-Max, giving agents the ability to read a screen and act on what they see without a separate OCR step.

Where it sits

  • $0.7/M$ per 1M tokensBlended price (3:1)Lower is better#21 / 64peer median $1.70/Mvendor price, checked by HokAI
  • --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
  • --% solvedSWE-bench VerifiedHigher is better-- / 28peer median 78.3%per source, see benchmark scores
  • --% correctGPQA DiamondHigher is better-- / 44peer median 88.3%per source, see benchmark scores

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: Alibaba Cloud · Family: Qwen3.7

More about Alibaba Cloud on HokAI

Context window: 1,000,000 tokens · Max output: 32,000

Input modalities: text, image, video, tool-calls · Output: text, tool-calls

About Qwen3.7-Plus

Qwen3.7-Plus is Alibaba Cloud's multimodal agent model, previewed at the Alibaba Cloud Summit in Hangzhou on May 20, 2026 and shipped to general availability on the Bailian platform (marketed internationally as Model Studio) on June 1, 2026. It is the perception-and-action sibling to the text-only flagship Qwen3.7-Max, built by the same Qwen team inside Alibaba Group: where Max is tuned for raw text-reasoning throughput, Plus is built around a single mandate, give an agent eyes. It accepts text, static images, and video as input and reasons over all three in the same context, though it still only produces text as output.

Independent tracking from Artificial Analysis puts Qwen3.7-Plus well above the 16-point average for tracked models on its Intelligence Index composite, though it also flags the model as comparatively slow and verbose for its price bracket. Alibaba has not published a full sub-benchmark breakdown (SWE-bench, AIME, MMLU-Pro) for the Plus variant, unlike Max, which posts its own GPQA Diamond and SWE-bench Pro scores publicly. The model carries a 1M-token context window, the same ceiling as Max, and ships closed-weight and API-only, a departure from Alibaba's historical open-weight releases for smaller Qwen models. It's reachable through Alibaba Cloud's Bailian (Model Studio) platform as well as third-party gateways including Fireworks, Together AI, and OpenRouter.

Pricing

$0.40 per 1M input tokens, $1.60 per 1M output tokens, and roughly $0.08 per 1M on cached repeat-context calls via Alibaba Cloud Bailian (Model Studio).

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.012$0.0016$0.014
Support reply$0.0008$0.0005$0.0013
One coding agent run$0.080$0.032$0.112

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Vision and video as first-class input: Reads UI screenshots, charts, handwritten pages, and video frames alongside text in the same reasoning context, not as a bolted-on separate call.
  • 1M-token context window: Holds an entire multi-hour screen-recording transcript or a large multimodal document set in a single call, without external retrieval chunking, matching Qwen3.7-Max's ceiling.
  • Autonomous agent loop: Combines deep reasoning, self-programming, tool invocation, verification, and autonomous iteration to read screens and act without a human step in between.
  • Anthropic API protocol support: Speaks the Anthropic Messages protocol alongside Alibaba's native Bailian API, letting Claude-Code-style tooling point at it with minimal changes.
  • About one-sixth Qwen3.7-Max's per-token rate: Puts multimodal agent capability within reach of high-volume, budget-sensitive pipelines that can't justify Max's reasoning-tier pricing for screen-reading work.

Pros

  • Scores well above the tracked-model average on independent benchmarks, one of the stronger results in its price bracket.
  • Only agent model in the Qwen3.7 line with native vision and video input, letting it read screens instead of requiring text-described UI.
  • A fraction of Max's per-token rate makes continuous, high-volume screen-reading agents financially viable rather than a specialty expense.

Cons

  • Rated comparatively slow and verbose for its price bracket by Artificial Analysis, which can inflate output-token spend on chatty responses.
  • No published SWE-bench, AIME, or MMLU-Pro breakdown for this specific variant, so head-to-head technical comparisons against Max or competitors are harder to source.
  • Closed-weight and API-only: no downloadable checkpoint, no self-hosting, no fine-tuning of the base model.

Benchmarks

  • AA Intelligence Index: 39 cited: Artificial Analysis · 15 Jun 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
  • AA blended price: $0.72/M cited: Artificial Analysis · 15 Jun 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

What does Qwen3.7-Plus actually cost?

Qwen3.7-Plus costs $0.40 per 1M input tokens and $1.60 per 1M output tokens, with a lower $0.08 per 1M rate on cached repeat-context calls. That's roughly one-sixth Qwen3.7-Max's $2.50/$7.50 rate card for the same 1M-token context ceiling. A support pipeline reading 2,000 UI screenshots a day, at roughly 500 tokens of image context and 200 output tokens per response, runs under $1 a day at these rates.

Can you use Qwen3.7-Plus without paying?

No, Qwen3.7-Plus has no free tier: every call is billed at the standard per-token rate through Alibaba Cloud Bailian (Model Studio) or a third-party gateway. It also has no downloadable checkpoint to self-host as a free alternative, unlike some smaller Qwen models Alibaba has released with open weights. Teams wanting a free or self-hostable option should look at the openly licensed Qwen3.6 series instead.

What are Qwen3.7-Plus's closest competitors?

Qwen3.7-Plus's closest competitor is its own sibling, Qwen3.7-Max, which drops vision and video input for a much higher pure-reasoning ceiling. Outside the Qwen family, Google's Gemini 3.7 Flash and Gemini 3.1 Flash-Lite both accept image and video input at a comparable or lower price point, with fuller public benchmark disclosure than Plus currently has. Pick Plus when the workload needs Alibaba's Bailian ecosystem or Anthropic-protocol compatibility with Claude Code.

How does Qwen3.7-Plus compare to Qwen3.7-Max in 2026?

Qwen3.7-Max scores 56.6 on the Artificial Analysis Intelligence Index versus Plus's 39, and its GPQA Diamond and SWE-bench Pro results (92.4 and 60.6% respectively) have no published counterpart for Plus. Max is also text-only, while Plus adds native image and video input for screen-reading agents. Choose Max for hard reasoning and coding work judged on published benchmarks; choose Plus when the task genuinely needs vision or video and the budget doesn't support Max's higher per-token rate.

How do you set up Qwen3.7-Plus?

Qwen3.7-Plus is available through Alibaba Cloud's Bailian/Model Studio API, or through gateways like OpenRouter and Fireworks AI if you'd rather not create an Alibaba Cloud account directly. Official SDKs cover Python, TypeScript, and Java, and the model also accepts calls through the Anthropic Messages API schema, so an existing Claude Code or Anthropic-protocol client can point at the Bailian endpoint with only its base URL and model ID changed. For screen-reading or UI-testing use cases, pair it with a screenshot-capture tool like Playwright so images and instructions reach the model in the same call.

Top Alternatives

  • Qwen3.7-Max: Pick Qwen3.7-Max if you need pure text-reasoning throughput; pick Qwen3.7-Plus if the task requires reading images or video.
  • Gemini 3.7 Flash: Pick Gemini 3.7 Flash for audio input support and Google's ecosystem; pick Qwen3.7-Plus for a lower per-token rate on a matching 1M-token ceiling.
  • Gemini 3.1 Flash-Lite: Pick Gemini 3.1 Flash-Lite for a fully disclosed GPQA Diamond score at a similar price; pick Qwen3.7-Plus for Anthropic-protocol compatibility with Claude Code.

More AI Models on HokAI

Visit Qwen3.7-Plus Official Page