GLM-5.3-Flash is Z.ai's open-weight sibling to GLM-5.3, offering a 1,048,576-token context window and MIT-licensed weights published on Hugging Face for immediate self-hosting. It targets teams that need self-hosted, vision-and-video-capable reasoning without the parameter count or price tag of a full frontier model.
What changed
GLM-5.3-Flash is Z.ai's first natively multimodal GLM-5 model, released August 26, 2026 with 320 billion total parameters and 18 billion active per token. It scores 84.3 on Terminal-Bench 2.1, within a point of Claude Opus 4.8's 85.0, and ships as MIT-licensed open weights.
Where it sits
- $0.237/M$ per 1M tokensBlended price (3:1)Lower is better#16 / 79peer median $1.71/Mvendor price, checked by HokAI
- 53 tok/stokens/sOutput speedHigher is better#39 / 49peer median 90 tok/scited: Artificial Analysis
Cheaper than 81% of the 79 GA models with a published price, and rank 39 of 49 on output speed as cited from Artificial Analysis.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Z.ai · Family: GLM-5.3
Context window: 1,048,576 tokens · Max output: 131,072
Input modalities: text, image, video, pdf, tool-calls · Output: text, tool-calls
About GLM-5.3-Flash
GLM-5.3-Flash is a multimodal large language model built by Z.ai (formerly Zhipu AI), a Beijing-based AI lab that completed a Hong Kong IPO in January 2026. Z.ai confirmed the model's identity on August 26, 2026, after it had circulated for six days on OpenRouter, OpenCode and Cline under the codename "Ox Alpha." It is a Mixture-of-Experts transformer built from 320 billion total parameters, of which 18 billion activate per token, across 45 layers that mix KDA linear attention with NoPE sparse MLA attention and route each token through 8 of 288 experts. Z.ai's own release notes also credit Manifold-Constrained Hyper-Connections (mHC), an additional architectural technique used to improve scaling efficiency. It is the first natively multimodal release in the GLM-5 line, sitting alongside the larger, text-only GLM-5.3 (released August 14, 2026) as the faster, cheaper sibling to the earlier GLM-5.2.
On agentic and coding evaluations, GLM-5.3-Flash's Terminal-Bench 2.1 result of 84.3 sits a point under Claude Opus 4.8's 85.0 and just short of GPT-5.6 Terra's 87.4. On Z.ai's own Code Bench v1.0 it scores 29.0 against Opus 4.8's 29.5. On DeepSWE v1.1 it jumps to 63.4, up from 46.2 for the previous GLM-5.2, and AutomationBench nearly doubles to 48.8 from 26.2. Z.ai's launch disclosures also reported narrower scores on MVBench (77.8), OfficeQA-Pro (62.4), ToolAthlon (78.4), CharXiv reasoning (89.4), ChartoGraphy (78), NL2Repo (56.3), BabyVision (53.4), Agent's Last Exam (26.3) and HLE with tools (55.3), none of which Artificial Analysis tracks directly. Under Artificial Analysis's current Intelligence Index methodology (v4.3.2, re-benchmarked since launch), GLM-5.3-Flash's composite score is 42, placing it 4th of 116 tracked models in its class; Z.ai's original day-one figure of 57 was measured under the prior v4.1.1 methodology and is no longer the number Artificial Analysis reports. Standard academic benchmarks such as GPQA Diamond, MMLU-Pro and SWE-bench Verified still had not been separately published for the Flash variant as of this writing, and Artificial Analysis's own evaluation suite has since dropped those three benchmarks from its tracked set entirely; Z.ai's disclosures continue to focus on agentic and vision tasks instead. Z.ai has since begun selling GLM-5.3-FlashX, a separate, faster paid variant of the same line running at roughly 200 tokens per second, priced above standard Flash.
The model ships with a 1,048,576-token context window, built on an IndexPool mechanism that compresses indexer key vectors, cutting attention compute roughly 3x and KV-cache size 4.4x compared with the full GLM-5.3. Maximum completion length is 131,072 tokens (128K) on the official API, up from the 48,000-token cap Z.ai listed at launch. Z.ai has not published a separate long-context recall score specific to this model.
GLM-5.3-Flash accepts text, images, video and files as input and returns text, including structured JSON output. Z.ai's documentation describes tool-calling support for external tool integration, and the model runs in an always-on thinking mode: reasoning cannot be disabled, though Z.ai exposes a reasoning_effort parameter (low, high, max) to tune how much the model thinks before answering. max is the default and the setting Z.ai used for its own published benchmarks. A vision encoder handles image and video frames, letting the model read screenshots, interfaces, charts and documents inside an agent loop, then inspect its own rendered output and iterate.
Standard API metering runs at a fraction of the full-size GLM-5.3's rate (exact figures in the pricing tab). A launch-week discount that halved those prices through September 9, 2026 has since ended; Z.ai's standing rate is what the pricing tab now shows. The model is bundled into every tier of the GLM Coding Plan, from the entry Lite plan to the top Max plan, at three times the usable quota of GLM-5.3.
The primary channel is Z.ai's own OpenAI-compatible developer API at docs.z.ai. Open weights are published as zai-org/GLM-5.3-Flash on Hugging Face under the MIT license, with a native FP8 safetensors checkpoint (about 306 GiB) and a BF16 variant at roughly double that size; community GGUF and NVFP4 quantizations followed within days, the same self-hosting pattern seen with open-weight peers like Kimi K3 and DeepSeek V4. Self-hosting the FP8 checkpoint needs an 8-GPU node of NVIDIA Hopper-class or newer hardware, or AMD Instinct gfx950 via ROCm, with vLLM, SGLang, KTransformers and TokenSpeed all shipping day-one recipes. OpenRouter, Cloudflare Workers AI and Vercel's AI Gateway all added the model the same week as the official announcement. As of this writing, none of the big three hyperscaler catalogs (AWS Bedrock, Google Vertex AI, Azure AI Foundry) list GLM-5.3-Flash directly, though they carry earlier GLM-5 releases.
Z.ai's public governance disclosures describe supervised fine-tuning and RLHF alignment plus a hosted-API content filter, with no formal responsible-scaling-policy analog to Anthropic's RSP. For the wider GLM-5.3 release, Z.ai delayed shipping open weights to complete a safety evaluation after finding the model's code-vulnerability discovery ability had increased sharply; no equivalent delay has been reported for the Flash variant specifically. The MIT license lets outside researchers audit and red-team the weights directly, which is the company's main stated safety lever rather than a published red-team partner list.
Teams building agentic coding tools, browser or computer-use agents, or document and screenshot-reading pipelines get near-frontier agentic scores at open-weight prices, rivaling closed models like Claude Opus 4.8 and GPT-5.6 Terra on Terminal-Bench 2.1 specifically, and can self-host for data residency reasons closed frontier labs cannot match. Teams needing verified academic reasoning scores (GPQA, MMLU-Pro, formal SWE-bench Verified) should look elsewhere until Z.ai publishes those numbers for this specific checkpoint. Raw generation speed, now measured at roughly 52.5 tokens per second by Artificial Analysis, still trails the median for comparable open-weight models, so latency-sensitive, high-throughput chat products may prefer GLM-5.2 or a smaller model instead.
Z.ai trained the underlying architecture on a 30-trillion-token multimodal corpus; the company has not disclosed a specific training data cutoff date for the Flash checkpoint. Z.ai has not published SOC 2, HIPAA or GDPR certifications for its hosted API, and data residency for the hosted service defaults to China-based infrastructure plus whatever regions third-party API resellers offer. The MIT license carries no data-retention terms of its own; Z.ai's general API terms, not the model card, govern how the hosted service retains or trains on submitted data.
GLM-5.3-Flash follows GLM-5.2 (open weights, released in mid-June 2026) and the full-size GLM-5.3 (released August 14, 2026), arriving twelve days later as the multimodal, cheaper sibling. It had already been serving production traffic anonymously as "Ox Alpha" on OpenRouter, OpenCode, Cline and Nous Research's portal since August 20, 2026, before Z.ai confirmed the identity match on August 26 and published weights the same day. Readers comparing it against other coding-focused models can see HokAI's best AI coding assistants guide and the broader how to choose the right LLM guide, or browse the full slate of models HokAI tracks.
Pricing
List price is $0.15 per 1M input tokens, $0.50 per 1M output, and $0.03 per 1M cached input via Z.ai's official API, a fraction of the full [GLM-5.3](/hub/models/glm-5.3)'s $1.40/$4.40 rate. A 50%-off launch promotion (to $0.075/$0.25/$0.015) ended September 9, 2026; the price above is Z.ai's standing rate, not a temporary discount. Open weights can also be self-hosted for the cost of compute only.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.0045 | $0.0005 | $0.0050 |
| Support reply | $0.0003 | $0.0001 | $0.0004 |
| One coding agent run | $0.030 | $0.010 | $0.040 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Native Multimodal Input: Accepts text, images, video and files in a single request, the first GLM-5 release built multimodal from the ground up rather than bolted on.
- Million-Token-Class Context Window: A vision-capable long-context window built on Z.ai's IndexPool mechanism, which trims attention compute and KV-cache size well below the full GLM-5.3's footprint.
- MIT-Licensed Open Weights: Full-precision and quantized checkpoints published on Hugging Face under MIT, permitting unrestricted commercial use, fine-tuning and self-hosting.
- Sharp Agentic-Coding Gains Over GLM-5.2: DeepSWE v1.1 climbed to 63.4 from GLM-5.2's 46.2, and AutomationBench nearly doubled to 48.8, the biggest jump of any metric Z.ai disclosed at launch.
- Steep Discount Off the Flagship Rate: Priced far under what Z.ai charges for the full-size GLM-5.3, and a launch-week promo trims the bill even more into early autumn 2026.
Pros
- Artificial Analysis ranks its Intelligence Index third of 109 tracked models at launch, an unusually strong showing for an open-weight release.
- A vision-capable, million-token-class context window with a meaningfully lighter attention and KV-cache footprint than the full GLM-5.3.
- MIT license allows self-hosting, fine-tuning, and independent audit of the weights, unlike closed frontier models.
Cons
- No published GPQA Diamond, MMLU-Pro, or SWE-bench Verified score for this checkpoint as of launch.
- Raw output speed sits below the median for its open-weight peer group, so it is not the fastest choice for high-throughput chat.
- Self-hosting the FP8 checkpoint needs an 8-GPU Hopper-class node (about 306 GiB VRAM), out of reach for smaller teams.
Benchmarks
- GDPval-AA v2: 1,773 vendor-reported · 26 Aug 2026 — Real knowledge-work deliverables judged against professionals, run by Artificial Analysis.
- Terminal-Bench 2.1: 84.3% vendor-reported · 26 Aug 2026 — Multi-step tasks completed in a real command line, % solved.
- AA Intelligence Index: 42 cited: Artificial Analysis · 29 Sep 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
- AA blended price: $0.24/M cited: Artificial Analysis · 29 Sep 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
- Output speed: 53 tok/s cited: Artificial Analysis · 29 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What are GLM-5.3-Flash's API pricing plans in 2026?
Z.ai's standing price is $0.15 for every 1M input tokens, $0.50 for every 1M output tokens, and $0.03 for every 1M cached input tokens through the official API. A 50%-off launch promotion applied through September 9, 2026; it has since ended, and the rate above is what the API currently charges, not a temporary discount. Rates on third-party hosts, such as OpenRouter or Cloudflare's Workers AI, can differ from the official price above.
What has changed for GLM-5.3-Flash since its August 2026 launch?
The launch-week discount pricing ended in September 2026, so the official API now bills at Z.ai's standard rate rather than the promotional one. Z.ai's documentation also now lists a 131,072-token maximum output, well above the 48,000-token cap quoted at launch. Separately, Artificial Analysis has revised its Intelligence Index methodology since release; GLM-5.3-Flash's tracked composite score is now 42 under the current version, not the 57 Z.ai originally cited.
Is GLM-5.3-Flash open source or proprietary?
GLM-5.3-Flash is released under the MIT license, with FP8 and BF16 weight checkpoints published on Hugging Face as zai-org/GLM-5.3-Flash. MIT permits unrestricted commercial use, modification, and self-hosting, unlike the restricted open-weights licenses some rival open models use.
Does GLM-5.3-Flash's API retain or train on your inputs?
Z.ai has not published a model-specific data retention or training-on-inputs disclosure for GLM-5.3-Flash; that policy is set by the company's general API terms of service rather than the model card. Teams with strict data-residency requirements can avoid the question entirely by self-hosting the MIT-licensed weights.
What is GLM-5.3-Flash best used for?
It suits agentic coding, terminal-automation, and screenshot- or video-aware agent loops, where its strong Terminal-Bench score and 1M-token context do most of the work. It is a weaker choice for workloads that need verified academic reasoning benchmarks or high-throughput, low-latency chat, since neither GPQA-class scores nor top-tier raw speed have been demonstrated for this checkpoint.
HokAI guides covering GLM-5.3-Flash
- The Cheapest LLM APIs in 2026, and What Each Cheap Price Leaves Out: Compare six cheap LLM APIs on price per million tokens, quality score and fine print: Claude Haiku 5.5, GLM-5.3 Flash, Ling, Gemini, DeepSeek and Mistral.
- Best AI Models You Can Run Locally for Coding in 2026: Qwen3.8-27B beat Gemma 4 31B 12 to 6 on real coding tasks in a 24GB-GPU test. See which open-weight models actually fit your GPU, license and RAM in 2026.