Side-by-side comparison of GLM-5.2, Kimi K3: pricing, capabilities, integrations and compliance — from verified HokAI records.
Z.ai
Pick it if: GLM-5.2 is Z.ai's flagship released June 13, 2026 on a 744B MoE architecture with 40B active parameters and a 1,000,000-token context window. Priced at $1.40/$4.40 per 1M tokens under an MIT license, it scores 80.3% GPQA Diamond and 62.1% SWE-bench Pro.
Its edge: Tops open-source SWE-bench Pro at 62.1% as of June 2026, 3.7 points ahead of predecessor GLM-5.1.
The catch: No native image, audio, or video input — vision tasks require the separate GLM-5V-Turbo model.
Moonshot AI
Pick it if: Kimi K3 launched July 16, 2026 with 2.8 trillion total parameters (280B active) and a 1,048,576-token context window, the largest announced open-weight-class model to date. It targets teams doing agentic coding and long-document analysis who want an open alternative to GPT-5.6 Sol and Claude Opus 4.8.
Its edge: Highest published open-weight GPQA Diamond score at 93.5%, ahead of Opus 4.8's 91.0%.
The catch: Hallucination rate rose to 51% on AA-Omniscience versus 39% for predecessor K2.6, more confident fabrication despite higher accuracy.
| Input / 1M tokens | $1.40 | $3 |
|---|---|---|
| Output / 1M tokens | $4.40 | $15 |
| Cached input / 1M | $0.26 | $0.3 |
| Blended cost (3:1) | $2.15 | $6 |
| Free tier | false | false |
| Pricing model | per-token | per-token |
| Strengths | Tops open-source SWE-bench Pro at 62.1% as of June 2026, 3.7 points ahead of predecessor GLM-5.1.; 1-million-token context window at $1.40/1M input — the largest context available in any MIT-licensed model.; MIT license and OpenAI-compatibl | Highest published open-weight GPQA Diamond score at 93.5%, ahead of Opus 4.8's 91.0%.; Terminal-Bench 2.1 score of 88.3%, half a point behind GPT-5.6 Sol and ahead of every other open or closed model tested.; Program Bench 77.8% and SWE Mar |
|---|---|---|
| Limitations | No native image, audio, or video input — vision tasks require the separate GLM-5V-Turbo model.; No disclosed SOC 2, HIPAA, or ISO 27001 certifications, limiting adoption in regulated industries.; 1M-context recall quality at depth has no pu | Hallucination rate rose to 51% on AA-Omniscience versus 39% for predecessor K2.6, more confident fabrication despite higher accuracy.; Committed weight release (Modified MIT) not yet live on Hugging Face as of this writing, expected by July |
| Context window | 1M | 1M |
|---|---|---|
| Max output tokens | 131.1K | 128K |
| Input modalities | text; tool-calls | text; image |
| Output modalities | text; tool-calls | text; tool-calls; code |
| Capabilities | Tool use; Function calling; Structured output | Vision; Tool use; Code execution |
| Reasoning modes | standard; extended-thinking-high; extended-thinking-max | none; low; medium |
| Long-context recall | medium | high |
| Openness | open-source | open-weights |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.