Side-by-side comparison of DeepSeek-V4-Pro-0813, GLM-5.2: pricing, capabilities, integrations and compliance — from verified HokAI records.
DeepSeek
Pick it if: DeepSeek-V4-Pro-0813 ships MIT-licensed open weights, with 49B parameters active per token and non-think, think-high, and think-max reasoning tiers callers pick per request. It suits cost-sensitive agentic coding teams willing to add an external safety layer, since independent 2026 red-teaming found refusal collapses under free-form prompting.
Its edge: SWE-bench Verified 80.6%, matching Gemini-3.1-Pro's agentic coding score, per DeepSeek V4 Pro GA coverage.
The catch: Text-only: no native vision, audio, or video input, despite pre-launch reporting that expected multimodal training.
Z.ai
Pick it if: GLM-5.2 is Z.ai's flagship released June 13, 2026 on a 744B MoE architecture with 40B active parameters and a 1,000,000-token context window. Priced at $1.40/$4.40 per 1M tokens under an MIT license, it scores 80.3% GPQA Diamond and 62.1% SWE-bench Pro.
Its edge: Tops open-source SWE-bench Pro at 62.1% as of June 2026, 3.7 points ahead of predecessor GLM-5.1.
The catch: No native image, audio, or video input — vision tasks require the separate GLM-5V-Turbo model.
| Input / 1M tokens | $0.435 | $1.40 |
|---|---|---|
| Output / 1M tokens | $0.87 | $4.40 |
| Cached input / 1M | $0.004 | $0.26 |
| Blended cost (3:1) | $0.544 | $2.15 |
| Free tier | false | false |
| Pricing model | per-token | per-token |
| Strengths | SWE-bench Verified 80.6%, matching Gemini-3.1-Pro's agentic coding score, per DeepSeek V4 Pro GA coverage.; MMLU-Pro 87.5% and LiveCodeBench 93.5% pass@1, among the highest reported scores for an open-weight model.; Outputs at 80.0 tokens/s | Tops open-source SWE-bench Pro at 62.1% as of June 2026, 3.7 points ahead of predecessor GLM-5.1.; 1-million-token context window at $1.40/1M input — the largest context available in any MIT-licensed model.; MIT license and OpenAI-compatibl |
|---|---|---|
| Limitations | Text-only: no native vision, audio, or video input, despite pre-launch reporting that expected multimodal training.; FAR.AI's independent stress test found 98-100% jailbreak success across CBRN, cyber, and terrorism prompts using an unmodif | No native image, audio, or video input — vision tasks require the separate GLM-5V-Turbo model.; No disclosed SOC 2, HIPAA, or ISO 27001 certifications, limiting adoption in regulated industries.; 1M-context recall quality at depth has no pu |
| Context window | 1M | 1M |
|---|---|---|
| Max output tokens | 384K | 131.1K |
| Input modalities | text; tool-calls | text; tool-calls |
| Output modalities | text; tool-calls | text; tool-calls |
| Capabilities | Tool use; Function calling; Structured output | Tool use; Function calling; Structured output |
| Reasoning modes | non-think; think-high; think-max | standard; extended-thinking-high; extended-thinking-max |
| Long-context recall | — | medium |
| Openness | open-source | open-source |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.