Side-by-side pricing, features and compliance for any tools in the directory.
DeepSeek
Pick it if: DeepSeek-V4-Flash-0731 is DeepSeek's July 31, 2026 GA release of its efficiency-tier MoE model, holding a 1,048,576-token context window and Terminal-Bench 2.1 score of 82.7. It suits cost-sensitive coding agents and long-context pipelines more than latency-critical chat or confirmed vision workloads.
Its edge: DeepSWE jumped from 7.3 (preview) to 54.4 in the 0731 re-post-train, a 7.5x gain from post-training alone.
The catch: Vision support is reported inconsistently across providers as of the 0731 release; some list image input, others describe it as text-only.
Moonshot AI
Pick it if: Kimi K3 launched July 16, 2026 with 2.8 trillion total parameters (280B active) and a 1,048,576-token context window, the largest announced open-weight-class model to date. It targets teams doing agentic coding and long-document analysis who want an open alternative to GPT-5.6 Sol and Claude Opus 4.8.
Its edge: Highest published open-weight GPQA Diamond score at 93.5%, ahead of Opus 4.8's 91.0%.
The catch: Hallucination rate rose to 51% on AA-Omniscience versus 39% for predecessor K2.6, more confident fabrication despite higher accuracy.
| Input / 1M tokens | $0.14 | $3 |
|---|---|---|
| Output / 1M tokens | $0.28 | $15 |
| Cached input / 1M | — | $0.3 |
| Blended cost (3:1) | $0.175 | $6 |
| Free tier | false | false |
| Pricing model | per-token | per-token |
| Strengths | DeepSWE jumped from 7.3 (preview) to 54.4 in the 0731 re-post-train, a 7.5x gain from post-training alone.; 1,048,576-token context window at $0.14 per 1M input tokens, among the cheapest long-context models available.; Terminal-Bench 2.1 r | Highest published open-weight GPQA Diamond score at 93.5%, ahead of Opus 4.8's 91.0%.; Terminal-Bench 2.1 score of 88.3%, half a point behind GPT-5.6 Sol and ahead of every other open or closed model tested.; Program Bench 77.8% and SWE Mar |
|---|---|---|
| Limitations | Vision support is reported inconsistently across providers as of the 0731 release; some list image input, others describe it as text-only.; No published system card or disclosed red-team partners for this specific build.; Only DeepSeek's ow | Hallucination rate rose to 51% on AA-Omniscience versus 39% for predecessor K2.6, more confident fabrication despite higher accuracy.; Committed weight release (Modified MIT) not yet live on Hugging Face as of this writing, expected by July |
| Context window | 1M | 1M |
|---|---|---|
| Max output tokens | 65.5K | 128K |
| Input modalities | text; tool-calls | text; image |
| Output modalities | text; tool-calls | text; tool-calls; code |
| Capabilities | Tool use; Function calling; Structured output | Vision; Tool use; Code execution |
| Reasoning modes | standard | none; low; medium |
| Long-context recall | unverified | high |
| Openness | open-source | open-weights |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.