Smart Stack

Side-by-side pricing, features and compliance for any tools in the directory.

DeepSeek-V4-Flash-0731 vs Kimi K3

DeepSeek-V4-Flash-0731

DeepSeek

Pick it if: DeepSeek-V4-Flash-0731 is DeepSeek's July 31, 2026 GA release of its efficiency-tier MoE model, holding a 1,048,576-token context window and Terminal-Bench 2.1 score of 82.7. It suits cost-sensitive coding agents and long-context pipelines more than latency-critical chat or confirmed vision workloads.

Its edge: DeepSWE jumped from 7.3 (preview) to 54.4 in the 0731 re-post-train, a 7.5x gain from post-training alone.

The catch: Vision support is reported inconsistently across providers as of the 0731 release; some list image input, others describe it as text-only.

Kimi K3

Moonshot AI

Pick it if: Kimi K3 launched July 16, 2026 with 2.8 trillion total parameters (280B active) and a 1,048,576-token context window, the largest announced open-weight-class model to date. It targets teams doing agentic coding and long-document analysis who want an open alternative to GPT-5.6 Sol and Claude Opus 4.8.

Its edge: Highest published open-weight GPQA Diamond score at 93.5%, ahead of Opus 4.8's 91.0%.

The catch: Hallucination rate rose to 51% on AA-Omniscience versus 39% for predecessor K2.6, more confident fabrication despite higher accuracy.

Pricing

Input / 1M tokens$0.14$3
Output / 1M tokens$0.28$15
Cached input / 1M$0.3
Blended cost (3:1)$0.175$6
Free tierfalsefalse
Pricing modelper-tokenper-token

Verdict & fit

StrengthsDeepSWE jumped from 7.3 (preview) to 54.4 in the 0731 re-post-train, a 7.5x gain from post-training alone.; 1,048,576-token context window at $0.14 per 1M input tokens, among the cheapest long-context models available.; Terminal-Bench 2.1 rHighest published open-weight GPQA Diamond score at 93.5%, ahead of Opus 4.8's 91.0%.; Terminal-Bench 2.1 score of 88.3%, half a point behind GPT-5.6 Sol and ahead of every other open or closed model tested.; Program Bench 77.8% and SWE Mar
LimitationsVision support is reported inconsistently across providers as of the 0731 release; some list image input, others describe it as text-only.; No published system card or disclosed red-team partners for this specific build.; Only DeepSeek's owHallucination rate rose to 51% on AA-Omniscience versus 39% for predecessor K2.6, more confident fabrication despite higher accuracy.; Committed weight release (Modified MIT) not yet live on Hugging Face as of this writing, expected by July

Capability envelope

Context window1M1M
Max output tokens65.5K128K
Input modalitiestext; tool-callstext; image
Output modalitiestext; tool-callstext; tool-calls; code
CapabilitiesTool use; Function calling; Structured outputVision; Tool use; Code execution
Reasoning modesstandardnone; low; medium
Long-context recallunverifiedhigh
Opennessopen-sourceopen-weights

Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.