Choosing a model for your use case does not always mean choosing the highest score. Rule out what you cannot use, weight what you care about, and read why each model lands where it does. The ranking is deterministic: the same weights give the same list.
Open weights only · must not train on your data · a maximum blended price · a required region. These exclude models outright; only GA models are considered.
Four weights: Capability (SWE-bench Verified and GPQA Diamond, whichever the model publishes.); Price (Blended $ per 1M tokens at 3:1, lower is better, log-scaled.); Speed (Output tokens per second as cited from Artificial Analysis.); Data terms (Does not train on your data, documents a zero-data-retention option, states a retention period.). A model that does not publish a weighted figure is not ranked; set that weight to 0 to include it.
85 GA models pass the default constraints; 35 publish every weighted figure and are ranked here under the default weights (Capability 50 · Price 50 · Speed 0 · Data terms 25). Top ten:
| # | Model | Creator | Score | Why it sits here |
|---|---|---|---|---|
| 1 | Solar Pro 4 | Upstage | 76.5 | Rank 1 of 35: strong on price (rank 7 of 35). |
| 2 | Gemini 3.1 Flash-Lite | Google DeepMind | 74.4 | Rank 2 of 35: strong on price (rank 9 of 35). |
| 3 | Muse Spark 1.3 | Meta AI | 72.3 | Rank 3 of 35: strong on capability (rank 5 of 35). |
| 4 | Grok 4.20 | xAI | 72 | Rank 4 of 35: strong on data terms (rank 1 of 35). |
| 5 | Claude Opus 5 | Anthropic | 71.4 | Rank 5 of 35: strong on capability (rank 1 of 35) and data terms (rank 1 of 35), weaker on price (rank 26 of 35). |
| 6 | Claude Sonnet 5 | Anthropic | 69.5 | Rank 6 of 35: strong on capability (rank 11 of 35) and data terms (rank 1 of 35). |
| 7 | Phi-4 | Microsoft | 69.4 | Rank 7 of 35: strong on price (rank 1 of 35) and data terms (rank 1 of 35), weaker on capability (rank 33 of 35). |
| 8 | Ministral 3 14B | Mistral AI | 68.4 | Rank 8 of 35: strong on price (rank 5 of 35), weaker on capability (rank 26 of 35). |
| 9 | Seed 2.0 Pro | ByteDance | 68.1 | Rank 9 of 35: mid-pack on every weighted axis. |
| 10 | Claude Opus 4.8 | Anthropic | 68 | Rank 10 of 35: strong on capability (rank 9 of 35) and data terms (rank 1 of 35), weaker on price (rank 26 of 35). |
Not ranked because a weighted figure is not published: Bosun (no capability or price); Claude Fable 5.1 (no capability); Cohere Command A+ (no price); Cohere Parse 5 (no capability or price); DiffusionGemma 26B-A4B (no price); FLUX.2 (no capability or price); GLM-5.3-Flash (no capability); GPT Image 2 (no capability); GPT-5.6-Cyber (no capability); Gemini 3 Pro Image (no capability); Gemini 3.6 Flash (no capability); Gemini 3.7 Flash (no capability); Gemini 3.8 Flash (no capability); Gemini Nano (no capability); Gemini Omni 1.1 Flash (no capability or data terms); Gemma 4 12B (no price); Gemma 4 31B (no price); Grok 4.3 (no capability or data terms); Grok 4.5 (no capability); Grok 4.6 (no capability); Grok Imagine Image 2.0 (no capability or price); Kimi K2.7-Code (no capability); Kling 3.0 (no capability or price); LFM2.5-VL-3B (no capability or price); LTX-2.5 (no capability or price); Ling 3.0 Flash Fin (no capability); MAI-Code-1.1-Flash (no data terms); MAI-Image-2.5 (no capability); MiniMax H3 (no capability or price); MiniMax M2.7 (no capability); MiniMax M3 (no capability); MiniMax Music 3.0 (no capability or price); Ministral 3 3B (no capability); Mistral OCR 4 (no capability or price); Muse Glimmer (no price); North Micro Vision Instruct (no capability); Osmosis-Structure-0.6B (no capability or price or data terms); Pulsar 16B (no price); Qwen3.6-27B (no price); Qwen3.7-Plus (no capability); Qwen3.8-27B (no price); Qwen3.8-Max (no capability); Recraft V3 (no capability or price); Recraft V4.1 (no capability or price); Riverflow 2.0 Pro (no capability or price); Seed 2.1 Turbo (no capability); SeedRealtime (no capability or price); Seedance 2.5 (no capability or price or data terms); Wan-Animate-2 (no capability or price); WeatherNext 3 (no capability or price).
Smart Match asks about your team and budget and returns a ranked stack with reasons. Open Smart Match · Full leaderboard · All models