Side-by-side pricing, features and compliance for any tools in the directory.
Anthropic
Pick it if: Claude Opus 5 launched July 24, 2026 with a 1 million token context window as the default tier and a new xhigh reasoning effort mode. It replaces Opus 4.8 as Anthropic's flagship Opus model for long-running agentic coding and research workloads.
Its edge: Leads SWE-bench Verified at 97.0%, the highest published score on the benchmark to date.
The catch: Output speed averages about 54 tokens per second, below the 73.7 t/s median for comparable reasoning models.
Google DeepMind
Pick it if: Gemini 3.6 Flash replaces the previous Flash generation as Google's cost-efficient workhorse for agentic coding and desktop automation, finishing tasks with 17% fewer output tokens along the way. Its knowledge cutoff advances a full 14 months to March 2026, making it the default pick for teams that outgrew the old Flash tier but don't need a full Pro-tier model.
Its edge: DeepSWE score jumped from 37% to 49% between Gemini 3.5 Flash and 3.6 Flash, the largest single-generation coding gain in the Flash line.
The catch: Architecture details, including dense vs mixture-of-experts and parameter count, are undisclosed.
OpenAI
Pick it if: Sol replaces manual multi-agent orchestration for teams running the hardest coding and research workloads, not everyday chat. It's the default model inside the enterprise Copilot suite, so most business users already run it without choosing it directly, while cost-sensitive teams should route to a cheaper sibling model instead.
Its edge: State-of-the-art across coding, knowledge work, cybersecurity, and science benchmarks with 2x token efficiency vs competing flagships.
The catch: Premium pricing at $5/$30 per 1M tokens — 2.5x Terra, 5x Luna — limits high-volume use cases.
| Input / 1M tokens | $5 | $1.50 | $5 |
|---|---|---|---|
| Output / 1M tokens | $25 | $7.50 | $30 |
| Cached input / 1M | $0.5 | $0.15 | $0.625 |
| Blended cost (3:1) | $10 | $3 | $11.25 |
| Free tier | false | false | false |
| Pricing model | per-token | per-token | per-token |
| Strengths | Leads SWE-bench Verified at 97.0%, the highest published score on the benchmark to date.; Scores 30.16% on ARC-AGI-3 at high effort, about 20x Opus 4.8 and roughly 4x GPT-5.6 Sol Max.; 1M token context window is both the default and the cei | DeepSWE score jumped from 37% to 49% between Gemini 3.5 Flash and 3.6 Flash, the largest single-generation coding gain in the Flash line.; Supports parallel tool calls and configurable reasoning effort for multi-step agent orchestration.; S | State-of-the-art across coding, knowledge work, cybersecurity, and science benchmarks with 2x token efficiency vs competing flagships.; Ultra multi-agent mode coordinates parallel subagents for faster complex task completion — unique at thi |
|---|---|---|---|
| Limitations | Output speed averages about 54 tokens per second, below the 73.7 t/s median for comparable reasoning models.; Accepts text, image, and PDF input only; no native audio or video input.; UK AI Security Institute testing found Opus 5 completed | Architecture details, including dense vs mixture-of-experts and parameter count, are undisclosed.; Priority-tier latency guarantees cost roughly 1.8x standard pricing.; MLE-Bench score of 63.9% still trails what frontier Pro-tier models rep | Premium pricing at $5/$30 per 1M tokens — 2.5x Terra, 5x Luna — limits high-volume use cases.; Ultra mode and multi-agent beta features may have rough edges during early GA period.; No native audio/video I/O — vision only; multimodal laggin |
| Context window | 1M | 1M | 200K |
|---|---|---|---|
| Max output tokens | 128K | 64K | 64K |
| Input modalities | text; image; pdf | text; image; video | text; image |
| Output modalities | text; tool-calls | text | text; tool-calls; code |
| Capabilities | Vision; Tool use; Code execution | Vision; Tool use; Audio input | Vision; Tool use; Code execution |
| Reasoning modes | low; medium; high | standard; configurable-thinking | none; low; medium |
| Long-context recall | high | high at 128K tokens, moderate at the full 1M window | high |
| Openness | proprietary | proprietary | proprietary |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.