Side-by-side pricing, features and compliance for any tools in the directory.
Anthropic
Pick it if: Claude Opus 5 launched July 24, 2026 with a 1 million token context window as the default tier and a new xhigh reasoning effort mode. It replaces Opus 4.8 as Anthropic's flagship Opus model for long-running agentic coding and research workloads.
Its edge: Leads SWE-bench Verified at 97.0%, the highest published score on the benchmark to date.
The catch: Output speed averages about 54 tokens per second, below the 73.7 t/s median for comparable reasoning models.
Google DeepMind
Pick it if: Gemini 3.6 Flash replaces the previous Flash generation as Google's cost-efficient workhorse for agentic coding and desktop automation, finishing tasks with 17% fewer output tokens along the way. Its knowledge cutoff advances a full 14 months to March 2026, making it the default pick for teams that outgrew the old Flash tier but don't need a full Pro-tier model.
Its edge: DeepSWE score jumped from 37% to 49% between Gemini 3.5 Flash and 3.6 Flash, the largest single-generation coding gain in the Flash line.
The catch: Architecture details, including dense vs mixture-of-experts and parameter count, are undisclosed.
| Input / 1M tokens | $5 | $1.50 |
|---|---|---|
| Output / 1M tokens | $25 | $7.50 |
| Cached input / 1M | $0.5 | $0.15 |
| Blended cost (3:1) | $10 | $3 |
| Free tier | false | false |
| Pricing model | per-token | per-token |
| Strengths | Leads SWE-bench Verified at 97.0%, the highest published score on the benchmark to date.; Scores 30.16% on ARC-AGI-3 at high effort, about 20x Opus 4.8 and roughly 4x GPT-5.6 Sol Max.; 1M token context window is both the default and the cei | DeepSWE score jumped from 37% to 49% between Gemini 3.5 Flash and 3.6 Flash, the largest single-generation coding gain in the Flash line.; Supports parallel tool calls and configurable reasoning effort for multi-step agent orchestration.; S |
|---|---|---|
| Limitations | Output speed averages about 54 tokens per second, below the 73.7 t/s median for comparable reasoning models.; Accepts text, image, and PDF input only; no native audio or video input.; UK AI Security Institute testing found Opus 5 completed | Architecture details, including dense vs mixture-of-experts and parameter count, are undisclosed.; Priority-tier latency guarantees cost roughly 1.8x standard pricing.; MLE-Bench score of 63.9% still trails what frontier Pro-tier models rep |
| Context window | 1M | 1M |
|---|---|---|
| Max output tokens | 128K | 64K |
| Input modalities | text; image; pdf | text; image; video |
| Output modalities | text; tool-calls | text |
| Capabilities | Vision; Tool use; Code execution | Vision; Tool use; Audio input |
| Reasoning modes | low; medium; high | standard; configurable-thinking |
| Long-context recall | high | high at 128K tokens, moderate at the full 1M window |
| Openness | proprietary | proprietary |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.