Side-by-side pricing, features and compliance for any tools in the directory.
Google DeepMind
Pick it if: Gemini 3.6 Flash replaces the previous Flash generation as Google's cost-efficient workhorse for agentic coding and desktop automation, finishing tasks with 17% fewer output tokens along the way. Its knowledge cutoff advances a full 14 months to March 2026, making it the default pick for teams that outgrew the old Flash tier but don't need a full Pro-tier model.
Its edge: DeepSWE score jumped from 37% to 49% between Gemini 3.5 Flash and 3.6 Flash, the largest single-generation coding gain in the Flash line.
The catch: Architecture details, including dense vs mixture-of-experts and parameter count, are undisclosed.
OpenAI
Pick it if: Sol replaces manual multi-agent orchestration for teams running the hardest coding and research workloads, not everyday chat. It's the default model inside the enterprise Copilot suite, so most business users already run it without choosing it directly, while cost-sensitive teams should route to a cheaper sibling model instead.
Its edge: State-of-the-art across coding, knowledge work, cybersecurity, and science benchmarks with 2x token efficiency vs competing flagships.
The catch: Premium pricing at $5/$30 per 1M tokens — 2.5x Terra, 5x Luna — limits high-volume use cases.
| Input / 1M tokens | $1.50 | $5 |
|---|---|---|
| Output / 1M tokens | $7.50 | $30 |
| Cached input / 1M | $0.15 | $0.625 |
| Blended cost (3:1) | $3 | $11.25 |
| Free tier | false | false |
| Pricing model | per-token | per-token |
| Strengths | DeepSWE score jumped from 37% to 49% between Gemini 3.5 Flash and 3.6 Flash, the largest single-generation coding gain in the Flash line.; Supports parallel tool calls and configurable reasoning effort for multi-step agent orchestration.; S | State-of-the-art across coding, knowledge work, cybersecurity, and science benchmarks with 2x token efficiency vs competing flagships.; Ultra multi-agent mode coordinates parallel subagents for faster complex task completion — unique at thi |
|---|---|---|
| Limitations | Architecture details, including dense vs mixture-of-experts and parameter count, are undisclosed.; Priority-tier latency guarantees cost roughly 1.8x standard pricing.; MLE-Bench score of 63.9% still trails what frontier Pro-tier models rep | Premium pricing at $5/$30 per 1M tokens — 2.5x Terra, 5x Luna — limits high-volume use cases.; Ultra mode and multi-agent beta features may have rough edges during early GA period.; No native audio/video I/O — vision only; multimodal laggin |
| Context window | 1M | 200K |
|---|---|---|
| Max output tokens | 64K | 64K |
| Input modalities | text; image; video | text; image |
| Output modalities | text | text; tool-calls; code |
| Capabilities | Vision; Tool use; Audio input | Vision; Tool use; Code execution |
| Reasoning modes | standard; configurable-thinking | none; low; medium |
| Long-context recall | high at 128K tokens, moderate at the full 1M window | high |
| Openness | proprietary | proprietary |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.