Side-by-side pricing, features and compliance for any tools in the directory.
Side-by-side comparison of Claude Sonnet 5, Grok 4.5: pricing, capabilities, integrations and compliance — from verified HokAI records.
Anthropic
Pick it if: Sonnet 5 fits engineering teams running high-volume agentic coding and computer-use work who want near-Opus quality at Sonnet pricing. It's the wrong pick for voice or video-first products, since it has no native audio or video I/O, and for configs still pinned to non-default sampling parameters. Max output runs to 128,000 tokens on the synchronous API.
Its edge: First model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).
The catch: No native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.
xAI
Pick it if: Grok 4.5 replaces Grok 4.3 as xAI's coding-focused model, fitting teams running high-volume agentic coding who want low cost per resolved task. It averages 15,954 output tokens per SWE-bench Pro task, well under Claude Opus 4.8's 67,020, at faster output speed than xAI's own prior release.
Its edge: Beats GPT-5.5 head-to-head on SWE-bench Pro (64.7% vs 58.6%), the one coding benchmark xAI chose to publish at launch.
The catch: No GPQA Diamond, AIME 2025, MMLU-Pro, or ARC-AGI 2 scores published at launch, so pure academic-reasoning performance is unverified.
| Input / 1M tokens | $3 | $2 |
|---|---|---|
| Output / 1M tokens | $15 | $6 |
| Cached input / 1M | $0.3 | $0.5 |
| Blended cost (3:1) | $6 | $3 |
| Summarise a 20-page PDF | $0.105 | $0.066 |
| Support reply | $0.010 | $0.0058 |
| One coding agent run | $0.900 | $0.520 |
| Free tier | true | false |
| Strengths | First model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).; Computer use jumped to 81.2% on OSWorld-Verified and 80.4% on Terminal-Bench 2.1, up from 78.5% and 67.0% on Sonnet 4.6.; Priced a | Beats GPT-5.5 head-to-head on SWE-bench Pro (64.7% vs 58.6%), the one coding benchmark xAI chose to publish at launch.; Resolves the average coding-agent task in far fewer output tokens than Claude Opus 4.8, which compounds into lower real- |
|---|---|---|
| Limitations | No native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.; New tokenizer produces roughly 30% more tokens for the same text than Sonnet 4.6, inflating raw token counts and requiring max_t | No GPQA Diamond, AIME 2025, MMLU-Pro, or ARC-AGI 2 scores published at launch, so pure academic-reasoning performance is unverified.; Context window is the smallest of xAI's current model line-up, a real regression for teams with workflows |
| Context window | 1M | 500K |
|---|---|---|
| Max output tokens | 128K | -- |
| Input modalities | text; image; pdf | text; image; tool-calls |
| Output modalities | text; tool-calls | text; tool-calls |
| Capabilities | Vision; Tool use; Web browsing | Vision; Tool use; Web browsing |
| Reasoning modes | adaptive-thinking | low; medium; high |
| Long-context recall | high | unverified |
| Openness | proprietary | proprietary |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.