Side-by-side comparison of Claude Sonnet 5, Grok 4.5: pricing, capabilities, integrations and compliance — from verified HokAI records.
Anthropic
Pick it if: Claude Sonnet 5 is Anthropic's mid-tier flagship (released June 30, 2026) with a 1M-token context window and 82.1% SWE-bench Verified, the first model to clear 80% on that benchmark. Priced at $3 input / $15 output per 1M tokens (40% below Opus 4.8), it defaults to adaptive thinking and posts 96.2% on GPQA Diamond and 81.2% on OSWorld-Verified computer use.
Its edge: First model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).
The catch: No native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.
xAI
Pick it if: Grok 4.5 replaces Grok 4.3 as xAI's coding-focused model, fitting teams running high-volume agentic coding who want low cost per resolved task. It averages 15,954 output tokens per SWE-bench Pro task, well under Claude Opus 4.8's 67,020, at faster output speed than xAI's own prior release.
Its edge: Beats GPT-5.5 head-to-head on SWE-bench Pro (64.7% vs 58.6%), the one coding benchmark xAI chose to publish at launch.
The catch: No GPQA Diamond, AIME 2025, MMLU-Pro, or ARC-AGI 2 scores published at launch, so pure academic-reasoning performance is unverified.
| Input / 1M tokens | $3 | $2 |
|---|---|---|
| Output / 1M tokens | $15 | $6 |
| Cached input / 1M | $0.3 | $0.5 |
| Blended cost (3:1) | $6 | $3 |
| Free tier | true | false |
| Pricing model | per-token | per-token |
| Strengths | First model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).; Computer use jumped to 81.2% on OSWorld-Verified and 80.4% on Terminal-Bench 2.1, up from 78.5% and 67.0% on Sonnet 4.6.; Priced a | Beats GPT-5.5 head-to-head on SWE-bench Pro (64.7% vs 58.6%), the one coding benchmark xAI chose to publish at launch.; Resolves the average coding-agent task in far fewer output tokens than Claude Opus 4.8, which compounds into lower real- |
|---|---|---|
| Limitations | No native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.; New tokenizer produces roughly 30% more tokens for the same text than Sonnet 4.6, inflating raw token counts and requiring max_t | No GPQA Diamond, AIME 2025, MMLU-Pro, or ARC-AGI 2 scores published at launch, so pure academic-reasoning performance is unverified.; Context window is the smallest of xAI's current model line-up, a real regression for teams with workflows |
| Context window | 1M | 500K |
|---|---|---|
| Max output tokens | 128K | — |
| Input modalities | text; image; pdf | text; image; tool-calls |
| Output modalities | text; tool-calls | text; tool-calls |
| Capabilities | Vision; Tool use; Web browsing | Vision; Tool use; Web browsing |
| Reasoning modes | adaptive-thinking | low; medium; high |
| Long-context recall | high | unverified |
| Openness | proprietary | proprietary |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.