Side-by-side comparison of Claude Sonnet 5, Mistral Large 3: pricing, capabilities, integrations and compliance — from verified HokAI records.
Anthropic
Pick it if: Claude Sonnet 5 is Anthropic's mid-tier flagship (released June 30, 2026) with a 1M-token context window and 82.1% SWE-bench Verified, the first model to clear 80% on that benchmark. Priced at $3 input / $15 output per 1M tokens (40% below Opus 4.8), it defaults to adaptive thinking and posts 96.2% on GPQA Diamond and 81.2% on OSWorld-Verified computer use.
Its edge: First model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).
The catch: No native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.
Mistral AI
Pick it if: Mistral Large 3 suits teams that want a self-hostable, openly licensed alternative to closed frontier models for coding and long-document work, scoring about 92% on HumanEval. It's the wrong pick for graduate-level science reasoning, where newer open-weight reasoning models pull well ahead, and for teams lacking the multi-GPU infrastructure needed to self-host it at scale.
Its edge: Apache 2.0 open weights at frontier scale (675B total, 41B active MoE), with no commercial-use restrictions.
The catch: GPQA Diamond score of 43.9% trails DeepSeek-V3.2 and Kimi K2-Thinking, which score 70-85% on the same benchmark.
| Input / 1M tokens | $3 | $0.5 |
|---|---|---|
| Output / 1M tokens | $15 | $1.50 |
| Cached input / 1M | $0.3 | — |
| Blended cost (3:1) | $6 | $0.75 |
| Free tier | true | false |
| Pricing model | per-token | per-token |
| Strengths | First model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).; Computer use jumped to 81.2% on OSWorld-Verified and 80.4% on Terminal-Bench 2.1, up from 78.5% and 67.0% on Sonnet 4.6.; Priced a | Apache 2.0 open weights at frontier scale (675B total, 41B active MoE), with no commercial-use restrictions.; 256K-token context window for both input and output, double Mistral Medium 3's 128K.; Strong coding performance (~92% HumanEval pa |
|---|---|---|
| Limitations | No native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.; New tokenizer produces roughly 30% more tokens for the same text than Sonnet 4.6, inflating raw token counts and requiring max_t | GPQA Diamond score of 43.9% trails DeepSeek-V3.2 and Kimi K2-Thinking, which score 70-85% on the same benchmark.; Self-hosting requires holding all 675B parameters in VRAM (roughly 355GB at 4-bit, ~710GB at FP16) even though only 41B are ac |
| Context window | 1M | 256K |
|---|---|---|
| Max output tokens | 128K | 256K |
| Input modalities | text; image; pdf | text; image |
| Output modalities | text; tool-calls | text |
| Capabilities | Vision; Tool use; Web browsing | Vision; Tool use; Function calling |
| Reasoning modes | adaptive-thinking | standard |
| Long-context recall | high | high |
| Openness | proprietary | open-source |
Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.