Smart Stack

Side-by-side pricing, features and compliance for any tools in the directory.

Claude Sonnet 5 vs GPT-5.6 Luna

Claude Sonnet 5

Anthropic

Pick it if: Claude Sonnet 5 is Anthropic's mid-tier flagship (released June 30, 2026) with a 1M-token context window and 82.1% SWE-bench Verified, the first model to clear 80% on that benchmark. Priced at $3 input / $15 output per 1M tokens (40% below Opus 4.8), it defaults to adaptive thinking and posts 96.2% on GPQA Diamond and 81.2% on OSWorld-Verified computer use.

Its edge: First model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).

The catch: No native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.

GPT-5.6 Luna

OpenAI

Pick it if: Luna replaces expensive per-token inference for high-volume, latency-sensitive workloads like classification and routing, not hard reasoning tasks. It's the only model in the family on ChatGPT's free tier, so most casual users meet it first, while teams needing multi-agent orchestration should upgrade to a pricier sibling.

Its edge: Best cost-efficiency in GPT-5.6 family — $1/$6 per 1M tokens is 20% of Sol, 40% of Terra.

The catch: Reasoning capped at high effort — no xhigh, max, or pro mode; ceiling on complex multi-step reasoning.

Pricing

Input / 1M tokens$3$1
Output / 1M tokens$15$6
Cached input / 1M$0.3$0.125
Blended cost (3:1)$6$2.25
Free tiertruetrue
Pricing modelper-tokenper-token

Verdict & fit

StrengthsFirst model to break 80% on SWE-bench Verified at 82.1%, ahead of Gemini 3.1 Pro (80.6%) and GPT-5.4 (~80%).; Computer use jumped to 81.2% on OSWorld-Verified and 80.4% on Terminal-Bench 2.1, up from 78.5% and 67.0% on Sonnet 4.6.; Priced aBest cost-efficiency in GPT-5.6 family — $1/$6 per 1M tokens is 20% of Sol, 40% of Terra.; Fastest latency in family (~800ms p50, ~90 tokens/sec) — ideal for latency-sensitive production paths.; Full vision + text capability with function c
LimitationsNo native audio or video input/output; text and image only, requiring a separate ASR/TTS stack for voice apps.; New tokenizer produces roughly 30% more tokens for the same text than Sonnet 4.6, inflating raw token counts and requiring max_tReasoning capped at high effort — no xhigh, max, or pro mode; ceiling on complex multi-step reasoning.; No ultra multi-agent, no programmatic tool calling, no explicit cache breakpoints (auto caching only).; SWE-bench ~60%, GPQA ~55% — sign

Capability envelope

Context window1M200K
Max output tokens128K64K
Input modalitiestext; image; pdftext; image
Output modalitiestext; tool-callstext; tool-calls; code
CapabilitiesVision; Tool use; Web browsingVision; Tool use; Code execution
Reasoning modesadaptive-thinkingnone; low; medium
Long-context recallhighmedium
Opennessproprietaryproprietary

Compare up to four at a time, or run Smart Match to get a ranked shortlist. Browse the full AI directory.