GPT-5.6 Sol review, pricing and limits

OpenAI's flagship GPT-5.6 model for maximum capability — coding, agentic work, science, and cybersecurity with ultra multi-agent coordination and 2x token efficiency.

  • ga
  • proprietary
  • multimodal
  • GPT-5.6 family
checked

Sol replaces manual multi-agent orchestration for teams running the hardest coding and research workloads, not everyday chat. It's the default model inside the enterprise Copilot suite, so most business users already run it without choosing it directly, while cost-sensitive teams should route to a cheaper sibling model instead.

Sol is OpenAI's flagship model, released July 9, 2026, ranking #1 on LMArena Chatbot Arena and scoring 78.5% on SWE-bench Verified. It adds ultra multi-agent coordination and programmatic tool calling that its cheaper siblings, Terra and Luna, simply don't get access to.

Where it sits

  • $11.25/M$ per 1M tokensBlended price (3:1)Lower is better#55 / 64peer median $1.70/Mvendor price, checked by HokAI
  • 75 tok/stokens/sOutput speedHigher is better#24 / 39peer median 90 tok/scited: Artificial Analysis
  • 78.5%% solvedSWE-bench VerifiedHigher is better#14 / 28peer median 78.3%per source, see benchmark scores
  • 72.1%% correctGPQA DiamondHigher is better#36 / 44peer median 88.3%per source, see benchmark scores

Pricier than 90% of the 64 GA models with a published price, mid-pack on SWE-bench Verified (rank 14 of 28), and one of 22 that document a zero-data-retention option.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: OpenAI · Family: GPT-5.6

More about OpenAI on HokAI

Context window: 200,000 tokens · Max output: 64,000

Input modalities: text, image · Output: text, tool-calls, code

About GPT-5.6 Sol

GPT-5.6 Sol tops OpenAI's GPT-5.6 lineup, reaching general availability on July 9, 2026 after a limited preview that started June 26, 2026. It sets a new bar for both intelligence and efficiency, leading results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models with fewer tokens and at lower estimated cost. The model delivers stronger performance per dollar: more successful work for the same spend, or comparable results at lower total cost.

A key innovation is ultra mode, OpenAI's highest-capability setting that coordinates multiple agents across parallel workstreams to finish complex tasks faster. This multi-agent coordination, similar to the Codex ultra mode, reduces wall-clock time and improves performance for complex tasks that divide cleanly into independent workstreams. Sol also features stronger computer use and design judgment, making it the most polished collaborator for inspecting, refining, and delivering ready-to-use results, particularly for frontend development with improved layout, visual hierarchy, and design aesthetics.

The model introduces programmatic tool calling (PTC), allowing it to write JavaScript to call eligible tools, pass results between calls, and process intermediate outputs in a hosted runtime. PTC is ZDR-compatible with no additional container costs, ideal for bounded, tool-heavy workflows. Explicit prompt caching with cache breakpoints and a half-hour minimum cache life provides predictable caching economics, with cache reads retaining a large discount over uncached input. Persisted reasoning reuses reasoning items across turns via reasoning.context for multi-turn quality and cache efficiency.

Reasoning effort supports none, low, medium, high, xhigh, max, plus a pro mode that performs more model work for reliability on difficult tasks, returning a single final answer. Max reasoning effort is reserved for the hardest quality-first workloads. Sol's exact per-token numbers are covered in the pricing FAQ, not repeated here. Sol is available on the OpenAI API, ChatGPT, Codex, Microsoft 365 Copilot, Cerebras, and additional gateways; its training data runs through June 2026.

Pricing

Input $5/M, Output $30/M, Cached input $0.625/M (1.25x base). Cache writes at 1.25x, reads at 90% discount. Explicit cache breakpoints supported. No batch discount at launch. Ultra/multi-agent usage billed at standard token rates. Programmatic tool calling billed at standard rates. Server-side tools (if any) billed separately. Cerebras inference pricing separate.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.150$0.030$0.180
Support reply$0.010$0.0090$0.019
One coding agent run$1.00$0.600$1.60

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Ultra Multi-Agent Coordination: Coordinates multiple subagents in parallel across independent workstreams, cutting wall-clock time on complex tasks by 40-60%.
  • Programmatic Tool Calling (PTC): The model writes JavaScript to call tools, pass results, and process intermediates in a hosted runtime with no added container costs.
  • Explicit Prompt Caching with Breakpoints: Developers mark exact cacheable prefixes and get a heavy discount on cached reads, with a half-hour minimum cache life.
  • Pro Reasoning Mode: Performs extended model work for reliability on the hardest tasks, then returns a single final answer instead of a running trace.
  • 2x Token Efficiency: Reaches frontier performance using roughly half the output tokens of comparable models, which directly lowers cost and latency.

Pros

  • Frontier intelligence (LMArena rank 1) paired with 2x token efficiency gives the best performance per dollar at this capability tier.
  • Ultra multi-agent mode is a genuine workflow accelerator for parallelizable engineering tasks.
  • Programmatic tool calling eliminates turn overhead for tool-heavy loops and stays ZDR compatible.
  • Being the default enterprise Copilot model gives it instant, massive seat distribution and trust.

Cons

  • Premium pricing excludes high-volume and cost-sensitive workloads that don't need frontier capability.
  • No native audio or video modality; vision only, lagging Gemini 3.1 on multimodal breadth.
  • Pro mode and ultra mode add significant latency and token overhead compared to standard mode.
  • Multi-agent beta still has coordination inefficiencies, including occasionally overlapping sub-tasks.

Benchmarks

  • MATH: 85.6% vendor-reported · 26 Jun 2026 — Competition maths problems, % solved.
  • MMLU: 92.8% vendor-reported · 26 Jun 2026 — General-knowledge exam across 57 subjects, % correct.
  • MMLU-Pro: 84.7% vendor-reported · 26 Jun 2026 — A harder version of the 57-subject knowledge exam, % correct.
  • AIME 2025: 88.3% vendor-reported · 26 Jun 2026 — Competition-level maths problems from the 2025 exam, % solved.
  • ARC-AGI 2: 28.4% vendor-reported · 26 Jun 2026 — Abstract visual puzzles built to resist memorisation, % solved.
  • HumanEval: 96.2% vendor-reported · 26 Jun 2026 — Small programs that must pass hidden tests, % passing.
  • LiveBench: 68.9% vendor-reported · 26 Jun 2026 — A rolling set of fresh questions that cannot have been in training data, % correct.
  • LMArena Elo: 1415 independent · 12 Jul 2026 — Rating from blind human votes on which answer is better.
  • GPQA Diamond: 72.1% vendor-reported · 26 Jun 2026 — PhD-level science questions that are hard to search for, % correct.
  • LMArena rank: #1 independent · 12 Jul 2026 — Position on the blind human-preference leaderboard; #1 is best.
  • Aider Polyglot: 76.8% vendor-reported · 26 Jun 2026 — Code edits across several programming languages, % correct.
  • SWE-bench Verified: 78.5% vendor-reported · 26 Jun 2026 — Real GitHub issues fixed end to end, % solved.
  • Humanity's Last Exam: 22.1% vendor-reported · 26 Jun 2026 — Expert-written questions across many fields, % correct.
  • AA Intelligence Index: 62 cited: Artificial Analysis · 10 Jul 2026 — Composite of 10 evaluations run by Artificial Analysis, 0 to 100.
  • AA blended price: $17.50/M cited: Artificial Analysis · 10 Jul 2026 — Price per 1M tokens at a 3:1 input to output blend, as listed by Artificial Analysis.
  • Output speed: 75 tok/s cited: Artificial Analysis · 10 Jul 2026 — Median tokens written per second as measured by Artificial Analysis.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

How much does GPT-5.6 Sol cost per 1M tokens?

Sol lists at $5/M for input and $30/M for output, with cache reads dropping to $0.625/M (a 90% discount) and cache writes at 1.25x the base input rate. There's no batch discount at launch, and Cerebras inference is priced separately.

How does GPT-5.6 Sol compare on benchmarks vs GPT-5.6 Terra?

Sol scores 78.5% on SWE-bench Verified and 72.1% on GPQA Diamond, ahead of Terra on both. Sol also adds ultra multi-agent mode, programmatic tool calling, and a pro reasoning mode that Terra doesn't have, at roughly double Terra's per-token price.

Are GPT-5.6 Sol's weights open, or is it closed and API-only?

Sol is proprietary: OpenAI keeps the weights closed, with no license to run it yourself. Access runs only through OpenAI's own products and a short list of gateway partners.

Does GPT-5.6 Sol train on user data?

By default OpenAI doesn't train on API inputs or outputs, and Sol is Zero Data Retention eligible for enterprise customers. It's SOC2 Type II, ISO 27001, GDPR, and HIPAA eligible, with US and EU data residency options.

What is GPT-5.6 Sol good at, and where does it fall short?

Sol fits senior engineering teams running frontier coding, cybersecurity research, and complex multi-agent workflows where quality matters more than cost. Avoid it for high-volume, low-margin inference or simple conversational tasks; route those to Terra or Luna instead.

HokAI guides covering GPT-5.6 Sol

More AI Models on HokAI

Visit GPT-5.6 Sol Official Page