GPT-5.5 Pro review, pricing and limits

GPT-5.5 Pro is the highest-accuracy inference setting in the GPT-5.5 family, using parallel test-time compute on the same weights for correctness-critical work, at a 6x price premium over standard GPT-5.5.

  • ga
  • proprietary
  • multimodal
  • GPT-5.5 family
checked

GPT-5.5 Pro is for teams escalating their hardest code reviews and multi-file refactors past standard GPT-5.5, Gemini 3.1 Pro, or Claude Sonnet 4.6, not for everyday chat or latency-sensitive apps. OpenAI shaped its safety posture with roughly 200 early-access partners before release, positioning Pro strictly for correctness-critical work where a wrong answer is expensive.

OpenAI's GPT-5.5 Pro is the GPT-5.5 family's highest-accuracy inference setting, released April 24, 2026, with a 1,000,000-token context window and 128,000-token max output. It runs the same Mixture-of-Experts weights as standard GPT-5.5 but adds parallel test-time compute, exploring multiple reasoning paths before answering for correctness-critical work.

Where it sits

  • $67.50/M$ per 1M tokensBlended price (3:1)Lower is better#64 / 64peer median $1.70/Mvendor price, checked by HokAI
  • --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
  • 88.7%% solvedSWE-bench VerifiedHigher is better#4 / 28peer median 78.3%per source, see benchmark scores
  • --% correctGPQA DiamondHigher is better-- / 44peer median 88.3%per source, see benchmark scores

Pricier than 100% of the 64 GA models with a published price, in the top third on SWE-bench Verified (rank 4 of 28), and one of 65 whose vendor states it does not train on customer data.

Ranks are against GA models on HokAI that publish the same figure; ties share a rank.

Provider: OpenAI · Family: GPT-5.5

More about OpenAI on HokAI

Context window: 1,000,000 tokens · Max output: 128,000

Input modalities: text, image, audio, video, tool-calls · Output: text, tool-calls

About GPT-5.5 Pro

GPT-5.5 Pro is OpenAI's highest-accuracy inference setting within the GPT-5.5 family, released alongside the standard GPT-5.5 model on April 23, 2026, with API access opening April 24, 2026. It is not a separate set of weights: GPT-5.5 Pro runs the same Mixture-of-Experts transformer as GPT-5.5, with 128 expert routing groups and sparse activation per token, but applies parallel test-time compute so the model explores multiple reasoning paths before answering. It sits at the top of the GPT-5.5 lineup, above the standard GPT-5.5 model and GPT-5.5 Instant, and is positioned for correctness-critical work rather than everyday chat.

GPT-5.5 Pro shares GPT-5.5's headline benchmark results across the board, since Pro is a different inference setting on identical weights rather than a separate model. GPT-5.5 posts 82.7% on Terminal-Bench 2.0, ahead of Gemini 3.1 Pro's 68.5% and Claude's roughly 65%, and 51.7% on FrontierMath Tiers 1-3 with 35.4% on Tier 4. On Humanity's Last Exam, GPT-5.5 scores 41.4%, behind Claude Opus 4.7's 46.9% and Gemini 3.1 Pro's 44.4%, indicating raw academic-recall reasoning is not where the Pro premium pays off most. Against Gemini 3.1 Pro (80.6% SWE-bench Verified) and Claude Sonnet 4.6 (79.6%), GPT-5.5 Pro's coding lead is the clearest differentiator. OpenAI has not independently published separate GPQA Diamond or AIME 2025 scores for the Pro inference setting versus the standard model.

The API context window is 1,000,000 tokens with a maximum output of 128,000 tokens per completion. The Codex product caps GPT-5.5 (including Pro) at 400,000 tokens, so sessions needing the full 1M window must go through the Responses API directly. A context surcharge applies once a session's input exceeds 272,000 tokens: the entire session is billed at 2x input and 1.5x output rates, not just the portion over the threshold, which matters for retrieval-heavy workloads.

GPT-5.5 Pro processes text, image, audio, and video inputs in a single unified architecture rather than routing between separate stitched-together models. Vision input preserves up to 10,240,000 pixels or a 6,000-pixel dimension without resizing, which improves chart and document reading plus computer-use accuracy. Audio input supports transcription and translation across dozens of languages, but output remains text-only; there is no native audio-out in the chat completions or Responses API. Tool use, function calling, structured outputs, parallel tool calls, web browsing, and code execution via tools all carry over unchanged from GPT-5.5.

GPT-5.5 Pro is available through the OpenAI API with an API key, and reached general availability on Amazon Bedrock on June 1, 2026 as part of a $50 billion AWS-OpenAI partnership announced April 28, 2026 that ended OpenAI's prior Azure exclusivity; Bedrock pricing matches OpenAI's first-party rates with no markup. It also remains available through Azure OpenAI Service. Inside ChatGPT, GPT-5.5 Pro is restricted to Pro, Business, and Enterprise plans; Free and Plus users do not see it as a selectable model. It carries a steep price premium over standard GPT-5.5 given the added test-time compute (see pricing below for exact rates).

The GPT-5.5 system card was updated on April 24, 2026 to cover API deployment safeguards for both GPT-5.5 and the Pro inference setting, with separate evaluations noted where the parallel test-time compute setting could materially change risk posture. OpenAI ran its full pre-deployment safety evaluation suite and Preparedness Framework process, including targeted red-teaming for cybersecurity and biology uplift, and incorporated feedback from roughly 200 early-access partners ahead of release. The model's safety posture is balanced: it refuses clear-harm requests but is not unusually restrictive for legitimate technical or research use.

GPT-5.5 Pro suits teams that need the highest achievable accuracy on a specific hard problem and can absorb the cost and latency: scientific or legal document analysis and one-off research questions where a wrong answer is expensive. Latency-sensitive chat interfaces or high-volume customer support are a weaker match, since a cheaper model clears the accuracy bar there at a fraction of the cost.

GPT-5.5 Pro inherits GPT-5.5's data governance: API inputs are not used for training by default, with an enterprise zero-retention option available. OpenAI states SOC 2 Type II compliance, GDPR compliance, and HIPAA eligibility, with data residency options in the US and EU. Under the EU AI Act, GPT-5.5 (and therefore Pro) is classified as a general-purpose AI model with systemic risk obligations. OpenAI has not independently disclosed a training data cutoff date for GPT-5.5 or GPT-5.5 Pro.

GPT-5.5 Pro launched as part of the same release wave as standard GPT-5.5 (April 23, 2026) and GPT-5.5 Instant (May 5, 2026), replacing GPT-5.4 and GPT-5.2 Pro as OpenAI's top-accuracy offering; GPT-5.2 models, including GPT-5.2 Pro, were fully deprecated from ChatGPT by June 12, 2026, with existing conversations auto-migrating to GPT-5.5. Rumors as of late May 2026 point to a GPT-5.6 release later in the year with an even larger context window, though OpenAI has not confirmed a date.

Pricing

GPT-5.5 Pro bills at six times standard GPT-5.5's per-token rate. Cached input is estimated at $3 per 1M tokens by applying GPT-5.5's published 90% prompt-caching discount ratio; OpenAI has not independently confirmed a Pro-specific cached rate. Available only on ChatGPT Pro/Business/Enterprise plans and via the API.

What a real job costs

JobInputOutputTotal
Summarise a 20-page PDF$0.900$0.180$1.08
Support reply$0.060$0.054$0.114
One coding agent run$6.00$3.60$9.60

Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.

Key Features

  • Parallel Test-Time Compute: GPT-5.5 Pro explores multiple reasoning paths before answering, OpenAI's mechanism for squeezing extra accuracy out of the same GPT-5.5 weights on hard questions.
  • 1M-Token Context Window: Up to 1M input tokens and 128K output tokens via the Responses API, enough for whole-codebase or large-filing analysis in one call.
  • Matches GPT-5.5's Benchmark-Topping SWE-bench Score: Resolves the large majority of real GitHub issues end to end in third-party testing, the same headline coding result published for standard GPT-5.5.
  • Unified Multimodal Input: Text, image, audio, and video processed in one architecture, with vision preserving up to 10.24MP without resizing for accurate chart and document reading.
  • AWS Bedrock General Availability: Reached general availability on Amazon Bedrock in mid-2026 at OpenAI's first-party pricing, part of a major AWS-OpenAI partnership that ended prior Azure exclusivity.

Pros

  • Leads on agentic coding, matching GPT-5.5's benchmark-topping SWE-bench Verified score, ahead of Gemini 3.1 Pro and Claude Sonnet 4.6.
  • 1M-token context window for whole-codebase or large-document work in a single call.
  • No new vendor contract needed for teams already standardized on OpenAI, AWS, or Azure, since Pro runs on all three at matching rates.

Cons

  • Costs several times more than standard GPT-5.5, with no independently published benchmark showing the premium buys extra accuracy on most tasks.
  • No native audio output despite native audio input, and no published tokens-per-second figure for the Pro setting.
  • Locked out of ChatGPT Free and Plus plans; only Pro, Business, and Enterprise tiers can select it in the chat UI.

Benchmarks

  • MMLU: 92.4% vendor-reported · 23 Apr 2026 — General-knowledge exam across 57 subjects, % correct.
  • SWE-bench Verified: 88.7% vendor-reported · 23 Apr 2026 — Real GitHub issues fixed end to end, % solved.

A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.

Frequently Asked Questions

How much do you pay for GPT-5.5 Pro?

OpenAI charges $30 per 1M input tokens and $180 per 1M output tokens for GPT-5.5 Pro, six times standard GPT-5.5's $5/$30 rate. There's no independently confirmed Pro-specific cached-input price, though applying GPT-5.5's 90% caching discount would land it around $3 per 1M tokens. Access runs through the OpenAI API, AWS Bedrock, and Azure OpenAI Service at matching rates, or a ChatGPT Pro, Business, or Enterprise plan.

Does GPT-5.5 Pro have a free plan?

No. GPT-5.5 Pro is available only through paid API access on OpenAI, AWS Bedrock, or Azure OpenAI Service, or via a ChatGPT Pro, Business, or Enterprise subscription; the Free and Plus ChatGPT tiers can't select it. Standard GPT-5.5 or GPT-5.5 Instant cost less for anyone who doesn't need the Pro accuracy ceiling.

What are GPT-5.5 Pro's closest competitors?

The closest alternatives are Gemini 3.1 Pro, a lower-cost pick when the coding gap isn't worth the premium, and Claude Sonnet 4.6, a cheaper agentic-coding baseline for routine PRs. Claude Opus 4.7 is the better choice for deep academic-reasoning tasks rather than software engineering. Standard GPT-5.5 covers everything short of the hardest problems, at a sixth of Pro's price.

GPT-5.5 Pro or Gemini 3.1 Pro: which should you pick?

GPT-5.5 Pro's clearest edge over Gemini 3.1 Pro is coding: 88.7% vs 80.6% on SWE-bench Verified and 82.7% vs 68.5% on Terminal-Bench 2.0. The comparison flips on raw academic reasoning, where GPT-5.5 trails Gemini 3.1 Pro's 44.4% with a 41.4% score on Humanity's Last Exam. Gemini 3.1 Pro is the better pick for research-heavy work or tighter budgets; GPT-5.5 Pro wins for agentic coding regardless of cost.

What does it take to start using GPT-5.5 Pro?

Get an OpenAI API key and call the model string gpt-5.5-pro through the Responses API, or select it directly on a ChatGPT Pro, Business, or Enterprise plan. Enterprises already on AWS Bedrock or Azure OpenAI Service can call it there at matching first-party rates without a new vendor contract. Because Pro costs meaningfully more per token than standard GPT-5.5, route only escalations that fail a confidence check to it rather than defaulting to it.

Top Alternatives

  • GPT-5.5: GPT-5.5 Pro earns its price on a single hard problem where the accuracy ceiling matters. Standard GPT-5.5 costs a sixth as much and covers everyday work just as well.
  • Gemini 3.1 Pro: GPT-5.5 Pro holds a real SWE-bench Verified lead over Gemini 3.1 Pro. Gemini 3.1 Pro is the one to reach for when that coding gap isn't worth its premium.
  • Claude Sonnet 4.6: GPT-5.5 Pro's SWE-bench Verified gap over Claude Sonnet 4.6 is the widest of any comparison here. Sonnet 4.6 answers back with a cheaper baseline for routine agentic-coding work.
  • Claude Opus 4.7: Claude Opus 4.7 is the stronger choice for deep academic reasoning on Humanity's Last Exam. GPT-5.5 Pro's strength lies in agentic coding instead, where the comparison flips.

More AI Models on HokAI

Visit GPT-5.5 Pro Official Page