GPT-5.5 replaces separate ASR, vision, and text pipelines with a single model call, scoring 82.7% on Terminal-Bench for real agentic task completion. It suits multimodal agentic workflows and long-document research, but pure text-only, cost-sensitive workloads and voice-first apps needing low latency are better served elsewhere: OpenAI's Realtime API for voice, a cheaper text model for high-volume plain-text use.
GPT-5.5 is OpenAI's flagship multimodal model, released April 23, 2026, with a 1M-token context window through the API and built-in agentic computer use via Codex. In one session it handles written, visual, spoken, and video input, replacing the need to stitch separate ASR, vision, and generation models behind an agent.
Where it sits
- $11.25/M$ per 1M tokensBlended price (3:1)Lower is better#55 / 64peer median $1.70/Mvendor price, checked by HokAI
- 61 tok/stokens/sOutput speedHigher is better#27 / 39peer median 90 tok/scited: Artificial Analysis
- 88.7%% solvedSWE-bench VerifiedHigher is better#4 / 28peer median 78.3%per source, see benchmark scores
- 93.6%% correctGPQA DiamondHigher is better#8 / 44peer median 88.3%per source, see benchmark scores
Pricier than 90% of the 64 GA models with a published price, in the top third on SWE-bench Verified (rank 4 of 28), and one of 65 whose vendor states it does not train on customer data.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: OpenAI · Family: GPT-5
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, audio, video, tool-calls · Output: text, tool-calls
About GPT-5.5
GPT-5.5, codename Spud, is OpenAI's sixth-generation flagship multimodal language model, released April 23, 2026, with API access opening April 24. It is the first OpenAI model to process text, images, audio, and video in a single native architecture, removing the prior requirement to stitch GPT-5, Whisper, and Sora behind an agent. The model sits above GPT-5.4 in OpenAI's lineup and targets complex professional work: agentic coding, long-document analysis, multi-modal research, and cross-tool automation. OpenAI has not disclosed the parameter count; the architecture follows the Mixture-of-Experts lineage used in GPT-5, with 128 expert routing and sparse activation per token.
GPT-5.5 scores 88.7% on SWE-bench Verified, the agentic software engineering benchmark measuring real bug-fix success in production codebases. On MMLU it reaches 92.4%. For visual reasoning, it posts 92.1% on ChartQA and 96.2% on AI2D (infographic understanding). Terminal-Bench, measuring agentic task completion in a real terminal environment, comes in at 82.7%. BenchLM ranks GPT-5.5 fourth globally among 119 models as of June 2026 with an overall score of 91/100. GPT-5.5 Instant (the lighter default variant) scored 81.2% on AIME 2025, compared to 65.4% for GPT-5.3 Instant. The full GPT-5.5 model scores 93.6% on GPQA Diamond, per OpenAI's own published benchmark table, independently corroborated by BenchLM's evaluation. A specific AIME 2025 score for the full (non-Instant) GPT-5.5 model has not been independently published. Against Claude Opus 4.8 (88.6% SWE-bench Verified), GPT-5.5 is statistically even on agentic coding.
The API context window is 1M tokens with a maximum output of 128,000 tokens per completion. The Codex product caps GPT-5.5 at 400,000 tokens. A context surcharge applies once prompts exceed roughly a quarter-million tokens, billing the full session at a higher input and output rate. No publicly available needle-in-haystack evaluation at full context has been published for GPT-5.5 as of June 2026.
GPT-5.5's input modalities span text, image, audio, and video within one model session, with text and tool-calls as the only outputs. Computer use is live via Codex's hosted shell and apply-patch tooling. Function calling, structured outputs, MCP server integration, OpenAI Skills, web search, and a hosted shell are all available in the API. Batch API and Flex pricing tiers are live. Canvas was removed in a May 2026 update; writing and code features now surface through native response blocks.
GPT-5.5 is available via the OpenAI API directly. AWS Bedrock added GPT-5.5 on June 1, 2026, following a $50 billion partnership that ended OpenAI's prior exclusivity with Microsoft Azure. Microsoft retains a non-exclusive license to OpenAI's IP through 2032 and Azure remains an active deployment target. Authentication is via API key for the direct endpoint and AWS IAM on Bedrock. Together AI and Fireworks availability had not been confirmed as of June 9, 2026.
The GPT-5.5 System Card was published April 24, 2026, at deploymentsafety.openai.com/gpt-5-5. OpenAI reports that GPT-5.5 Instant produces roughly half the hallucinated claims on high-stakes prompts (medicine, law, finance) compared to GPT-5.3 Instant. Safety training follows reinforcement learning with a long internal chain-of-thought before output, plus safety classifiers applied to filter harmful and personal content. OpenAI's Preparedness Framework (ASL-3 scope) was applied during evaluation. Specific Harmbench refusal rates have not been published for GPT-5.5.
Native audio, video, and image processing in a single model call makes GPT-5.5 well suited to multimodal agentic pipelines, replacing separate ASR and vision integrations, and to complex professional research and long-document analysis at full context. Voice-first applications needing sub-200ms latency should use OpenAI's Realtime API rather than GPT-5.5's standard endpoint, where TTFT runs 400-700ms. On agentic coding benchmarks, Claude Opus 4.8 posts comparable SWE-bench Verified scores and should be evaluated in parallel.
Training data includes publicly available internet text, licensed third-party datasets, and human trainer inputs. OpenAI's default policy excludes API inputs from training; enterprise zero-retention agreements are available separately. The exact training cutoff for GPT-5.5 has not been published; GPT-5.2's was August 31, 2025 as a reference point. SOC 2 Type II certification applies. API inputs sent via AWS Bedrock are governed by AWS data terms separately from OpenAI's direct API policy.
GPT-5.5 Instant replaced GPT-5.3 Instant as ChatGPT's default model on May 5, 2026. A May 2026 style update reduced overly long and bullet-heavy responses in the Instant variant. Canvas was deprecated in the same window, with writing and code blocks replacing it natively. GPT-5.5 Pro launched alongside the standard model, targeting the heaviest professional workloads. As of June 2026, community sources reference GPT-5.6 with a rumored larger context window.
Pricing
$5 per 1M input, $30 per 1M output. Cached input $0.50 per 1M (90% off). Prompts over 272K tokens billed at 2x input / 1.5x output for the full session. Batch API: $2.50 input / $15 output. GPT-5.5 Pro: $30 input / $180 output per 1M.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.150 | $0.030 | $0.180 |
| Support reply | $0.010 | $0.0090 | $0.019 |
| One coding agent run | $1.00 | $0.600 | $1.60 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Unified Multimodal Architecture: Native text, image, audio, and video processing in one model session. Removes the need for separate ASR and vision pipelines behind an orchestrating agent.
- 1M-Token Context Window: Largest context in the GPT-5 family. Handles entire codebases or book-length documents in a single API call (API only; Codex caps at 400K).
- Agentic Computer Use: Built-in computer use via Codex hosted shell and apply-patch, handling multi-step software tasks end-to-end without manual orchestration.
- Native MCP and Skills Integration: MCP server integration and OpenAI Skills allow tool-augmented agent loops without custom orchestration code.
- Prompt Caching Discount: Cached input tokens cost a small fraction of the standard input rate, delivering meaningful savings on repeat-context workloads like long system prompts.
Pros
- Ranks among the highest-scoring models on SWE-bench Verified, near the top of publicly benchmarked agentic coding models.
- First OpenAI model with native audio and video input, eliminating separate pipeline integrations.
- Prompt caching cuts costs substantially on repeat-context agent loops via a steep discount on cached input tokens.
Cons
- A context surcharge applies once a session's prompts cross roughly a quarter-million tokens, doubling input cost and raising output cost for the entire session.
- No native audio output: speech synthesis requires a separate model or the Realtime API.
- For pure text-only workloads, specialized low-cost models like Gemini 2.5 Flash charge a small fraction of GPT-5.5's per-token rate.
Benchmarks
- MMLU: 92.4% vendor-reported · 10 Sep 2026 — General-knowledge exam across 57 subjects, % correct.
- GPQA Diamond: 93.6% vendor-reported · 10 Sep 2026 — PhD-level science questions that are hard to search for, % correct.
- SWE-bench Verified: 88.7% vendor-reported · 10 Sep 2026 — Real GitHub issues fixed end to end, % solved.
- Output speed: 61 tok/s cited: Artificial Analysis · 10 Sep 2026 — Median tokens written per second as measured by Artificial Analysis.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
How much does GPT-5.5 cost in 2026?
GPT-5.5's standard API pricing is $5.00 per 1M input tokens and $30.00 per 1M output tokens, with cached input at $0.50 per 1M (a 90% discount). The Batch API tier costs $2.50 input / $15.00 output per 1M with results within 24 hours, and the GPT-5.5 Pro tier runs $30.00 input / $180.00 output per 1M. A context surcharge doubles input pricing and raises output pricing by 1.5x for the whole session once a prompt exceeds 272K tokens. A 100K-token coding task with 10K output costs about $0.80 at standard rates.
Does GPT-5.5 have a free plan?
GPT-5.5 has no free tier: it is API-only, and OpenAI has not published a free-usage allowance or trial period for it. Cost-conscious teams can lean on prompt caching to cut repeat-context costs, or route pure text-only tasks to a cheaper specialized model such as Gemini 2.5 Flash.
Which models compete with GPT-5.5 in 2026?
For coding-focused workloads, Claude 4.7 Opus offers a smaller context window but competitive coding depth. For teams already inside Google's ecosystem, Gemini 3.1 Pro offers a large context window with tighter Google integration. For pure text workloads where cost matters most, Gemini 2.5 Flash costs $0.15 per 1M input and $0.60 per 1M output, a fraction of GPT-5.5's rate.
What separates GPT-5.5 from Claude Opus 4.8?
GPT-5.5 scores 88.7% on SWE-bench Verified versus Claude Opus 4.8's 88.6%, putting the two models statistically tied on agentic coding. GPT-5.5 posts 92.4% on MMLU compared to Claude Opus 4.8's roughly 81.2% on the differently-scoped MMLU-Pro benchmark, and GPT-5.5 adds native audio and video input that Claude Opus 4.8 does not have. Claude Opus 4.8 leads on the harder SWE-bench Pro variant at 69.2%, a score GPT-5.5 has not published. In practice the choice comes down to multimodal needs and deployment platform more than the sub-1-point gap on standard SWE-bench.
How do you get started with GPT-5.5?
GPT-5.5 is accessed via the OpenAI API directly (api.openai.com) with an API key, through AWS Bedrock using AWS IAM authentication since June 1, 2026, or via Azure OpenAI Service. Official SDKs cover Python, TypeScript, JavaScript, Java, Go, and C#. There is no self-hosting or open-weight option: GPT-5.5 is fully proprietary under OpenAI's Terms of Use, so getting started means creating an API key and choosing a hosting endpoint rather than downloading model weights.
Top Alternatives
- Claude 4.7 Opus: Pick GPT-5.5 for Codex agent and multimodal breadth; pick 4.7 Opus for 200K context and coding depth.
- Gemini 3.1 Pro: Pick GPT-5.5 for Codex agent; pick 3.1 Pro for 1M context and Google integration.
HokAI guides covering GPT-5.5
- AIML API vs OpenRouter: Which AI Gateway Should You Use in 2026?: OpenRouter is being bought by Stripe for a reported .5B. AIML API isn't. A pricing, scale and ownership comparison for teams picking an AI gateway in 2026.