GPT-5.4 fits agentic coding teams and anyone automating desktop workflows who has outgrown a plain chat model. It unifies chat and coding into a single model scoring around 80% on SWE-bench Verified, so teams no longer need a separate coding-only specialist. Budget-conscious or audio-first teams should look elsewhere; a mini variant or a competing model fits better.
GPT-5.4 is OpenAI's flagship large language model, released March 2026, built for agentic coding and desktop automation. Its Computer Use API scores 75% on OSWorld-Verified, ahead of the 72.4% average human baseline, making it the first OpenAI model to fold coding-specialist and general-chat capability into one unified architecture.
Where it sits
- $5.63/M$ per 1M tokensBlended price (3:1)Lower is better#46 / 64peer median $1.70/Mvendor price, checked by HokAI
- --tokens/sOutput speedHigher is better-- / 39peer median 90 tok/scited: Artificial Analysis
- 80%% solvedSWE-bench VerifiedHigher is better#12 / 28peer median 78.3%per source, see benchmark scores
- 74.8%% correctGPQA DiamondHigher is better#32 / 44peer median 88.3%per source, see benchmark scores
Pricier than 73% of the 64 GA models with a published price, mid-pack on SWE-bench Verified (rank 12 of 28), and one of 65 whose vendor states it does not train on customer data.
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: OpenAI · Family: GPT-5
Context window: 1,000,000 tokens · Max output: 128,000
Input modalities: text, image, tool-calls, code · Output: text, tool-calls, code
About GPT-5.4
GPT-5.4 is OpenAI's flagship model for professional work, announced March 5, 2026 alongside GPT-5.4 Thinking (a reasoning-focused variant) and GPT-5.4 Pro (a higher-cost, deep-reasoning tier). GPT-5.4 mini and nano followed on March 17, 2026 for high-volume workloads. It is the first mainline GPT-5 model to fold in GPT-5.3-Codex's coding capabilities directly, replacing the earlier split between a general chat model and a dedicated coding specialist, and it adds five reasoning-effort levels (none, low, medium, high, xhigh). GPT-5.5 superseded it as OpenAI's flagship six weeks later, on April 23, 2026, leaving GPT-5.4 in a mid-lineup position above mini/nano and below GPT-5.4 Pro and GPT-5.5.
The headline context window is 1M tokens, the largest OpenAI has shipped in its API. In practice the pricing and rate-limit ceiling sits well below that: sessions under a fixed input threshold use standard rates, while longer sessions pay a premium across the whole session, not just the overflow. Retrieval quality also degrades in the back half of the window, with details buried in the middle of very long prompts getting missed most often.
Modalities are text and image input with text and tool-call output; there is no native audio or video despite the multimodal framing of its document-understanding work, which covers dense scans, handwritten forms, and chart-heavy reports in a single pass. The model ships a Computer Use API for desktop automation, plus deep research and a tool-search feature for selecting from large tool libraries.
GPT-5.4 is available through the OpenAI API and ChatGPT (Plus and Pro tiers, with nothing in between), and reached general availability on AWS Bedrock on June 1, 2026 in US East, US West, and AWS GovCloud, ending Azure's prior exclusivity. The GPT-5.4 Thinking system card, published on OpenAI's Deployment Safety Hub and last updated April 24, 2026, documents red-teaming for the reasoning variant; OpenAI also reports a 33% reduction in factual errors relative to its immediate predecessor, and has not disclosed a parameter count or training cutoff for GPT-5.4.
Pricing
Under 272K input tokens: $2.50 per 1M input, $15 per 1M output, $1.25 per 1M for cached input (a 50% discount). Beyond that threshold, the whole session reprices at $5 per 1M input and $22.50 per 1M output. A separate GPT-5.4 Pro tier costs $30 per 1M input and $180 per 1M output for higher-stakes enterprise reasoning.
What a real job costs
| Job | Input | Output | Total |
|---|---|---|---|
| Summarise a 20-page PDF | $0.075 | $0.015 | $0.090 |
| Support reply | $0.0050 | $0.0045 | $0.0095 |
| One coding agent run | $0.500 | $0.300 | $0.800 |
Budgets: 20-page PDF = 30k in / 1k out · Support reply = 2k in / 300 out · Coding agent run = 200k in / 20k out. Computed from the vendor's per-token prices at render time; cached-input discounts are not applied.
Key Features
- Native Computer Use API: Takes screenshots and issues cursor, click, and keyboard actions against a desktop environment, scoring 75% on OSWorld-Verified versus a 72.4% human baseline.
- 1M-Token Context Window: Combines 922K input tokens with up to 128K output, the largest context OpenAI has shipped in the API; Codex users can opt into the full window via context-compaction settings.
- Five Reasoning-Effort Levels: Configurable from none to xhigh, letting developers trade cost and latency against accuracy on a per-request basis.
- Unified Coding Capability: Folds GPT-5.3-Codex's coding specialty into the mainline model, hitting roughly 80% on SWE-bench Verified without a separate coding-only model to switch to.
Pros
- Folds a dedicated coding specialist into the mainline model, so agent stacks no longer need to juggle a separate model for coding-heavy steps.
- Graduate-level reasoning (GPQA Diamond) paired with a Computer Use API that clears the average human baseline on desktop automation makes it a genuine generalist, not just a chat model with a benchmark bump.
- Prompt caching roughly halves the cost of resending the same system prompt or document context across a long agentic session.
Cons
- No native audio or video input or output, despite the multimodal framing of its document-understanding work.
- Crossing the long-context pricing threshold roughly doubles the input rate and raises output pricing for the whole session, not just the extra tokens, a cost cliff worth designing around.
- No free API tier; ChatGPT access starts at $20/month (Plus) and jumps straight to $200/month (Pro), with no plan in between.
Benchmarks
- MATH: 97.2% vendor-reported · 06 Mar 2026 — Competition maths problems, % solved.
- ARC-AGI 2: 62.1% vendor-reported · 06 Mar 2026 — Abstract visual puzzles built to resist memorisation, % solved.
- HumanEval: 95.1% vendor-reported · 06 Mar 2026 — Small programs that must pass hidden tests, % passing.
- GPQA Diamond: 74.8% vendor-reported · 06 Mar 2026 — PhD-level science questions that are hard to search for, % correct.
- OSWorld Verified: 75% vendor-reported · 06 Mar 2026 — Tasks completed by operating a real desktop, % solved.
- SWE-bench Verified: 80% vendor-reported · 06 Mar 2026 — Real GitHub issues fixed end to end, % solved.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
How much does GPT-5.4 cost in 2026?
Standard pricing is $2.50 per 1M input tokens and $15 per 1M output tokens, as long as a session stays under 272K input tokens; cached input costs $1.25 per 1M, half price. Cross that limit and the entire session bills at a higher $22.50 per 1M output rate instead of the standard one. There's also a separate GPT-5.4 Pro tier, priced well above standard rates, for higher-stakes enterprise reasoning.
Does GPT-5.4 have a free plan?
GPT-5.4 has no free tier through the OpenAI API. ChatGPT access starts at $20/month for Plus and jumps to $200/month for Pro, with nothing in between; free ChatGPT users have also been reported to lose access to GPT-5.4 in Codex CLI even when the site advertises a free trial.
Which tools compete with GPT-5.4 in 2026?
GPT-5.4's closest alternatives are GPT-5.5 (OpenAI's newer flagship, stronger on coding benchmarks and native audio/video input), Claude Opus 4.8 (better for long-horizon coding sessions before its price tier increases), and Gemini 3.1 Pro (a mixture-of-experts model with a graduate-level reasoning edge). Teams already invested in OpenAI's Computer Use API or ChatGPT ecosystem have the least reason to switch.
What separates GPT-5.4 from GPT-5.5?
GPT-5.5 is OpenAI's newer flagship, released six weeks after GPT-5.4, and scores higher on SWE-bench Verified at 88.7% while adding native audio and video input that GPT-5.4 lacks entirely. GPT-5.4 stays the cheaper, text-and-image-only option and keeps its lead on desktop computer-use automation through the Computer Use API.
How do you get started with GPT-5.4?
Getting started needs an OpenAI API key for direct access, or a ChatGPT Plus or Pro subscription for the hosted product; there is no self-hosting option. Developers running agentic workloads typically start at a lower reasoning-effort level before stepping up to high or xhigh for complex multi-step tasks, and enterprise users can reach GPT-5.4 through AWS Bedrock or Azure instead of the direct API.
Top Alternatives
- GPT-5.5: Pick GPT-5.5 for higher SWE-bench Verified accuracy and native audio/video input; pick GPT-5.4 for roughly half the per-token output price if text-and-image input is enough.
- Claude Opus 4.8: Pick Claude Opus 4.8 for stronger long-horizon coding and a larger context ceiling before pricing steps up; pick GPT-5.4 for its native Computer Use API and cheaper entry-level output pricing.
- Gemini 3.1 Pro: Pick Gemini 3.1 Pro for its mixture-of-experts efficiency and a graduate-level QA benchmark lead; pick GPT-5.4 for desktop and browser automation through the Computer Use API.
HokAI guides covering GPT-5.4
- Our llms.txt Has 708 Links. Google Ignores All of Them.: We fetched eight AI directories' llms.txt on 23 August 2026. Five had none. Ours lists 708 links that match our API exactly, and Google still ignores it.