OpenAI Agents API Pricing: What You Pay, What It Replaces, and When Claude or Bedrock Fits Better
The OpenAI Agents API is a managed service, in public beta, that runs long-lived cloud agents on OpenAI's Codex runtime. OpenAI handles sessions, orchestration, context compaction and recovery. Billing uses published model, tool and container rates with no separate platform fee, and the /v1/agents endpoint is not eligible for Zero Data Retention.
The short version
The OpenAI Agents API has no fee of its own. You pay model tokens, tool rates and sandbox container time, and tokens dominate: one long run costs about $25 on GPT-6 Astra and $0.25 on Luna. Claude Managed Agents adds $0.08 per running hour. None of the three managed runtimes offers zero data retention.
OpenAI's Agents API has no price of its own. You pay the model's per-token rate, the standard rate for any tool the agent calls, and container rates for the sandbox it runs in. That sentence comes straight from OpenAI's overview page, and it is the whole pricing story and also the reason a budget is hard to write.
The OpenAI Agents API is in public beta. It hands a long-running agent to OpenAI as a managed service: sessions, orchestration, context compaction and crash recovery all run on OpenAI's side. Your code supplies the tools and picks where the agent executes.
The short answer for a team lead deciding this month: use it if your agents already run on OpenAI models and you are tired of maintaining a loop, a sandbox and a retry layer. Look at Anthropic's Claude Managed Agents if you want a runtime billed by the hour. Look at Amazon's Bedrock version if procurement runs through AWS. None of the three is ready for strict zero-retention work.
We did not run any of these products for this piece. Every figure comes from the vendors' own documentation, opened on 1 October 2026, and every comparison is arithmetic on published list prices.
What changed: a managed agent runtime from OpenAI
Before this API, OpenAI sold you a model and a set of tools, and you wrote the loop that tied them together. The Agents API moves that loop to OpenAI. It exposes the same runtime that powers Codex, the company's coding agent, as something your own application can call.
You create a session with one call to client.beta.agents.sessions.create, or a POST to /v1/agents/sessions with the OpenAI-Beta: agents=v1 header. The call names a model, a list of tools and an environment. OpenAI's quickstart treats the environment as the main design choice.
- OpenAI-hosted: OpenAI provisions and manages the sandbox where the agent runs commands and code.
- Self-hosted: the agent works on your own machines and network.
- Sandbox partners: OpenAI also names nine partners, among them E2B and Modal.
The pricing section of the Agents API overview, captured 1 Oct 2026: no separate platform fee.
Tools include MCP servers, custom functions, web search and computer use, which drives a browser in an OpenAI-hosted environment. Turning on multi_agent lets the main agent hand pieces of work to subagents. The overview example sets max_concurrent_subagents to 4, so a single session can run five agents at once, counting the main one.
Three rivals arrived in the same window, which is why the question is live. Anthropic ships Claude Managed Agents, also in beta. Amazon previews Bedrock Managed Agents, powered by OpenAI. Each sells the same promise: you stop writing the agent loop.
Who this affects
Three groups should read the details, and they have different problems.
Teams already calling GPT-6 Astra or a smaller sibling through the Responses API are the obvious audience. They can try a managed session without changing vendors. Teams on open-source frameworks have the harder call, because a managed runtime trades portability for less code.
Platform and procurement leads are the second group. A beta header, a 30-day log and a state store with no expiry date all land on their desk. The third group is anyone on AWS who wants OpenAI models without a separate vendor contract, which is the case Bedrock was built for.
Solo builders and small teams should treat this as a decision about maintenance, not capability. If you only run one short task at a time, a plain model call is cheaper and simpler than any managed runtime.
What it costs: tokens, tools and sandbox time
There are three meters, and they run at once. OpenAI's pricing page lists each one, read on 1 October 2026.
Model tokens come first. Short-context standard rates per million tokens, taken from the pricing table:
| Model | Input | Cached input | Output |
|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $50.00 |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.50 |
Long prompts cost more. For Astra, the long-context rate is $20.00 input and $75.00 output. Because an agent session re-reads its history on every step, long context is where a quiet bill grows.
Tools come second. Web search is $10.00 per 1,000 calls, plus the search content billed as tokens at the model's rate. The tools table on the pricing page has no separate line for computer use. We read that as computer use being billed through model tokens, but OpenAI does not say so in as many words, so confirm it before you commit a budget.
Sandbox time comes third. Hosted containers are billed per 20-minute session, and the price climbs with memory.
| Container size | Price per 20-minute session |
|---|---|
| 1 GB | $0.03 |
| 4 GB | $0.12 |
| 16 GB | $0.48 |
| 64 GB | $1.92 |
A one-hour run in the smallest container is three of those blocks, about nine cents. Even a 64 GB box for a full hour is $5.76. The container is rarely the expensive part.
A worked example: one long run at three model tiers
Tokens decide the bill, so here is the arithmetic. Suppose an agent works through ten steps and reads a 200,000-token context at each one, with nothing cached. That is two million input tokens. Suppose it writes 100,000 tokens of output across the run.
| Model | Input cost | Output cost | Run total |
|---|---|---|---|
| GPT-6 Astra | $20.00 | $5.00 | $25.00 |
| GPT-6.1 Sol | $4.00 | $1.00 | $5.00 |
| GPT-6 Luna | $0.20 | $0.05 | $0.25 |
This is list-price arithmetic, not a measurement. Real runs cache part of the history, and caching cuts the input line sharply: at Astra's cached rate, the same two million tokens would cost $2.00 instead of $20.00.
The gap between tiers is a hundredfold. Picking the model matters more than any other setting you control, and the sandbox choice barely registers beside it. A team that sends every step to the top model pays for capability most steps do not need.
For a sense of how the tiers compare on quality, the model leaderboard shows published scores side by side. Cost per step only means something next to what the step has to get right.
Where the Agents API wins
The API earns its place on four counts, and they are about work you stop doing, not about new capability.
First, you stop owning the loop. OpenAI runs sessions, orchestration, compaction and recovery. Automatic compaction means one task can continue across several context windows without summarising code you wrote yourself.
Second, the cost model is legible. Every one of the three meters is on a published page, so a run can be estimated before you ship. The cost of a failure is also bounded: a crashed session recovers on OpenAI's side, not in your retry queue.
Third, it inherits the Codex runtime. That runtime has a public codebase, which means you can read how model calls, tools and context are coordinated. For a team that has to explain its agent to a security reviewer, that is a real asset.
Fourth, the same agents reach into the rest of OpenAI's product line. The always-on agent covered in our guide to what OpenAI's dots actually do and the security agent Codex Security both sit on the same vendor, so a single contract covers them.
Where Claude Managed Agents or Bedrock fits better
Anthropic's runtime differs on three points that matter to a buyer.
Billing. Claude Managed Agents bills tokens at standard Claude rates plus $0.08 per session-hour. The clock runs only while the session is running. Time spent idle, waiting for your next message or a tool confirmation, does not count. For an agent that waits on humans, that is a cleaner meter than a container block.
Model choice. Every request carries the managed-agents-2026-04-01 beta header, and it runs Claude models only, which suits a team that already prefers them. Anthropic also lets sessions run in a cloud sandbox or on your own infrastructure, and offers scheduled deployments for recurring runs.
Where it runs. It is also available on Claude Platform on AWS, so AWS-committed spend can reach it. Our separate record for Claude in Amazon Bedrock covers the Claude models side of the same relationship.
Amazon's offer is different again. Bedrock Managed Agents combines OpenAI models with the Codex agent runtime and Amazon Bedrock AgentCore, and each agent gets its own AWS identity and permissions. AWS says the runtime and model inference stay inside AWS, according to its product page.
The catch is that it is a public preview. AWS's documentation lists three US endpoints: N. Virginia, Ohio and Oregon. The Region codes are us-east-1, us-east-2 and us-west-2, and the preview signs requests with the bedrock-mantle service name. AWS says model availability differs by Region and account, and that a deployed stack can keep billing when no turn is running, because the example creates a NAT gateway and storage. Read AWS's preview limitations page before you build.
| Question | OpenAI Agents API | Claude Managed Agents | Bedrock (OpenAI) |
|---|---|---|---|
| Status | Public beta | Beta | Public preview |
| Runtime charge | Container blocks | $0.08 per session-hour | AWS resource costs |
| Models | OpenAI only | Claude only | OpenAI on Bedrock |
| Zero retention | No | No | Not stated |
The last row deserves its own section, because it is where the three products turn out to be alike.
Data handling: none of the three offers zero retention
OpenAI's data controls page has a row for /v1/agents. Customer data is not used for training. Abuse monitoring logs are kept for 30 days. Application state is kept until you delete it. The endpoint is not eligible for Zero Data Retention.
Anthropic's documentation is just as direct. Managed Agents is stateful by design, with conversation history, sandbox state and outputs stored server-side. So it is not currently eligible for Zero Data Retention or HIPAA Business Associate Agreement coverage.
For Bedrock, the AWS pages we read state that the agent runtime and inference remain inside AWS, and they list no compliance attestations for the preview. That is an absence on the pages reviewed, not a statement that none exist. Procurement should ask AWS directly.
What this means in practice: if your contract requires zero data retention, none of these managed runtimes fits today. A stateless pattern, where you keep state yourself and call a model API, is the path. The managed convenience is paid for in stored state.
Switching cost and lock-in
A managed runtime moves code from your repository into someone else's service. That is the point, and also the exit problem.
OpenAI's documentation shows OpenAI models only. Anthropic's runtime runs Claude only. If you want to switch model vendors per task, or hedge against a price change, an open framework keeps that option. Mastra and LlamaIndex are two code-first options in the same directory, alongside LangChain, which we describe but do not link here because it is already well covered.
Durable execution is the other route. Trigger.dev and Inngest run long jobs with retries and state under your control, and either can call any model. They give you the recovery behaviour without a vendor's agent loop.
The honest cost of leaving is rewriting the session, tool and event handling around a different shape. Both OpenAI and Anthropic use their own event model, so a move between them is a port, not a config change. Plan for that before you build ten agents on one beta.
You can browse the full AI orchestration category to see where the managed offers sit against the frameworks.
The case against a managed runtime
The best argument against all of this is blunt: wait. Three betas are shipping with changing endpoints, a state store you cannot empty on a schedule, and a retention row that rules out regulated work. A team that waits six months loses nothing and sees which design survives.
There is real force in that. OpenAI says plainly that the API will change before general availability, and the beta header is the proof. Building production flows on it now means accepting rework.
Our answer is that the risk depends on what you hand it. A research agent that reads public pages and writes a draft is cheap to rebuild. A payments flow is not. Start with work where a rewrite costs a week, keep inputs free of regulated data, and keep your own copy of anything the agent produces.
The other half of the answer is that the maintenance burden is real today. A team running its own loop, sandbox and retry layer pays every week, not in some future quarter.
What to do this week
Pick one agent task that already exists as a script, and run it three ways: once on a small model, once on the top model, and once on whichever managed rival you are curious about. Record tokens, minutes and a pass or fail against a fixed check.
- Cap the spend: set a hard token ceiling per session before you start.
- Cache first: structure prompts so the stable history is cached.
- Route by step: send planning to a larger model and routine steps to a smaller one.
- Log everything: keep your own record of tool calls, because the vendor's event log is not your audit trail.
If the choice between runtimes still depends on your stack, Smart Match can narrow it from what you describe. For wider context on agent categories, see our guides to the best agentic AI, AI agent builders and agents for business operations. The consumer side of the same market is covered in Manus against ChatGPT's agent and in the always-on agent comparison.
Who should pick which
| Situation | Pick | Reason |
|---|---|---|
| Already on OpenAI models | Agents API | Same vendor, one contract |
| Hour-billed runtime, Claude models | Claude Managed Agents | $0.08 per running hour |
| Spend and identity live in AWS | Bedrock Managed Agents | Own IAM identity, AWS billing |
| Need model portability | Open framework | No vendor runtime |
| Zero data retention required | None of the three | Stored state is not optional |
What would change this verdict in 2026
Our advice flips if any of these happen in the next two quarters, and the first three are in the vendors' hands. General availability for any of the three would end the beta-churn objection. A Zero Data Retention option on the /v1/agents endpoint would reopen regulated work. A published price for computer use would remove the one gap in OpenAI's own meter.
The fourth is outside the vendors' control: independent testing. Every performance claim at launch, including a customer's reported evaluation gain from subagents, is vendor-published. Until someone outside reproduces them, treat speed and accuracy numbers as marketing.
What comes next
The practical question is no longer whether to build an agent loop but which vendor holds it. Watch the three status labels. The day any of them drops the word beta, the cost-and-retention comparison above becomes a real procurement decision instead of a pilot.
For OpenAI and Amazon Web Services, the next test is whether Bedrock's preview adds the regions and attestations a European or Asia-Pacific buyer needs. For Anthropic, it is whether a stateful runtime can ever sit inside a zero-retention contract. Watch those two, and rerun the three-way test above when either changes.
Frequently asked questions
Does the OpenAI Agents API have its own price?
No. OpenAI bills the selected model's per-token rate, the standard rate for any tool the agent uses, and container rates for OpenAI-hosted sandboxes. Containers run from $0.03 for 1 GB to $1.92 for 64 GB per 20-minute session. Tokens are usually the largest line.
How much does a long agent run cost?
It depends mostly on the model. As list-price arithmetic, two million input tokens and 100,000 output tokens cost about $25 on GPT-6 Astra, $5 on GPT-6.1 Sol and $0.25 on GPT-6 Luna, before any caching. Cached input is priced at a fraction of the normal rate, so real runs usually cost less.
How does Claude Managed Agents pricing compare?
Anthropic bills tokens at standard Claude rates plus $0.08 per session-hour. The runtime clock runs only while the session is in the running state, not while it is idle. OpenAI bills sandbox time in 20-minute container blocks instead, so the two meters are not directly comparable.
Can I use these managed agents with zero data retention?
Not today. OpenAI's data controls page lists /v1/agents as not eligible for Zero Data Retention, with 30-day abuse logs and application state kept until deleted. Anthropic says Managed Agents is not currently eligible for Zero Data Retention or HIPAA coverage. AWS lists no attestations for its preview.
Is Amazon Bedrock Managed Agents a way to get OpenAI agents inside AWS?
Yes, as a public preview. It combines OpenAI models with the Codex harness and Amazon Bedrock AgentCore, and each agent runs under its own AWS identity. AWS lists three US Regions, so teams needing EU or Asia-Pacific data residency cannot use it yet.
Covered in this guide
- OpenAI Agents API: Managed API for cloud agents on the Codex runtime: hosted sandboxes, 9 sandbox partners, subagents and computer use, in public beta since Sept 2026.
- Bedrock Managed Agents: AWS preview service that runs stateful OpenAI agents through bedrock-mantle endpoints in 3 US Regions, with per-agent IAM roles and your own compute.
- Amazon Web Services: Amazon Web Services (AWS), launched 2006, is the world's largest cloud platform: $128.7B FY2025 revenue, roughly 29% global market share, and 200+ services.
- Anthropic: Anthropic, founded 2021 by 7 ex-OpenAI researchers, builds Claude and was valued near $965B after its May 2026 Series H round.
- Claude in Amazon Bedrock: Anthropic's Claude models served by AWS in 27 regions, billed per token on your AWS invoice and secured with IAM, KMS keys and CloudTrail logs.
- Codex Security: OpenAI's application security agent scans connected GitHub repositories, tests likely bugs in a sandbox, and drafts fixes for review as pull requests.
- E2B: Enterprise sandbox infrastructure for AI agents to execute code securely. Firecracker microVMs, sub-second cold starts, Python and JavaScript SDKs. Free tier with $100 credit; paid plans scale up to enterprise BYOC deployment.
- GPT-6 Astra: OpenAI's flagship model, launched September 2026 as the first ever rated at the Preparedness Framework's Critical cybersecurity level.
- GPT-6 Luna: GPT-6 Luna (Sept 2026) is OpenAI's lowest-cost GPT-6 model, tuned for high-volume chat, extraction and classification with six adjustable reasoning-effort levels.
- GPT-6.1 Sol: GPT-6.1 Sol is OpenAI's September 2026 mid-tier reasoning model for coding, computer use and document work, built to sit just below GPT-6 Astra.
- Inngest: Inngest runs event-driven background jobs and AI workflows with no queue or worker fleet to manage. A Free plan that needs no credit card covers development and small production apps.
- LlamaIndex: The world's most accurate agentic OCR and document-specific AI workflows for enterprise automation
- Mastra: Open-source TypeScript framework for building AI agents and workflows, routing to 90+ model providers through one interface.
- Modal: AI infrastructure that developers love: serverless compute for ML inference, training, and batch processing
- OpenAI: OpenAI builds the GPT-5.6 model family (Sol, Terra, Luna), o3, ChatGPT (900M+ weekly users), and the OpenAI API. Closed a $122B round at an $852B valuation in March 2026, the largest private funding round in history.
- Codex: Clones a GitHub repo into a sandbox, writes and tests code, then opens a pull request; OpenAI reported 20 million active users in August 2026.
- Trigger.dev: Open-source TypeScript platform for durable background jobs and AI agents. Checkpoint-resume execution survives crashes and redeploys, auto-retries, real-time traces. Free tier available.
Sources
Still deciding?
This guide covers a handful of options. Smart Match checks every listing in the directory against how you actually work and what you can spend, then hands you the shortlist and the reason behind each pick.
Start Smart MatchRelated guides
- AI Development Services in 2026: Which Layer You Actually NeedBuyer's guideHow to pick, across a category
- The AI Tool Ecosystem in 2026: Buy the Meter, Not the CategoryBuyer's guideHow to pick, across a category
- Best AI Chatbots in 2026: Pick by the Job, Not the LeaderboardUpdatedRechecked against current sources
- Best AI Coding Assistants in 2026: Pick the Job, Not the BrandBuyer's guideHow to pick, across a category
- Best AI Companies in 2026: Who Is Actually LeadingBuyer's guideHow to pick, across a category
- Best AI for Writing a Business Plan in 2026: Tested Picks and a VerdictBuyer's guideHow to pick, across a category