Last updated: 2026-09-30
OpenAI Agents API entered public beta on 10 September 2026 as a managed service for running long-lived cloud agents. It runs the Codex agent runtime, so OpenAI handles sessions, sandboxes, context compaction and subagents. Developers choose an OpenAI-hosted sandbox, their own infrastructure, or one of nine named sandbox partners.
About OpenAI Agents API
OpenAI Agents API is a developer API from OpenAI that runs long-lived cloud agents on your behalf. It was announced on 10 September 2026 and is in public beta. It exposes the same agent runtime that powers Codex as a managed service: OpenAI keeps the sessions, orchestration, context compaction and recovery, while your application supplies the tools and chooses where the agent runs.
You create a session with one call to client.beta.agents.sessions.create (or a POST to /v1/agents/sessions with the OpenAI-Beta: agents=v1 header). The call names a model such as GPT-6 Astra, a list of tools and an environment. Tools include MCP servers, custom functions, web search, programmatic tool calling, where the agent writes code to chain several tool calls and trim their output, and computer use, which drives a browser in an OpenAI-hosted environment. Turning on multi_agent lets the main agent hand pieces of a task to subagents; the launch examples set max_concurrent_subagents to 3 and 4.
The environment is the main design choice. The openai_hosted type gives you an OpenAI-managed sandbox where the agent runs code, edits files and produces artifacts. The self_hosted type points the agent at your own machines and private network. OpenAI also names nine sandbox partners: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop and Vercel. As a session nears its context limit, the API compacts earlier context automatically, so one task can span several context windows.
It suits teams that want a Codex-style coding or research agent inside their own product without building a session store, a sandbox pool and compaction logic. OpenAI's cookbook ships five sample apps: an incident response agent, a Slack bot, a data analyst that runs read-only SQL, a GitHub issue investigator and a document reviewer. Against open-source frameworks such as LangChain or Mastra, you trade portability for a managed runtime, since the documentation shows OpenAI models only. For visual workflow building, n8n is the closer fit.
Because the product is in public beta, endpoints, parameters and the beta header can change before general availability. Related OpenAI products on HokAI include Codex Security, dots and the GPT-6 Sol model.
Pricing
There is no separate Agents API fee: you pay the selected model's per-token rate plus the standard rates for any OpenAI tools and hosted containers. 50 output. 00 output).
00 per 1,000 calls plus search content tokens at model rates. 92 (64 GB) per 20-minute session. Read from the OpenAI pricing page on 2026-09-30.
| Tier | Monthly price | What it includes |
|---|---|---|
| GPT-6 Astra | Custom | per-token |
| GPT-6.1 Sol | Custom | per-token |
| GPT-6 Luna | Custom | per-token |
Key Features
- Sessions from one call: A single sessions.create call sets the model, tools, environment and first task, and the launch example needs 29 lines of JavaScript.
- Three environment routes: Run agents in an OpenAI-hosted sandbox, on your own infrastructure with the self_hosted type, or on one of 9 named sandbox partners.
- Automatic context compaction: Earlier context is summarized as a session nears its limit, so one task can continue across several context windows without your own compaction code.
- Tool search and programmatic tool calling: Tool definitions load on demand to cut token use, and agents can run tool calls in parallel and filter results in code before they reach the model.
- Subagents: The multi_agent setting lets a main agent delegate to subagents that each keep their own context; the launch examples cap concurrency at 3 and 4.
- Computer use: The computer_use tool runs a browser in an OpenAI-hosted environment, with optional screenshots and an approval step for each website the agent wants to reach.
- Open-source runtime: The Codex agent runtime behind the API has a public codebase, so developers can read how model calls, tools and context are coordinated.
Pros
- OpenAI runs four things you would otherwise build: sessions, orchestration, context compaction and crash recovery.
- Billing uses rates that are already published for models, tools and containers, so a run's cost can be estimated from the pricing page before you ship.
- OpenAI's launch post quotes the CTO of Ciridae, a customer, reporting an eval score rise from 0.71 to 0.85 and a 4x latency cut after adopting subagents (vendor-published testimonial, not independently tested).
Cons
- The API is in public beta and OpenAI says it will iterate before general availability, so endpoints and the OpenAI-Beta: agents=v1 header can change.
- OpenAI's data controls page lists /v1/agents as not eligible for Zero Data Retention, with application state kept until deleted and 30-day abuse monitoring logs.
- The documentation shows OpenAI models only, so teams that want to switch model vendors per task will find open-source frameworks more portable.
Data Handling
- Training-data policy
- OpenAI states that data sent to the API is not used to train or improve its models unless you explicitly opt in.
- Data retention
- 30 days
Frequently Asked Questions
What are OpenAI Agents API's pricing plans in 2026?
There are no plans and no separate Agents API fee. You pay the chosen model's per-token rate, with GPT-6 Astra at $10.00 input and $50.00 output per 1M tokens on short context, plus standard rates for tools: $10.00 per 1,000 web search calls and $0.03 to $1.92 per 20-minute hosted container session. Long-context prompts and larger sandboxes raise the bill.
Is OpenAI Agents API free to use?
The documentation describes no free plan or free allowance for the Agents API. Public beta access is open to all developers and adds no fee of its own, but every run is billed at model, tool and container rates. The cheapest way to test is a small model and a short task.
What are the best alternatives to OpenAI Agents API?
LangChain and Mastra suit teams that want an open-source framework they host themselves and that can call several model vendors. E2B fits when you only need sandboxes for code your own agent writes. n8n is the pick for visual workflows with little code.
OpenAI Agents API or LangChain: which should you pick?
Pick the Agents API when you want OpenAI to run sessions, the sandbox and context compaction, and you are happy to use OpenAI models. Pick LangChain when portability across model vendors matters more than a managed runtime, or when your data cannot go through an endpoint that is not Zero Data Retention eligible.
What does it take to start using OpenAI Agents API?
You need an OpenAI API key and the OpenAI SDK for your language. The quickstart runs a directory-tree script in an OpenAI-hosted sandbox, and the cURL examples require Bash and jq. A first session is one sessions.create call.
Top Alternatives
- LangChain: Pick LangChain if you want to host the loop yourself and swap model vendors; pick the Agents API if you want OpenAI to run sessions and sandboxes.
- Mastra: Choose Mastra when your team writes TypeScript workflows and wants self-hosting; choose the Agents API for a managed Codex-style runtime.
- E2B: Choose E2B when you only need sandboxes for your own agent code; the Agents API already lists E2B as one of its sandbox partners.
- n8n: Pick n8n if you build visual workflows without much code; pick the Agents API if you write the agent in code.