Automation and Agents
OpenAI Agents API vs. Responses API: Which Should You Use?
OpenAI now offers a managed Codex harness through Agents API. Here is how to decide whether your workflow needs it, an application-owned SDK loop, or the direct Responses API.

On this page
- What OpenAI announced
- Three interfaces represent three control choices
- When a managed harness earns a place
- When the SDK or direct Responses API is cleaner
- The managed API does not take your application off duty
- Compare the options on one fair task
- Count cost and retention as architecture inputs
- A practical choice table
- Migrate in small, reversible steps
- Common questions
- Is Agents API the same as the Agents SDK?
- Does the Agents API require a sandbox?
- Does self-hosting make Agents API eligible for ZDR?
- Does the managed harness make an agent production-ready?
- The decision to make
If your agent needs durable sessions, orchestration, context management and an optional place to run code, the Agents API is worth evaluating. If your task is a direct model call, or you want your own application to own the agent loop, start with the Responses API or Agents SDK. OpenAI's September 10, 2026 announcement made its Codex harness available through a public-beta API. It did not make every task an agent task, or remove the application work around permissions, evaluation, interface and safe stopping.
The practical question is not “Which API is newest?” It is “Which responsibilities do I want this application to own?” The answer depends on the task you actually want done, its failure cost, and the parts your application must control. This guide compares those boundaries and gives you a bounded test to run against real work. If you already know the managed harness fits, use the companion guide to choose an Agents API execution environment. The current Agents API documentation also says data residency is limited to the United States and the API is not eligible for Zero Data Retention, even with a self-hosted sandbox. Treat that as an early eligibility check, not a footnote. Before choosing any automation, use the Work Worth Doing test to confirm the task should be automated at all.
View image detailWhat OpenAI announced
On September 10, OpenAI introduced the Agents API in public beta as a way to use the Codex harness from an application. OpenAI describes that harness as the part that manages sessions, orchestration, context compaction and recovery. The application still supplies tools and chooses where the agent executes. The service can connect the model to a sandbox that runs code, edits files, uses MCP servers and produces artifacts. Those are product descriptions from OpenAI, not results from a comparison we ran. OpenAI's announcement and current Agents API overview are the sources for those claims.
It helps to divide a deployed agent into four pieces. The model interprets instructions and produces decisions. The harness repeats the work cycle, preserves session context and coordinates tools. The environment provides files, shell commands or other execution space. The application connects that machinery to people, data, business rules and permissions. An API can change who operates one part without taking ownership of the others.
The Agents API changes the harness boundary. It gives an app a managed session and a set of ways to steer work, receive events and continue a task. It does not know whether a proposed change is authorized for your business, whether a customer record is the right one, or whether a final artifact satisfies your team's bar. Those judgments need explicit application behavior and human review.
The launch page says there is no separate Agents API fee, while current documentation describes charges for model usage, tools and OpenAI-hosted sandboxes at their applicable rates. That is different from saying an agent run is free. A workflow can make repeated model and tool calls, allocate a container, run retries, and take staff time to inspect. Without measuring those inputs, a rate-card comparison is not a cost estimate.
Three interfaces represent three control choices
OpenAI's Agents guide distinguishes three starting points. Its table is a vendor-authored map, useful for identifying the design question but not an independent effort or quality benchmark.
- Agents API: Who owns the loop?: OpenAI runs a managed Codex harness; State and tools: Managed sessions, service-connected and app-handled tools, optional sandbox; Strongest reason to consider it: You need a longer-running agent and want the harness operations handled by the service
- Agents SDK: Who owns the loop?: Your application runs the agent loop; State and tools: Reusable agent definitions, tools and handoffs in your app; your storage and runtime choices; Strongest reason to consider it: You want an agent framework but need to keep execution and control flow in your product
- Responses API: Who owns the loop?: Your application builds the loop, or makes direct calls; State and tools: Direct model responses, tools, response chaining or conversation state; Strongest reason to consider it: You need the model/API building block and want the smallest useful mechanism
View image detailThe distinctions matter because “API” can sound like one tool competing with another on intelligence. In this case, the important difference is how much runtime machinery is managed for you. A direct Responses call may be all a summarization feature needs. A task that can be steered across steps might benefit from an SDK. A service that must preserve a session and coordinate longer work may be a candidate for the managed API.
These are not exclusive forever. A product could use direct model calls for simple classification and a managed session for a longer task. It could also use an SDK for application-controlled work and a separate service for a narrow background process. The boundary should follow the job and its controls, not a desire to standardize prematurely on one interface.
There is a hidden cost on each side. With a direct API, you may need to build the loop, state transitions, retries and observability that fit your application. With a managed harness, you accept a service boundary and its documented limits, and still have to build your application's tool handlers, policy, user experience, test suite and auditing. The vendor handles some plumbing, not the product decision.
When a managed harness earns a place
The strongest reason to test the Agents API is when your existing task requires persistent work across turns and the orchestration around it is becoming a real part of your product. Imagine a hypothetical internal release assistant. It needs to inspect several public release notes, ask a specialist subtask to compare changed APIs, keep a running evidence log, and pause for a reviewer before creating an internal summary. The work could span enough steps that the application must track progress and recover after a timeout.
The relevant benefit would not be that the assistant “thinks better.” It would be that the service already offers session management, resuming and steering, compaction and coordination. You would test whether those features remove enough integration work to justify the service dependency. The example is a proposed design, not a test we performed.
A smaller one-shot task is weaker evidence for adopting a managed runtime. If the prompt takes one input, calls one allowed function and returns a result, an extra session lifecycle may be needless. If your app already has an event loop and a mature job queue, you may already own the responsibilities the Agents API aims to manage. Compare against that working baseline, not against an imagined pile of custom infrastructure.
Long duration alone is not a sufficient reason either. A long-running task that waits for a human or an external event may be better represented as explicit states in your application. You might want to persist an approval decision, wait on a webhook, and then resume. The Agents API documents sessions and events, but your service still decides how these transitions fit the user's account, notification preferences and failure policy.
Ask three questions. First, does the task actually need iterative tool use or context that grows across turns? Second, will the managed session replace work your team must otherwise build and support? Third, can you still impose the controls that matter, such as a restricted tool list, visible progress, a hard stop and an inspectable result? If the answers are yes, the API is a serious candidate. If not, choose the smaller surface or keep the task manual.
View image detailWhen the SDK or direct Responses API is cleaner
If your product team needs to own each transition, the SDK may be a better starting point. The SDK gives you a framework for agent definitions, tools and handoffs while leaving the loop inside your application. That may fit a product with custom queues, application-specific session storage or established approval services. It also means your developers remain responsible for runtime operations and behavior. OpenAI's comparison guide describes that division; it does not guarantee that every SDK project takes a particular amount of work.
If you are not yet building an agent, a direct Responses API call may be clearer still. A task like “classify this incoming request against four categories and return a structured object” may not need a durable session, shell or subagents. If a tool call is needed, your application can define the specific operation. Then an ordinary test can verify the input, output and permission boundary.
The right baseline is the strongest simple solution that could solve the actual problem. That may be a query, a deterministic API, a small function, an existing product integration or a human review step. Adding an agent should solve a limitation in that path. An agent is not inherently better just because a new API is named for it.
There are also control and eligibility reasons to stay with another path. The current Agents API documentation says it supports U.S. data residency only and is not eligible for Zero Data Retention. It says that choosing a self-hosted sandbox does not make the Agents API ZDR-eligible. If a policy requires ZDR or residency outside the documented region, a self-hosted environment is not a workaround. Evaluate another product/API configuration or do not send that data through this service.
OpenAI documents WebSocket mode for long-running, tool-heavy workflows using the Responses API and says it supports store=false and ZDR. That does not make it a drop-in substitute for the Agents API's Codex harness. It is a separate way to manage a persistent connection and continue Responses work. If the underlying need is direct model/tool orchestration and the data terms matter, inspect that documented route with the same scrutiny.
The managed API does not take your application off duty
The application remains the part that turns model capability into a product. It decides which person can launch a session, which record the agent can see, which tools exist and what those tools are allowed to change. It also has to show the user what is happening, distinguish progress from completion and handle the point where a task requires approval.
That last distinction is practical. “The model asked to run a tool” and “the operation succeeded” are separate events. Your system should record the request, apply authorization outside the model's free-form answer, execute only the allowed action, and verify the changed state. A completed assistant message does not prove that an external system reached the intended state.
The harness may preserve context, but the application still needs to choose the source of truth. If a task receives an instruction in a chat, a customer record in a database and a policy from a document, which one wins when they disagree? How do you prevent a tool from reading a neighboring account? What does the app show when the agent cannot resolve a conflict? A model's summary of an earlier turn is not an audit log by itself.
Likewise, recovery has two meanings. The platform can recover or continue the session. Your application must decide whether an operation is safe to retry. A read-only query is often safe to repeat. A payment, publication, permission change or message send may not be. Your code needs an idempotency key, external state check or human confirmation where the operation demands it. That responsibility does not disappear when the loop is managed.
Build stop conditions into the application boundary. A user should be able to pause, cancel or ask what is pending. When the agent reaches an irreversible action, the product should present the exact action, target and evidence to a reviewer. A useful agent system can still say “I have a proposed update, but I have not applied it.” That is an honest intermediate state, not a product failure.
View image detailCompare the options on one fair task
Do not begin with a demo that gives one option a convenient prompt and another a production integration. Write an acceptance contract before you run anything. It should describe the starting state, source data, allowed actions, expected result and cases where the correct outcome is to stop. If the contract is vague, a fluent answer can hide a failed task.
Use the same task for each candidate. Keep the prompt, model where possible, tools, data and acceptance rules aligned. Candidates may have different session mechanics, but the product outcome should be the same. Include the best current manual or deterministic path as a baseline. A pilot can be small and still be fair.
For a hypothetical release-note comparison, the job might be: “Compare the two named official documents, identify changed requirements, cite the passages and produce a review draft. If a claim lacks direct support, mark it unresolved. Do not send or publish anything.” The result is accepted only if every material statement maps to source text, no unsupported claim is upgraded, links resolve, and the final artifact stays a draft. That is a proposal, not a run we performed.
Record more than whether the run completed. Track whether the answer met the rubric, whether the system stopped when required, number of retries, tool calls, tokens, container time if any, repair needed, reviewer minutes and whether the trace let someone reproduce the result. A candidate that returns faster but needs extensive cleanup has not necessarily won. One successful run is a clue, not a reliability rate.
Test ordinary and awkward cases. What happens when an input document is missing? What if two sources disagree? Does a tool time out? Can the user cancel during a long task? Does a retry duplicate a write? Can a reviewer see which source supported a conclusion? How does the application recover when a browser or environment disconnects? The failure path should be part of the task definition, not the small print after a showcase.
Then compare the work your team must maintain. With direct Responses, count your state machine, queues, retries, traces and environment management. With the SDK, count the app-owned runner and integrations. With the Agents API, count configuration, handlers, access policy, product UX and service-specific operations still required. Do not call one route “low-code” until the responsibilities and ongoing maintenance have been counted.
View image detail
View image detailCount cost and retention as architecture inputs
Model price is only one line. The total for an accepted result can include input and output tokens, repeated calls, tool use, sandbox time, logs, observability, data transfer, engineering effort, review and correction. If a workflow retries or sends too much context, the list price of one model call tells little about the delivered job.
You do not need to invent a cost estimate before testing. Start by logging usage and labeling accepted versus rejected outcomes. Count how much reviewer time follows each option. If container billing applies, record compute duration and configuration. If a tool provider charges separately, record that separately. Keep the measurement period and model/version attached so a later model or rate change does not quietly contaminate the comparison.
Also list what state exists and who can delete it. OpenAI's current Agents API guide says session state is retained so a task can continue, and sessions and published artifacts can be deleted. The data policy also says the API is not ZDR-eligible. Your team still needs to decide which inputs should enter a session, what application logs retain, what output artifacts are useful, and how deletion is requested and verified.
Do not collapse retention, residency and execution location into one checkbox. Residency identifies where a service's data processing or storage is covered under a policy. A self-hosted sandbox describes where code executes. ZDR eligibility is a specific service/data-control status. An application can run code in its infrastructure while the managed harness still processes session information under the API's documented terms. Verify the exact terms with current official docs and your organization's reviewer.
For a sensitive workflow, eligibility comes before a bake-off. If the data terms disqualify a route, do not spend effort comparing its completion rate. Remove it from the shortlist. If your policy allows the API, then decide what minimum information can be sent and where files, traces and outputs may live. Use synthetic or redacted data for early evaluation when that is enough to test the loop.
View image detail
View image detailA practical choice table
- One prompt, one structured answer: First route to evaluate: Responses API; Why: It avoids adding a durable orchestration layer before the task needs one; What must still be checked: Input/output validation, appropriate storage and policy
- A tool call with a fixed application-controlled sequence: First route to evaluate: Responses API or SDK; Why: The app can keep its explicit order and authorization; What must still be checked: Tool errors, state and retry rules
- Reusable agent definitions and handoffs within your product: First route to evaluate: Agents SDK; Why: The loop remains in your application while common agent patterns are packaged; What must still be checked: Runtime ownership, storage, evaluations and maintenance
- Multi-step work that needs persistent sessions and managed orchestration: First route to evaluate: Agents API; Why: The service may remove harness work that the team otherwise owns; What must still be checked: Public-beta status, service terms, app controls, actual cost and failure behavior
- Need ZDR or residency outside the currently documented supported region: First route to evaluate: Do not select Agents API on the assumption that self-hosting fixes it; Why: Current docs say it is non-ZDR and U.S.-only, including with self-hosted sandbox; What must still be checked: Verify another eligible configuration and current terms
Take a team that wants to turn source material into an internal draft. If it only needs to summarize one upload and return citations, a direct request could be enough. If an existing app coordinates reviewer approvals, preserving that state in the app may be simpler with its own loop. If researchers want to pause, add another source and resume a session while the product delegates independent comparisons, the managed harness becomes more relevant. Even then, the app must validate source grounding and keep publication under a separate approval.
That sequence is the useful outcome of the table: you should be able to explain what requirement makes the next layer worthwhile. “We want an AI agent” is not a requirement. “The task needs a resumable multi-step session and our team does not want to operate the context and orchestration loop” is much more testable.
Migrate in small, reversible steps
If the Agents API passes a small pilot, migrate one bounded task first. Keep the old route available while you compare output, failures and review work. Do not migrate every workflow just because the API can support it. Add only the tools needed for that job, and leave actions that can create harm behind a separate application approval.
Preserve the parts that make the application understandable: a versioned task definition, a stable acceptance rubric, source identifiers, tool permissions, a readable trace and a documented recovery policy. The exact conversation or harness internals may differ between systems, but your application can still keep the inputs, decisions, outputs and reviewer result needed to explain what happened.
Define an exit condition. If a supported feature changes, the service no longer fits a data requirement, or a task repeatedly fails its rubric, how can you turn the route off? Can a human do the work while the route is disabled? Can you export the useful artifact? Which logs are needed to explain a previous action? A migration is safer when the off-switch and fallback are designed alongside the first call.
The public beta label matters. OpenAI says the product will change while it gathers feedback. That can be a reason to experiment in a reversible, low-risk task. It is a reason not to make an untested dependency invisible inside a critical workflow. Read the current docs when implementation begins, record the date and test the behaviors your app relies on. A September 10 announcement is not a permanent API contract.
View image detailCommon questions
Is Agents API the same as the Agents SDK?
No. OpenAI describes Agents API as a service that manages a Codex-based harness and sessions. The SDK runs within your application and gives your app control over the agent loop and runtime integration. They can address related jobs, but their control boundaries differ. Check the current Agents guide before implementation.
Does the Agents API require a sandbox?
No. The documentation includes an environment option called none for work that can use configured tools or remote MCP without a shell and workspace. If the agent must execute code or work with files, choose an environment deliberately. A sandbox is an execution boundary, not a complete access policy.
Does self-hosting make Agents API eligible for ZDR?
No. Current OpenAI documentation says the Agents API is not ZDR-eligible, even when you choose a self-hosted sandbox. Verify current terms and your organization's requirements before sending production data.
Does the managed harness make an agent production-ready?
It supplies orchestration features. Your application still owns authorization, tool design, evaluation, user experience, business rules and safe handling of consequential actions. A vendor description or customer story does not establish that your own workflow is reliable.
The decision to make
Choose the Agents API when a durable managed harness solves a specific application problem and the current service terms fit your data. Choose the SDK when you want agent patterns but need your application to run the loop. Choose Responses API when a direct model/tool call or a small explicit loop is the simpler answer. Then compare your best candidates on one task with a written success rule and a safe failure path.
The useful thing is not to adopt the newest interface. It is to stop rebuilding work you do not need to own while keeping the controls your users and team do need. OpenAI's separate computer-use decision guide applies a similar task-first question to screen automation, but it covers a separate product update.
Checked for this article



