The managed harness OpenAI opened to everyone on 10 September runs the one step in an agent loop that rewrites what the agent is trying to do — context compaction — and it runs it on a version number you will never see. Everything else in the Agents API is convenience you could have built. Compaction is the piece you are handing over, and it is the piece your evaluations were silently holding constant.
At a glance
The public beta splits an agent deployment into three ownership zones, and the split is not the one the managed-runtime products of the last two years drew.
| Zone | Who runs it | What lives there | Can you version it? |
|---|---|---|---|
| Capability | You | Your tools, your MCP servers, your prompts, your data access. | Yes — it is your code. |
| Harness | OpenAI | Session lifecycle, subagent orchestration, recovery, context compaction. | No. |
| Execution | OpenAI, a partner, or you | The sandbox the agent writes files and runs commands in. | Yes — you choose the provider and the image. |
What actually shipped
On 10 September 2026 OpenAI put the Agents API into public beta. The pitch is a single call that runs the same managed harness behind Codex and ChatGPT for Work: durable sessions that continue work across turns, streaming progress, your own tools and MCP servers, coordination of subagents, automatic compaction of earlier context as a session approaches the window limit, and recovery when a step fails.
Execution is deliberately unbundled. An Environment — the optional sandbox where the agent runs code, edits files and reads output — can be OpenAI-hosted, supplied from your own infrastructure, or drawn from a list of nine partners: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop and Vercel.
There is no fee for the Agents API itself during the beta. You pay for model tokens, for tools, and — if you use the hosted sandbox — for container time, priced per twenty-minute session by memory tier and billed per minute with a five-minute minimum for eligible sessions.
Worth keeping straight, because the names collide: an Agents API session, an Agents SDK session, a Responses conversation and a sandbox are four different resources. The Responses API is a primitive you compose into a loop. The Agents API is the loop.
Somebody finally sold the loop
A month ago the honest summary of the managed agent platforms was that none of them would sell you the loop. AgentCore, Foundry, Vertex AI Agent Engine and Cloudflare Agents each sold a ring of services — identity, memory, a tool gateway, traces, durable execution — around a control loop you still had to write, and the reason was not shyness. The loop is where a vendor's opinions become your behaviour, and nobody wanted to own that.
The Agents API owns it. That is a genuine product decision and it removes real work: session durability, retries, subagent fan-out and the long tail of "what happens when the model returns garbage at step forty" are all things teams rebuild badly. If you have written an agent loop twice you know the second one was not better, just yours.
But the loop is not a uniform thing you either own or do not. It is a sequence of steps, and they differ enormously in how much they change the meaning of what is running. Retry policy is a knob. Streaming is plumbing. Subagent fan-out is an architecture you can reason about from the outside, because you can see the subagent's inputs and outputs. Compaction is none of those.
Compaction is the step that changes the objective
Every long-running agent eventually exceeds its window, and every implementation answers with some form of lossy rewrite: summarise the older turns, drop tool results that look spent, keep what appears load-bearing. The wiki's context compaction page treats this as a memory-hierarchy problem, which is how it is usually framed. The framing understates it.
Compaction is the only step in the loop that rewrites the agent's own statement of what it is doing. The objective arrived as text in the transcript. After compaction it is a paraphrase of that text, produced by a model, under a summarisation prompt you did not write. Do that four times across a long session and the agent is pursuing a fourth-generation restatement of your instruction — competently, confidently, and against acceptance criteria nobody can now recover.
This is the primary mechanism behind goal drift, and drift is specifically the failure that survives outcome evaluation: the final artefact is a good answer to the question the agent ended up asking. You cannot catch it by grading the result, because the result is fine on its own terms. You catch it by pinning the objective outside the transcript and re-injecting it verbatim — which is a thing you do to your own compaction step.
The concrete test: open a long session's transcript at the point where compaction fired, and find the sentence that states the goal. If it is a paraphrase rather than your original string, you have a drift surface. If you cannot open the transcript at that point at all, you have a drift surface you cannot measure.
So the question to ask about a managed harness is not "is its compaction good?" It is: can I tell when it changed? A harness version is not a model version. It can ship on a Tuesday, improve the median case, and move your agent's behaviour on the long tail — the very sessions where compaction fires — without a single identifier in your logs changing. Every eval score you have carries an implicit tuple of model version, prompt version, harness version and judge version, and the whole discipline of rollout and versioning exists to keep that tuple explicit. The beta hands you a tuple with one slot permanently marked "current".
None of this makes the product a mistake. It makes it a trade with a specific price, and the price is charged in your ability to attribute a regression. When your pass rate drops four points next month, the candidate explanations are your prompt, your tools, the model, and a fourth thing you cannot inspect, cannot pin and cannot roll back.
The harness is free, which is the tell
Read the pricing as a design statement rather than a discount. Container tiers scale linearly with memory — roughly three cents per gigabyte per twenty-minute session — and per-minute billing with a five-minute floor is a direct answer to the objection that a sandbox is idle most of the time. That part is well judged.
The harness, meanwhile, costs nothing and is paid for in tokens. That alignment is worth stating plainly: a harness that compacts later, keeps more history, or retries more often bills you more and earns the vendor more, and there is no line item that would let you notice. This is not an accusation — aggressive compaction also degrades quality, so the incentives are not purely one way — but it is the reason a free component is never free to reason about. The only defence is a number you compute yourself: tokens per completed task, tracked per harness change, which is the same discipline agent unit economics asks for and which almost nobody runs against a component they did not know had versions.
Nine sandbox partners is not generosity
The unbundled execution layer is the most interesting structural choice in the release, and it cuts against the obvious read of the news. A vendor consolidating the stack would host the sandbox and be done. Instead OpenAI named nine sandbox providers and let you bring your own infrastructure.
The reason is that compute is where the regulated objections live. Data residency, network egress, VPC peering, an auditor asking which jurisdiction a file was written in — those are the conversations that stall an enterprise deal, and they are all about where the code ran, not about who ordered it to run. Handing that layer to partners removes the blocker while keeping the layer that actually differentiates.
Which tells you what OpenAI thinks the durable asset is, and it agrees with the benchmark literature. The last year of agent evaluation has kept arriving at the same uncomfortable finding: a large share of measured agent performance belongs to the scaffold rather than the model — harness-inclusive scores and open trace corpora both landed there. If the scaffold carries the performance, the scaffold is the product. The Agents API is that thesis shipped as an endpoint.
When to take the trade
| Situation | Take the managed harness | Keep your own loop |
|---|---|---|
| Sessions are short and rarely compact | Yes — the risky step almost never fires. | Only if you have already built it. |
| Long autonomous runs, hours of horizon | Only with objective re-injection you control. | Yes — compaction policy is a product decision here. |
| You gate releases on an eval suite | Only if you also pin a canary and re-baseline often. | Yes — you need the version tuple complete. |
| Regulated data, strict residency | Yes, with your own or a partner sandbox. | Only if the harness must also stay inside. |
| You are prototyping | Yes, obviously. | No. |
If you adopt it, three things are worth doing in the first week, and they are cheap. Put your objective somewhere the harness cannot paraphrase — a tool the agent must call to re-read its acceptance criteria works, because a tool result arrives fresh in every window. Keep one frozen scenario you re-run daily against a fixed model and fixed prompts, so that a silent harness change shows up as an unexplained movement on a line you control. And log tokens per completed task from day one, because that series is the only early warning you get about a component whose behaviour you cannot read.
The exit question is the one to answer before adoption rather than during an incident: what does leaving look like? Exiting a managed agent runtime is mostly a data-gravity problem, and here the gravity is unusual — it is not your data that is captive, it is your calibration. Tools port in an afternoon. The knowledge of how your agent behaves at hour three does not.
FAQ
Is the Agents API a replacement for the Responses API?
No. The Responses API is a primitive you assemble into a loop; the Agents API runs the loop for you. They are separate resources with separate session objects, and an agent built on Responses does not automatically become an Agents API agent.
Does using the Agents API mean my code runs on OpenAI infrastructure?
Not necessarily. Execution is unbundled: you can use the OpenAI-hosted sandbox, one of nine named partners including Cloudflare, E2B, Modal and Vercel, or your own infrastructure. The harness runs at OpenAI regardless — that is the part you cannot relocate.
Why single out compaction rather than the whole harness?
Because the other steps are observable from outside. You can see what a retry retried and what a subagent was asked. Compaction rewrites the transcript itself, including the sentence stating the goal, so its effects are invisible in exactly the runs where they matter most.
What does the beta pricing actually cost me?
Model tokens and tools as usual, plus hosted-container time if you use the OpenAI sandbox — priced by memory tier per twenty-minute session, billed per minute with a five-minute minimum on eligible sessions. There is no separate charge for the API during the beta.
If I already run my own loop, is there any reason to switch?
Subagent orchestration and session durability are real work, and if yours are shaky the trade may be worth it. But do not switch during a quarter when you are also changing models or prompts — you will lose the ability to attribute whatever happens next.
Further reading
On this wiki:
- Agent harness — what the scaffold around a model actually does, and why it moves scores.
- Goal drift — why a long run substitutes an easier objective and pursues it competently.
- Context compaction & hierarchical memory — the mechanics of the rewrite.
- Managed agent runtimes — the category this release redefines.
- Exiting a managed agent runtime — what leaving costs, planned in advance.
- Claude Managed Agents: architecture — the same trade drawn by a different vendor.
Sources:
- OpenAI release notes — Agents API public beta, 10 September 2026.
- OpenAI API docs — Agents guide — sessions, environments, subagents, compaction.
- Introducing the Agents API and hosted sandboxes — OpenAI developer community announcement.