AI Blog

The Agents API sells you the harness — compaction included

OpenAI opened the Agents API in public beta on 10 September, putting the managed Codex harness — sessions, subagent orchestration, recovery and context compaction — behind one API call, with no fee beyond tokens and containers. The compaction step is the part worth arguing about: it is the transformation that quietly rewrites what your agent is trying to do, and it now runs on a version you cannot pin, diff or roll back. Your eval numbers stop describing a system you control the moment you adopt it.

By Agentic AI Wiki 13 min read

The managed harness OpenAI opened to everyone on 10 September runs the one step in an agent loop that rewrites what the agent is trying to do — context compaction — and it runs it on a version number you will never see. Everything else in the Agents API is convenience you could have built. Compaction is the piece you are handing over, and it is the piece your evaluations were silently holding constant.

At a glance

The public beta splits an agent deployment into three ownership zones, and the split is not the one the managed-runtime products of the last two years drew.

ZoneWho runs itWhat lives thereCan you version it?
Capability You Your tools, your MCP servers, your prompts, your data access. Yes — it is your code.
Harness OpenAI Session lifecycle, subagent orchestration, recovery, context compaction. No.
Execution OpenAI, a partner, or you The sandbox the agent writes files and runs commands in. Yes — you choose the provider and the image.
Three ownership zones in an Agents API deployment Three vertical zones. On the left, Capability, owned by you: tools, MCP servers, prompts and data access. In the centre, drawn as a solid accent block, the Harness owned by OpenAI: session lifecycle, subagent orchestration, recovery and context compaction, labelled as carrying no version you can pin. On the right, Execution: a sandbox that may be OpenAI-hosted, supplied by one of nine partners, or run on your own infrastructure. Arrows run from the harness out to the capability zone and back, and from the harness to the execution zone. Who runs which part of the loop Yours OpenAI's Your choice Capability Your tools Your MCP servers Your prompts Your data access Versioned in your repo Diffable, revertible Harness Session lifecycle Subagent orchestration Failure recovery Context compaction No version you can pin No diff, no rollback Execution OpenAI-hosted, or one of nine partners, or your own VPC Blaxel · Cloudflare · Daytona DigitalOcean · E2B · Modal Oracle · Runloop · Vercel The step that rewrites the objective Compaction paraphrases the transcript — including the sentence stating the goal — every time the window fills. It runs inside the accent block, on a version you never see.
Two of the three zones stayed yours. The one in the middle is the one that transforms your context.

What actually shipped

On 10 September 2026 OpenAI put the Agents API into public beta. The pitch is a single call that runs the same managed harness behind Codex and ChatGPT for Work: durable sessions that continue work across turns, streaming progress, your own tools and MCP servers, coordination of subagents, automatic compaction of earlier context as a session approaches the window limit, and recovery when a step fails.

Execution is deliberately unbundled. An Environment — the optional sandbox where the agent runs code, edits files and reads output — can be OpenAI-hosted, supplied from your own infrastructure, or drawn from a list of nine partners: Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop and Vercel.

There is no fee for the Agents API itself during the beta. You pay for model tokens, for tools, and — if you use the hosted sandbox — for container time, priced per twenty-minute session by memory tier and billed per minute with a five-minute minimum for eligible sessions.

Worth keeping straight, because the names collide: an Agents API session, an Agents SDK session, a Responses conversation and a sandbox are four different resources. The Responses API is a primitive you compose into a loop. The Agents API is the loop.

Somebody finally sold the loop

Where the agent loop runs, across three generations of product Three columns. Roll your own: you write the loop and own compaction, and every version in the stack can be pinned. Managed agent runtimes such as AgentCore, Foundry and Vertex AI Agent Engine: they sell identity, memory, a tool gateway and traces around a loop you still write yourself. The Agents API: OpenAI writes and runs the loop, and you supply tools and choose where code executes. Who writes the control loop Roll your own You write the loop. You own the compaction prompt and its policy. Managed runtimes Identity, memory, gateway, traces — sold around a loop you still write. Agents API OpenAI writes and runs the loop. You supply tools and pick the sandbox. Version tuple Version tuple Version tuple Model, prompt, harness, judge — all pinnable. All pinnable; the ring services carry SLAs. One slot reads "current", permanently.
The managed-runtime generation sold everything around the loop. This one sells the loop.

A month ago the honest summary of the managed agent platforms was that none of them would sell you the loop. AgentCore, Foundry, Vertex AI Agent Engine and Cloudflare Agents each sold a ring of services — identity, memory, a tool gateway, traces, durable execution — around a control loop you still had to write, and the reason was not shyness. The loop is where a vendor's opinions become your behaviour, and nobody wanted to own that.

The Agents API owns it. That is a genuine product decision and it removes real work: session durability, retries, subagent fan-out and the long tail of "what happens when the model returns garbage at step forty" are all things teams rebuild badly. If you have written an agent loop twice you know the second one was not better, just yours.

But the loop is not a uniform thing you either own or do not. It is a sequence of steps, and they differ enormously in how much they change the meaning of what is running. Retry policy is a knob. Streaming is plumbing. Subagent fan-out is an architecture you can reason about from the outside, because you can see the subagent's inputs and outputs. Compaction is none of those.

Compaction is the step that changes the objective

Every long-running agent eventually exceeds its window, and every implementation answers with some form of lossy rewrite: summarise the older turns, drop tool results that look spent, keep what appears load-bearing. The wiki's context compaction page treats this as a memory-hierarchy problem, which is how it is usually framed. The framing understates it.

Compaction is the only step in the loop that rewrites the agent's own statement of what it is doing. The objective arrived as text in the transcript. After compaction it is a paraphrase of that text, produced by a model, under a summarisation prompt you did not write. Do that four times across a long session and the agent is pursuing a fourth-generation restatement of your instruction — competently, confidently, and against acceptance criteria nobody can now recover.

This is the primary mechanism behind goal drift, and drift is specifically the failure that survives outcome evaluation: the final artefact is a good answer to the question the agent ended up asking. You cannot catch it by grading the result, because the result is fine on its own terms. You catch it by pinning the objective outside the transcript and re-injecting it verbatim — which is a thing you do to your own compaction step.

The concrete test: open a long session's transcript at the point where compaction fired, and find the sentence that states the goal. If it is a paraphrase rather than your original string, you have a drift surface. If you cannot open the transcript at that point at all, you have a drift surface you cannot measure.

So the question to ask about a managed harness is not "is its compaction good?" It is: can I tell when it changed? A harness version is not a model version. It can ship on a Tuesday, improve the median case, and move your agent's behaviour on the long tail — the very sessions where compaction fires — without a single identifier in your logs changing. Every eval score you have carries an implicit tuple of model version, prompt version, harness version and judge version, and the whole discipline of rollout and versioning exists to keep that tuple explicit. The beta hands you a tuple with one slot permanently marked "current".

None of this makes the product a mistake. It makes it a trade with a specific price, and the price is charged in your ability to attribute a regression. When your pass rate drops four points next month, the candidate explanations are your prompt, your tools, the model, and a fourth thing you cannot inspect, cannot pin and cannot roll back.

The harness is free, which is the tell

Hosted sandbox price per twenty-minute session by memory tier Horizontal bar chart with four bars. One gigabyte costs three cents per twenty-minute session, four gigabytes twelve cents, sixteen gigabytes forty-eight cents, and sixty-four gigabytes one dollar and ninety-two cents. Bar lengths scale linearly with memory, so the price is a flat rate of roughly three cents per gigabyte per session. USD per 20-minute container session 1 GB $0.03 4 GB $0.12 16 GB $0.48 64 GB $1.92 $0.48 $0.96 $1.44 $1.92 Flat ~$0.03 per GB per session. Eligible sessions bill per minute, five-minute minimum.
Container time is the only line item the Agents API adds. The harness itself is priced at zero.

Read the pricing as a design statement rather than a discount. Container tiers scale linearly with memory — roughly three cents per gigabyte per twenty-minute session — and per-minute billing with a five-minute floor is a direct answer to the objection that a sandbox is idle most of the time. That part is well judged.

The harness, meanwhile, costs nothing and is paid for in tokens. That alignment is worth stating plainly: a harness that compacts later, keeps more history, or retries more often bills you more and earns the vendor more, and there is no line item that would let you notice. This is not an accusation — aggressive compaction also degrades quality, so the incentives are not purely one way — but it is the reason a free component is never free to reason about. The only defence is a number you compute yourself: tokens per completed task, tracked per harness change, which is the same discipline agent unit economics asks for and which almost nobody runs against a component they did not know had versions.

Nine sandbox partners is not generosity

The unbundled execution layer is the most interesting structural choice in the release, and it cuts against the obvious read of the news. A vendor consolidating the stack would host the sandbox and be done. Instead OpenAI named nine sandbox providers and let you bring your own infrastructure.

The reason is that compute is where the regulated objections live. Data residency, network egress, VPC peering, an auditor asking which jurisdiction a file was written in — those are the conversations that stall an enterprise deal, and they are all about where the code ran, not about who ordered it to run. Handing that layer to partners removes the blocker while keeping the layer that actually differentiates.

Which tells you what OpenAI thinks the durable asset is, and it agrees with the benchmark literature. The last year of agent evaluation has kept arriving at the same uncomfortable finding: a large share of measured agent performance belongs to the scaffold rather than the model — harness-inclusive scores and open trace corpora both landed there. If the scaffold carries the performance, the scaffold is the product. The Agents API is that thesis shipped as an endpoint.

When to take the trade

SituationTake the managed harnessKeep your own loop
Sessions are short and rarely compact Yes — the risky step almost never fires. Only if you have already built it.
Long autonomous runs, hours of horizon Only with objective re-injection you control. Yes — compaction policy is a product decision here.
You gate releases on an eval suite Only if you also pin a canary and re-baseline often. Yes — you need the version tuple complete.
Regulated data, strict residency Yes, with your own or a partner sandbox. Only if the harness must also stay inside.
You are prototyping Yes, obviously. No.

If you adopt it, three things are worth doing in the first week, and they are cheap. Put your objective somewhere the harness cannot paraphrase — a tool the agent must call to re-read its acceptance criteria works, because a tool result arrives fresh in every window. Keep one frozen scenario you re-run daily against a fixed model and fixed prompts, so that a silent harness change shows up as an unexplained movement on a line you control. And log tokens per completed task from day one, because that series is the only early warning you get about a component whose behaviour you cannot read.

The exit question is the one to answer before adoption rather than during an incident: what does leaving look like? Exiting a managed agent runtime is mostly a data-gravity problem, and here the gravity is unusual — it is not your data that is captive, it is your calibration. Tools port in an afternoon. The knowledge of how your agent behaves at hour three does not.

FAQ

Is the Agents API a replacement for the Responses API?

No. The Responses API is a primitive you assemble into a loop; the Agents API runs the loop for you. They are separate resources with separate session objects, and an agent built on Responses does not automatically become an Agents API agent.

Does using the Agents API mean my code runs on OpenAI infrastructure?

Not necessarily. Execution is unbundled: you can use the OpenAI-hosted sandbox, one of nine named partners including Cloudflare, E2B, Modal and Vercel, or your own infrastructure. The harness runs at OpenAI regardless — that is the part you cannot relocate.

Why single out compaction rather than the whole harness?

Because the other steps are observable from outside. You can see what a retry retried and what a subagent was asked. Compaction rewrites the transcript itself, including the sentence stating the goal, so its effects are invisible in exactly the runs where they matter most.

What does the beta pricing actually cost me?

Model tokens and tools as usual, plus hosted-container time if you use the OpenAI sandbox — priced by memory tier per twenty-minute session, billed per minute with a five-minute minimum on eligible sessions. There is no separate charge for the API during the beta.

If I already run my own loop, is there any reason to switch?

Subagent orchestration and session durability are real work, and if yours are shaky the trade may be worth it. But do not switch during a quarter when you are also changing models or prompts — you will lose the ability to attribute whatever happens next.

Further reading

On this wiki:

Sources: