Operations / AgentOps: Deploy & Operate
AgentOps: Deploy & Operate
Running agents in production: rollout, versioning, scaling, idempotent retries, cost control, incident response.
- Durable State & ResumabilityMake the agent loop a durable computation — event-sourced history, journal-before-effect, and resume that replays rather than re-derives, so a crash or redeploy never restarts a half-done task.
- Concurrency, Queues & ScalingAgents are batch jobs, not requests: a queue with leased workers, per-tenant concurrency caps, journal-as-state for horizontal scale, and bounded fan-out are what survive production load.
- Idempotency, Retries & Side-Effect SafetyFour stacked retry sources mean every write tool will fire twice unless you construct exactly-once with intent-derived idempotency keys, failure classification, and a durable side-effect ledger.
- Cost Control at the Loop LevelAgent cost is unbounded by default; treat the per-task token/step/dollar ceiling as a fail-closed circuit breaker, then tune model cascades, prompt and tool caching, and early-exit against a quality metric.
- Rollout, Versioning & PinningBehavior is the (model, prompt, tools) triple; pin it to dated snapshots, stamp it on every run, and promote new versions only through shadow/canary plus an eval gate with instant config-flip rollback.
- Feature flags for agentsFlags scoped to prompts, models, tools, and policies — what to gate, how to roll, and why "off by default" is a non-negotiable for agent flags.
- Kill switchesA button that stops a running agent fleet — what it must actually stop (in-flight calls, queued work, scheduled retries), and how to test it before you need it.
- Incident Response & Runaway ContainmentA runaway agent fails open and keeps acting; detect from rate and progress, contain with in-loop fail-closed kill switches the resume path respects, rely on pre-installed blast-radius bounds, and turn every incident into a regression test.
- Rate Limits & Provider CapacityA 429 is a capacity contract, not a transient error, and retrying it turns a shortfall into an outage; agents burn tokens-per-minute quadratically, so the fix is a shared admission-control bucket, deliberate load shedding, and treating cross-provider failover as a behavior change your evals must cover.
- Model Deprecation & MigrationA model ID is a dependency with an expiry date you cannot vendor: notice floors of 60 days or less, why the swap is a re-qualification rather than a string replacement, silent platform auto-upgrades as the worst failure mode, and the generated inventory plus warm candidate lane that turn a retirement into a one-day operation.
- Self-Hosted Inference for AgentsLeaving the provider API changes the currency from tokens to KV-cache bytes: capacity is concurrent sequences times context length, prefix-cache hit rate is a routing problem rather than an engine one, autoscaling does not work at agent timescales, and the break-even is a utilisation number.
- Multi-Tenancy for AgentsAn agent adds five stores your row-level policy never touched — prompt cache, semantic cache, vector index, memory and the rate-limit pool — and only the semantic cache can hand one tenant another tenant’s answer.
- SLOs & Error Budgets for AgentsCorrectness fails every test an SLI must pass, so budget the mechanical indicators you can compute deterministically, run a second harm budget denominated in actions taken rather than requests served, and demote judge-scored quality to a control chart that never pages.
- Graceful Degradation & FallbackThe fallback is a different agent — different tool dialect, context ceiling and refusal profile — running your most traffic down your least-tested path under thresholds calibrated for a model that stopped answering: decide what to shed before what to swap, fail closed on side effects, and make degraded mode a named state with an exit.
- Exiting a Managed Agent RuntimeMaintenance mode promises existing workloads keep running, which is exactly what makes teams wait — the real deadline is the day the frozen model catalog stops carrying a model you need, and meanwhile the loop everyone assumes is the hard part ports in a sprint while the conversation state nobody inventoried is the migration: prove the export on day one, dual-write, then port code against a backlog that has stopped growing.
- Third-Party Tool DriftA tool description is part of your prompt and someone else owns the text, so behaviour changes with no commit and no error — the loud structural breaks are the harmless kind, while a reworded docstring moves tool-selection rates silently: snapshot the catalog, stamp its hash on every trace, and absorb the change in a facade rather than the prompt.
- Scheduled & Triggered AgentsWith no user present, ask becomes abstain, retry becomes reconcile, and the default failure is silence rather than an error — so put a dead-man’s switch on every schedule, end each firing in an acted / no-op / blocked verdict, and rate-limit the notification channel independently of the agent’s own judgement.
- Load-Testing an Agent SystemThe first decision is what to do about side effects, and every answer changes the measurement — mocked tools delete the seconds of real latency that dominate a trajectory. Size the run from Little’s Law, generate sampled tasks rather than one repeated prompt, grade quality under load because degradation is silent, and report the concurrency where success starts falling plus which dependency returned the first 429.
- Re-indexing & Embedding MigrationsAn embedding model is a schema with no in-place migration — two models occupy different spaces, so there is no canary and no dual-read, only a complete second index and an atomic cutover — and the rebuild that discovers your chunker drifted is the one where the quality delta can no longer be attributed to anything.
- Repairing What the Agent Already DidContainment fires in four seconds against a fault that landed three weeks ago, and almost nothing tells you what to do about the four thousand wrong actions already committed — each was a judgement rather than a row, so scope by decision instead of by record, sort the damage into recompute, compensate, irreversible and derived, replay against the pinned versions in force at the time, and run the backfill as a staged compute-then-apply migration keyed on the original run id.
- Staging Environments for AgentsMocking the tools deletes the latency, error shapes and schema drift you were trying to catch, while the model is the one component you can pin for free — so invert the instinct, run four fidelity tiers where each may only gate a release step whose failures it could have caught, and put the boundary at an egress write-blocker that returns a success the agent believes.
- Tool Catalog LifecycleSelection is a function of the whole catalog, so the fortieth tool changes behaviour on the thirty-nine tasks that were working and the definitions are a permanent per-step cost in every cached prefix — which makes the gate for adding a tool the existing eval set rather than a new one, scoping per task type worth more than dynamic tool retrieval, and removal a migration with a tombstone rather than a delete.
- Protocol Revisions & Deprecation WindowsA protocol revision is a dependency that expires on someone else’s calendar and sits on both ends of a connection you own one of, so there is no cutover — only a dual-revision window you run on purpose, and the input to every decision in it is a number almost nobody records: negotiated revision by share of traffic and by distinct caller. Normalise both revisions at the edge rather than forking the deployment, track deprecated capabilities on the calendar that holds certificate expiry, and choose libraries on their historical revision lag rather than their throughput.
- Long-Lived Sessions & Zero-Downtime DeploysRolling updates, connection draining and a thirty-second grace period were designed for sub-second requests; an agent session is a phone call, so your release cadence is now bounded by the p99 of your session-length distribution rather than by your pipeline. Pin the build to the session instead of migrating it, and expect the real breakage to be session state the new version cannot read — then bound the tail deliberately with a maximum session age derived from how long you are willing to drain.
- Sandbox Pools & Cold StartsEvery sub-100ms sandbox number is a snapshot restore measured one at a time, and neither half survives a fan-out: on an open benchmark one provider records an 83ms median sequentially and 14.8 seconds at concurrency 100, with the ordering between providers almost inverted. Warm and clean are one knob — Firecracker's own docs call resuming the same state more than once insecure, because entropy, identifiers and cached secrets repeat — so key the sandbox by your trust boundary, size the pool with Little's law at a hold time measured in minutes, and check whether your provider bills idle at all.
- Serving Agent TrafficYour API design assumes a frustrated client stops and a confused client reads documentation; an agent does neither, so a 429 is a pause rather than a signal and your error body is a prompt. Automated requests crossed 57.5% of HTML traffic in 2026 — label the classes before anything else, make errors machine-actionable, publish an idempotency contract for writes because the caller will retry, and price the operation rather than the session.
- Timeouts & Deadline BudgetsEvery timeout in your agent was chosen by someone who could not see the others, and their product is the worst case you ship: a 30-second tool ceiling in a twenty-step loop, under SDK defaults of ten minutes and two silent retries, is a run measured in hours — and a step that takes the slow path one time in a hundred makes a slow run one time in six. Pass an absolute deadline down the run and derive every timeout from what remains, the way gRPC does; then remember that a timeout abandons work rather than cancelling it, so the expiry path for a write is a reconcile by idempotency key, never a retry the model gets to propose.
- Fault Injection for Agent StacksWhen a dependency fails inside an ordinary service you get a 500; inside an agent you get a fluent wrong answer and a green dashboard, because the component handling the error is a model trained to keep going. Inject at the tool boundary rather than the network — empty success, slow-but-correct, error-in-a-200, stale data, mid-run credential expiry — with a seeded per-run fault plan, then assert on the trajectory and grade every run as correct-degraded, honest stop, or silent fabrication. Only the third blocks a release, and one mechanical check catches most of it: a write must never follow a faulted read in the same run.
- Unpinned Vendor DefaultsYour effective config is the union of what you set and what a provider, SDK, gateway and harness chose for you — and a pinned snapshot pins weights, not the fields your request omits: Claude Opus 5.5 defaults effort to medium where every other model defaults to high, so a model-string swap moved reasoning down a level with no diff. Log the resolved config as a fingerprint, assert it daily in CI against the live API, set explicitly whatever moves cost or tool-calling, and treat a default change as a release.
- Regional Failover for AgentsRTO and RPO assume the unit of recovery is a request, and an agent run is not one: when the region dies, a forty-minute task has already applied k of n external actions and k is unrecorded unless you wrote a side-effect ledger. Classify tasks as read-only, keyed or unkeyed and let the class decide resumption; separate admission control from the in-flight decision; and expect the real failure to be capacity, because caches are cold, rate limits are per-region and commitments may not follow you.
- Air-Gapped Agent DeploymentsEveryone asks where the model will run, which is the one question with a vendor answer; the expensive surprises are the implicit internet dependencies — package index, tool registry, hosted judge, telemetry, CRL — each of which fails inside the gap as an unexplained quality regression rather than a connection error. Separate sovereign-cloud from self-hosted from air-gapped before the architecture review, plan for a different model tier (IBM's self-hosted Bob kept the harness and ships Nemotron and Laguna, not the hosted frontier models), and accept the acceptance gate: a day of blocked egress in staging, counting confidently-wrong answers rather than failed tool calls.