Managed Agents

Anthropic’s Managed Agents article is about infrastructure for long-horizon agent work. As models improve, the assumptions encoded in harnesses go stale. Managed Agents is designed around a small set of stable interfaces that can outlast any particular harness implementation.

The design follows an old computing problem: how to build systems for “programs as yet unthought of.” Operating systems solved this by virtualizing hardware into abstractions such as processes and files. Managed Agents applies the same idea to agent infrastructure by virtualizing Session, Harness, and Sandbox.

Harness Assumptions Go Stale

Anthropic’s agent engineering work repeatedly shows that harnesses encode assumptions about what Claude cannot do on its own.

Those assumptions need to be questioned because they can become obsolete as models improve.

One example is context anxiety. In earlier work, Claude Sonnet 4.5 would wrap up tasks prematurely as it sensed the context limit approaching. Anthropic addressed this with context resets. But when the same harness ran on Claude Opus 4.5, the behavior was gone, and the reset logic became dead weight.

Managed Agents is Anthropic’s hosted service for running long-horizon agents through interfaces meant to outlast any specific implementation, including the ones Anthropic runs today.

Managed Agents virtualizes three components:

AbstractionMeaning
SessionThe append-only log of everything that happened.
HarnessThe loop that calls Claude and routes Claude’s tool calls to infrastructure.
SandboxThe execution environment where Claude can run code and edit files.

Each implementation can be swapped without disturbing the others. The system is opinionated about interface shape, not about what runs behind the interface.

Do Not Adopt A Pet

The first architecture placed all agent components into a single container: session, harness, and sandbox shared one environment.

This had benefits. Model-generated code could modify a file through direct syscalls, and there were no service boundaries to design.

But coupling everything into one container created a classic infrastructure problem: the system had adopted a pet rather than cattle. A pet is a named, hand-tended individual you cannot afford to lose. Cattle are standardized, interchangeable instances.

In the early architecture, the container became the pet. If it failed, the session was lost. If it became unresponsive, engineers had to nurse it back to health.

The Coupled Architecture Single container — every failure is total, every component inseparable Single Container Session (state) Harness (logic) Sandbox (execution) Container Crash → all state lost → no recovery possible Debug Access → shell into container → privacy compromise VPC Integration → network peering → architecture debt Each container is a pet — unique, irreplaceable, and fragile

This caused several problems:

  • Failures were hard to distinguish: the WebSocket event stream could not tell whether a failure came from a harness bug, packet drop, or container going offline.
  • Debugging conflicted with privacy: troubleshooting required shell access inside containers that often held user data.
  • Infrastructure location was baked in: the harness assumed everything Claude worked on lived in the same container. Customers wanting VPC access had to peer networks with Anthropic or run the harness in their own environment.

The root cause was that Session, Harness, and Sandbox shared one fate.

Decouple The Brain From The Hands

Anthropic’s solution was to decouple the brain from the hands:

  • Brain: Claude and its harness.
  • Hands: sandboxes and tools that perform actions.
  • Session: the log of session events.

Each became an interface with few assumptions about the others. Each could fail or be replaced independently.

Decoupled Architecture Stateless brain, durable log, disposable hands — replaceable by design Session Durable Event Log • Append-only • Source of truth • Survives harness/sandbox failure getSession(id) Harness Brain · Stateless • Calls model • Routes tool calls • Zero local state wake(sessionId) Sandbox Hands · Disposable • Runs code • Isolated from secrets • Interchangeable execute(name, input) getEvents emitEvent execute result Structural Guarantees ① Crash recovery — new Harness reads the session log, resumes from last event ② Security — credentials never enter the sandbox; secrets flow through a proxy ③ Composability — any brain can use any hands; hands can serve multiple brains

The Harness Leaves The Container

Decoupling means the harness no longer lives inside the container. It calls the container like any other tool:

execute(name, input) -> string

The container becomes cattle. If it dies, the harness catches the failure as a tool-call error and passes it to Claude. Claude can decide whether to retry. If it retries, the system can initialize a fresh container with a standard recipe:

provision({resources})

Failed containers no longer need to be nursed back to health.

Recovering From Harness Failure

The harness also becomes cattle. Because the session log sits outside the harness, nothing in the harness needs to survive a crash.

When a harness fails, a new harness can:

  1. Start with wake(sessionId).
  2. Retrieve the event log with getSession(id).
  3. Resume from the last event.
  4. Write durable events during the loop with emitEvent(id, event).

Persistent state lives in the Session, so the harness can be replaced at any time.

Security Boundary

In the coupled design, untrusted code generated by Claude ran in the same container as credentials. A prompt injection only needed to convince Claude to read environment variables to access tokens. Once an attacker had those tokens, they could spawn new unrestricted sessions and delegate work to them.

Narrowing token scopes helps, but it still encodes an assumption about what Claude cannot do with a limited token. As models become smarter, that assumption becomes less stable.

The structural fix is: tokens are never reachable from the sandbox where Claude-generated code runs.

Security Boundary — Credential Isolation Tokens never enter the execution environment — structural, not policy Harness Brain Proxy Auth gateway Vault Credential store ① session ID ② credential ③ Proxy calls external service with credential → returns result to Harness Structural Isolation Boundary Sandbox Code runs here ✕ Never reachable Vault Secrets live here Prompt injection with sandbox control → no direct credential access

Managed Agents uses two patterns:

  • Auth bundled with a resource: for Git, the repository access token is used during sandbox initialization to clone the repo and wire it into the local git remote. git push and git pull work from inside the sandbox, but the agent never handles the token.
  • Vault outside the sandbox: for custom tools and MCP, OAuth tokens live in a secure vault. Claude calls MCP tools through a dedicated proxy. The proxy receives a session-associated token, fetches credentials from the vault, calls the external service, and returns the result to the harness. The harness never sees the credentials.

The Session Is Not Claude’s Context Window

Long-horizon tasks often exceed Claude’s context window. Standard approaches require irreversible decisions about what to keep and what to discard.

Anthropic has explored these techniques in context engineering. Compaction lets Claude save a summary of its context window. The memory tool lets Claude write context to files for learning across sessions. These can be combined with context trimming, which selectively removes tokens such as old tool results or thinking blocks.

But irreversible retention or deletion can fail. It is hard to know which tokens future turns will need. If messages are transformed by compaction and removed from Claude’s context window, they are recoverable only if stored elsewhere.

In Managed Agents, the Session provides the same benefit as a context object outside the model context window. But instead of living in the sandbox or a REPL, context is durably stored in the Session log. The getEvents() interface lets the brain query positional slices of the event stream:

  • Continue from wherever it last stopped reading.
  • Rewind before a specific event to see the lead-up.
  • Re-read context before a specific action.

The harness can also transform fetched events before passing them into Claude’s context window. Those transformations can include prompt-cache-friendly organization or other context engineering strategies.

This separates two concerns:

  • Session provides durable, recoverable context storage.
  • Harness performs arbitrary context management, deciding which events enter Claude’s context window and how they are organized.

Anthropic designed it this way because they cannot predict which context engineering strategies future models will need. The interface guarantees that the Session is durable and queryable; context management remains in the harness.

Many Brains, Many Hands

Many Brains

Decoupling the brain from the hands solved an early customer problem. When teams wanted Claude to work against resources in their own VPC, the original container design forced network peering because the container holding the harness assumed every resource was nearby. Once the harness left the container, that assumption disappeared.

The same change improved performance.

In the early architecture, putting the brain in a container meant every brain required a container. No inference could happen until that container was provisioned. Every session paid the full setup cost up front: clone the repo, boot processes, and fetch pending events, even if the session never touched the sandbox.

That dead time shows up as time-to-first-token (TTFT), the delay between accepting work and producing the first response token.

After decoupling, containers are provisioned only when the brain calls a tool and needs one. Sessions that do not need a container immediately do not wait for one. Inference can start as soon as the orchestration layer pulls pending events from the Session log.

Result: p50 TTFT dropped roughly 60%, and p95 dropped over 90%.

Scaling many brains becomes starting many stateless harnesses and connecting them to hands only when needed.

Many Hands

Anthropic also wanted each brain to connect to many hands. In practice, Claude must reason about many execution environments and decide where to send work. That is harder than operating in a single shell.

The early single-container design existed because earlier models were not capable of this. As intelligence scaled, the single container became the limitation: if it failed, state for every hand the brain was reaching into was lost.

After decoupling, each hand is a tool with the same interface:

execute(name, input) -> string

This supports custom tools, MCP servers, and Anthropic’s own tools. The harness does not need to know whether the sandbox is a container, phone, browser, or another environment.

Because no hand is coupled to a specific brain, brains can pass hands to one another.

Conclusion: Managed Agents As Meta-Harness

Managed Agents faces the old systems problem of designing for programs that do not exist yet. Operating systems lasted by virtualizing hardware into abstractions general enough for future programs. Managed Agents aims to do the same for future harnesses, sandboxes, and other components around Claude.

Managed Agents is not one specific harness. It is a meta-harness: a system of general interfaces that can accommodate many harnesses.

Claude Code is one harness Anthropic uses widely. Task-specific harnesses can excel in narrow domains. Managed Agents is meant to accommodate these forms and evolve with Claude’s intelligence over time.

Meta-harness design means:

  • Being opinionated about interfaces around Claude: Claude needs to manipulate state, the Session, and perform computation, the Sandbox.
  • Expecting Claude to scale to many brains and many hands.
  • Designing interfaces that can run reliably and securely over long horizons.
  • Making no assumptions about how many brains or hands Claude will need, or where they should live.

Source

Scaling Managed Agents: Decoupling the Brain from the Hands — Lance Martin, Gabe Cemaj & Michael Cohen, Anthropic, April 2026.

Was this page helpful?