Execution Environment and Sandbox (E)

This page corresponds to §3 of Agent Harness Engineering: A Survey. Execution is the physical substrate of an agent harness: the environment in which actions actually run. In LLM-agent systems, execution and sandboxing are tightly coupled because the agent may write files, run shell commands, install packages, browse the web, or operate a desktop.

Why Sandboxing Is First-Class

The paper argues that sandboxing serves three purposes in agent systems:

  • Security: model-generated code is too large and too dynamic to rely on static review. Multi-step autonomy means harmful actions may happen without a human in the loop. Prompt injection can also redirect an otherwise benign agent into an attack carrier.
  • Reproducibility: long-horizon tasks and evaluations need resettable execution state. SWE-bench, OSWorld, and similar benchmarks are practical because the environment can be restored to a known baseline.
  • Liveness: without a sandbox, every risky action needs a permission prompt. At scale, users either abandon the workflow or reflexively approve everything. A sandbox creates a bounded area where the agent can act freely.

The paper’s useful phrase is that a sandbox is both a cage and a license. It restricts blast radius, but it also makes autonomous execution possible.

Sandbox Categories

The survey classifies sandboxes by workload rather than by isolation primitive:

CategoryRepresentative systemsDesign signal
General-purpose managed sandboxesDaytona, E2B, Modal, Northflank, OpenSandbox, Docker SandboxesAPI-exposed execution for arbitrary OCI images; increasingly moving toward dedicated-kernel microVMs.
Computer-use infrastructureAnthropic Computer Use, CUA, OSWorldFull GUI/OS interaction through mouse, keyboard, and screenshots; high fidelity, high attack surface.
Code and repository execution sandboxesJudge0, OpenAI Code Interpreter, sandboxed.sh, langchain-sandbox, Repo2RunLanguage/runtime preinstallation, high concurrency, often stateless and request-scoped.
Framework-integrated runtimesOpenHands, GoEX, agent-infra sandbox, smolagents executorsBundled into an orchestration loop; convenient but tightly coupled.
Browser evaluation environmentsWebArena, VisualWebArena, BrowserGym, WorkArenaBoth sandbox and benchmark substrate; natural surface for indirect prompt injection.
OS-level permission sandboxesAnthropic sandbox-runtime, Claude Code sandboxing, IsolateGPT, AgentBound, transactional sandboxingNarrow the host view rather than creating a full new OS image.
Sandbox abstraction layersSWE-ReX, smolagents executor, K8s Agent Sandbox CRD, R2E-Gym, EnvScaler, SandMLEUnify multiple execution backends so agent logic is decoupled from where it runs.

Isolation technology is orthogonal: containers, gVisor-style user-space kernels, Firecracker/Kata microVMs, WebAssembly, bubblewrap, Seatbelt, seccomp, and full VMs can appear under multiple workload categories.

Threat Model

The agent setting amplifies traditional sandbox threats:

  • prompt injection can cause the agent to run attacker-directed actions;
  • goal misalignment can make sandbox escape instrumentally useful to the agent;
  • compositional amplification lets one weak tool combine with others into a larger breach.

SandboxEscapeBench reports that frontier models can exploit Docker sandbox weaknesses under realistic configurations. This makes sandbox escape an agent-relevant threat today, not a purely theoretical concern.

The paper distinguishes infrastructure isolation from semantic/capability isolation. A container bounds damage after an action runs. A tool permission policy decides whether the action should have been allowed at all. Robust harnesses need both.

Deployment Modes

ModeStrengthCost
Self-hostedLowest latency, tight local iteration, strong control over dataThe developer owns security and operations.
Cloud/SaaSElastic scale and managed isolationNetwork round trips and provider trust.
Hybrid/BYOCKeeps data locality while using remote capacityMore complex identity, routing, and audit boundaries.

Compliance, data residency, and auditability often push systems toward hybrid deployment even when latency and scale would otherwise favor a single mode.

Open Problem

Execution environments need to become both measurable and composable. The paper highlights three gaps:

  • common security evaluations for prompt injection, goal misalignment, and compositional amplification;
  • cheap reset and replay for large-scale training/evaluation, including learned surrogate environments whose fidelity still needs to be proven;
  • portability layers that preserve semantics across Linux containers, macOS, Windows, browser environments, desktop VMs, cloud, self-hosted, and hybrid deployments.

The bundle-vs-compose choice between framework-integrated runtimes and sandbox abstraction layers should be treated as an empirical harness design question, not just a product preference.

Was this page helpful?