Overview
This section follows Agent Harness Engineering: A Survey (Li et al., TMLR under review). The survey treats the agent execution harness as an independent system layer: the engineered wrapper around a language model that turns model calls into bounded, stateful, tool-mediated task execution. A harness is what makes long-running agent behavior controllable, inspectable, recoverable, measurable, and governable.
The paper maps 170+ public technical artifacts onto a seven-layer taxonomy, ETCLOVG. The pages in this section follow that taxonomy: first the paper’s thesis and methodology, then the seven layers.
Binding Constraint Thesis
The survey starts from the binding-constraint thesis: for long-horizon tasks evaluated across comparable frontier models, benchmark variance may be driven as much by the execution harness as by the model itself.
The paper cites three harness-only results:
- Changing only the edit-tool format and surrounding tool harness produced up to 10x gains on coding benchmarks across 15 models.
- A fixed GPT-5.2-Codex agent improved from 52.8% to 66.5% on Terminal-Bench 2.0 through system prompt restructuring, middleware context injection, and self-verification hooks.
- Meta-Harness reached 76.4% on Terminal-Bench-2 through automated harness optimization, surpassing hand-engineered approaches without changing model weights.
Each gain is larger than the 2-4 percentage point changes often treated as meaningful model progress on the same class of benchmark. The point is not that models no longer matter. The point is that, once models are strong enough to attempt long tasks, reliability is often bounded by the controller around the model.
Practitioner-Research Gap
Practitioners already use the vocabulary of harness engineering. OpenAI frames it as designing environments, constraints, documentation, and feedback loops around coding agents. Anthropic’s engineering posts arrive at similar principles from adjacent directions: keep architectures simple and inspectable, design tools for agents rather than copying human APIs, progressively disclose context instead of eagerly loading it, and preserve durable handoff artifacts for long-running work.
Research has studied the pieces: memory, tool use, planning, multi-agent coordination, security, and evaluation. What has been missing is a vocabulary for the system that composes those pieces into reliable operation. ETCLOVG is the paper’s attempt to name that system surface.
Scope And Corpus
The paper does not use “harness” to mean “all software around an LLM.” A project is in scope only when it implements or specifies a concrete mechanism for bounded, stateful, tool-mediated agent execution.
Included:
- agent frameworks with reusable orchestration or tool-routing logic;
- benchmarks that instantiate executable agent environments;
- sandboxes packaged for agent execution;
- memory, observability, evaluation, or governance systems that operate over agent state, traces, actions, or policies.
Excluded:
- simple chatbot demos;
- prompt packs;
- thin model-client wrappers;
- static datasets or leaderboards without an agent runtime;
- generic infrastructure not adapted to agent execution;
- product pages whose technical behavior cannot be inspected from public documentation.
Borderline cases are resolved by mechanism, not by label. A repository called an “agent” is not sufficient; an evaluation or sandbox project is included when it supplies reusable harness machinery.
The corpus is built from four streams: prior surveys and benchmark papers, reproducible GitHub searches, curated lists and package registries, and company engineering blogs or release notes. Each retained artifact is coded against the seven layers using public evidence. Coding is multi-label: the primary layer marks the artifact’s dominant mechanism, while secondary layers are assigned only when the documentation exposes an independent capability.
The paper is explicit about limitations. The corpus maps the visible agent-harness ecosystem, not every production system. It is biased toward English-language sources, GitHub-visible projects, open-source artifacts, and coding-agent infrastructure. Absence from a layer means “not publicly evidenced,” not “not implemented.”
Three Engineering Phases
The paper describes a shift in the marginal engineering surface:
| Phase | Main lever | What changed |
|---|---|---|
| Prompt engineering (2022-2024) | The input prompt | Optimize a mostly static text input to one model call. |
| Context engineering (2025) | The visible token set | Decide what the model should see at each step of a longer run. |
| Harness engineering (2026) | The surrounding control system | Maintain state, mediate tools, inject feedback, enforce constraints, evaluate progress, and recover execution. |
These phases overlap. Harness engineering includes context engineering; context engineering includes prompt design. The timeline is not a clean replacement story. It marks where the binding constraint moved.
The ETCLOVG Taxonomy
ETCLOVG divides the harness into a structural core and a control plane:
| Layer | Name | Responsibility |
|---|---|---|
| E | Execution environment and sandbox | Where agent code runs and what isolation bounds it. |
| T | Tool interface and protocol | How external capabilities are described, discovered, and invoked. |
| C | Context and memory | What the model can see across short-term, session-level, and persistent horizons. |
| L | Lifecycle and orchestration | The control flow that reads and writes state: inner loops, multi-agent patterns, and task pipelines. |
| O | Observability and operations | Traces, costs, failures, runtime signals, and reliability monitoring. |
| V | Verification and evaluation | Turning tasks and traces into scores, attribution, regression feedback, and deployment evidence. |
| G | Governance and security | Permissions, identity, policy, hardening, audit, and human oversight. |
The first four layers form the structural core. O, V, and G form the control plane around it.
Two taxonomy choices matter:
- Observability is promoted to a first-class layer rather than treated as a side effect of lifecycle hooks.
- Governance is promoted to a first-class layer because permissions, identity, policy, hardening, audit, and oversight are owned by distinct production systems and teams.
State management stays inside Lifecycle and Orchestration because state belongs next to the control flow that reads and writes it.
Cross-Layer Synthesis
The paper’s main contribution is not just the seven labels. It is the claim that a harness behaves as a coupled control system.
- Cost-quality-speed trilemma: stronger sandboxes, richer context, deeper evaluation, stricter governance, and more complete telemetry usually increase latency and cost. Production systems must decide which checks run synchronously, which run offline, and which failures justify expensive recovery paths.
- Capability-control tradeoff: every increase in authority expands the control problem. Larger tool menus, persistent memory, and permissive sandboxes all increase capability while enlarging the injection surface, privacy risk, stale state risk, and blast radius.
- Harness coupling problem: local optimizations are fragile because layers interact. Execution environments affect evaluation, tool descriptions consume context budget, observability traces become governance evidence only when they include identity and permission state, and verifiers reward some orchestration loops while penalizing others.
The ecosystem is moving from agent frameworks to agent platforms. Frameworks package local abstractions such as tools, memory stores, agents, and loops. Platforms add durable workspaces, managed sandboxes, identity, billing, observability, evaluation, governance, and human handoff across many users and many runs.
Open Problems
The survey closes with five cross-layer problems:
- Hardening and scaling execution environments: common security evaluations, cost models for choosing containers vs. microVMs vs. OS permissions vs. browser/desktop environments, and portability across deployment modes.
- Maintaining reliable state in long-running agents: treating context management as state estimation, with uncertainty-aware summaries, provenance, contradiction handling, staleness markers, and recovery from durable artifacts.
- Diagnosing failures from traces: moving from final-score-centric evaluation to trace-native evaluation that computes outcome, trajectory quality, attribution, and regression cases from spans.
- Standard handoffs across agents, tools, and humans: transferring intent, constraints, permissions, artifacts, provenance, budget state, risk level, trace history, and unresolved decisions, not just a text summary.
- Keeping harnesses useful as models improve: continuously re-evaluating wrappers, resets, verifiers, planners, memory rules, and permission gates as model capabilities change.
The paper is descriptive today. Turning ETCLOVG into a normative design framework is the next step.