Verification and Evaluation (V)

This page corresponds to §8 of Agent Harness Engineering: A Survey. Verification and evaluation turns behavior into engineering evidence. The paper’s core claim is that reported scores are properties of a model-harness pair, not a model alone.

An agent evaluation measures an execution episode: a task grounded in an environment, an agent interacting with tools and state over time, a captured trace, and a judgement over both the final outcome and the path taken.

Five-Stage Task-To-Feedback Lifecycle

1. Task And Benchmark Grounding

For LLM agents, a task is not just a natural-language prompt. It includes environment state, available tools, allowed actions, constraints, termination conditions, and success criteria.

DomainBenchmarksGrounding signal
Software/terminalSWE-bench, Terminal-BenchReal repositories, issue states, tests, patches.
Web/browser/computer-useWebArena, VisualWebArena, BrowserGym, WorkArena, OSWorldBrowser, desktop, or app state changes.
Cross-domain enterprise workflowsAgentBench, GAIA, TheAgentCompany, WorkArena++Breadth across tools, systems, and tasks.

Strong outcome verification requires strong task grounding.

2. Pre-Execution Readiness Validation

Before execution, the harness must validate environment dependencies, tools, context state, permissions, budgets, and judges. Otherwise downstream failures cannot be attributed.

If tool descriptions change across runs, the benchmark no longer tests the same action space. If memory is not reset, the agent may benefit from leaked state. If judge prompts change, reported scores are not comparable.

3. Controlled Execution And Trace Capture

A rollout is the unit of evaluation: task, model config, harness config, action sequence, observations, final state, and score. Controlled rollouts fix environment state, tool availability, timeouts, budgets, permissions, and judge versions.

Traces should record model outputs, tool calls, tool results, state changes, context snapshots, errors, retries, recovery behavior, tokens, latency, and cost. The trace distinguishes failures that look identical at the outcome level.

4. Multi-Level Judgement And Failure Attribution

The paper separates three judgement levels:

  • Outcome-level: did the final state satisfy the task?
  • Trajectory-level: was the path safe, efficient, authorized, recoverable, and consistent?
  • Evaluator-level: was the judge itself reliable?

Failure attribution is a diagnostic process over the trace, not a single-label classification. A failed rollout may stem from model reasoning, tool schema, sandbox state, stale context, flaky tests, benchmark ambiguity, judge instability, or orchestration loops.

5. Continuous Regression And Deployment Feedback

Regression should be triggered by harness changes as well as model changes: tool schemas, compaction policies, sandbox images, permissions, judges, and recovery loops can all cause global regressions.

A practical evaluation suite is layered:

  • unit-style tests for tool schemas;
  • one-step tests for local decisions;
  • full rollout tests;
  • long-horizon multi-session simulations;
  • production traces converted into regression cases.

Evaluation As Measurement Instrument

The paper warns against final-score-centric evaluation. A single pass/fail score can hide infrastructure noise and rollout variance.

Evaluation itself must be studied as a measurement instrument:

  • repeat rollouts to expose variance;
  • version environment, tool, permission, budget, and judge conditions;
  • report success-cost-latency frontiers, not just success rate;
  • compute trajectory metrics over traces;
  • run factorial model-by-harness comparisons.

Trace-native evaluation is the next step: traces become the primary object from which systems compute outcome, trajectory quality, failure attribution, and regression tests.

Was this page helpful?