Blog
Engineering August 24, 2026 6 min read OpenAI

Harness Engineering (5): When agent output exceeds human attention

Once agents increase code throughput, one-by-one human review becomes the next bottleneck. Checks, review, merge policy, and recovery must be configured by risk.

J

Jonathan

Founder

When one agent submits one small change a day, the existing Pull Request process can cope. A human reads the diff, runs tests, leaves comments, and waits for an update.

When several agents search, implement, test, and respond to feedback in parallel, the constraint reverses. Generation finishes, but Pull Requests wait for human attention. The queue grows, branches conflict, and correct work stalls because no one has looked at it.

In OpenAI’s experiment, the initial three-engineer team processed about 3.5 PRs per engineer per day. Over time, most review shifted toward agent-to-agent workflows, with Codex reading feedback, updating the change, repairing build failures, and continuing the PR lifecycle.

The lesson is not “remove humans from review.” It is this:

When execution throughput changes by an order of magnitude, quality control must move from uniform human handling to risk tiers and system verification.

Human review should not carry every quality responsibility

“Someone looked at it” is not a stable safety guarantee. Large diffs, context switching, time pressure, and uneven domain knowledge all affect review quality.

A scalable system separates responsibilities:

  • types, formatting, and dependency boundaries go to static checks
  • known behavior goes to unit and integration tests
  • user journeys go to browser and end-to-end verification
  • performance and reliability go to logs, metrics, and traces
  • recurring defect patterns go to specialized review agents
  • product tradeoffs, security boundaries, and irreversible risk stay with people

Human attention moves toward what the system cannot yet express mechanically.

Agent review needs distinct perspectives

Asking the implementing agent to “review your work” often finds only surface defects. The agent carries the same assumptions that shaped the solution.

OpenAI explicitly describes self-review plus additional agent reviews, but it does not prescribe the role taxonomy below. A practical extension is to give review agents different objectives and context:

  • an architecture reviewer checks dependency direction and public interfaces
  • a test reviewer searches for uncovered failures and assertions fitted too closely to the implementation
  • a security reviewer checks permissions, exposure, and untrusted input
  • a product reviewer checks behavior against acceptance criteria
  • a simplification reviewer looks for unnecessary abstraction and duplicated work

They do not need different models. They need different roles, tools, and evidence requirements.

Merge speed must be designed with repair capacity

OpenAI made a claim that is easy to misuse: in a high-throughput environment, the cost of correction can be lower than the cost of waiting, so every flaky test need not block indefinitely.

That logic holds only when changes are small, regressions are detected quickly, flaky failures are distinguishable from real failures, rollback is cheap, and every failure has an owner.

Payments, authorization, destructive actions, and infrastructure migrations need a different policy. They require stronger approval, staged rollout, and explicit recovery.

RiskTypical changeMerge policy
LowDocumentation, internal tools, small stylingFast merge after automated checks and agent review
MediumReversible product or API behaviorFull tests, runtime verification, selective human review
HighPermissions, payments, migrations, deletionHuman approval, staged release, explicit rollback

Speed does not mean lowering quality. It means spending quality-control effort in proportion to actual risk. The three-tier policy above is our operating framework, not a policy reported by OpenAI.

A Pull Request should carry verification evidence

A PR containing only a diff transfers most understanding work to the reviewer. A better result states what changed, how the issue was reproduced, which checks ran, what UI or telemetry evidence supports the fix, what risk remains, and whether human judgment is required.

The evidence need not be a long report. Commands, key output, before-and-after screenshots, and explicit residual risk can substantially reduce review cost.

The agent submits not only code, but a verifiable completion claim. The reviewer no longer has to reconstruct the entire task; they judge whether the evidence supports merge.

Autonomy should rise in measured stages

Teams can expand authority gradually:

  1. The agent opens a change; a human reviews and merges.
  2. The agent self-checks and answers feedback; a human decides.
  3. Review agents handle routine feedback; people resolve disagreement and high-risk cases.
  4. Low-risk changes auto-merge after required checks; high-risk changes remain gated.
  5. Agents detect build failures, repair or roll back, and escalate only when judgment is needed.

Each stage should be supported by regression rate, time to repair, human intervention rate, flaky-test rate, and rollback success. If the metrics worsen, the missing piece may be the harness rather than the model.

OpenAI adds an essential caveat: this level of autonomy depends heavily on the structure and tooling of its specific repository. It should not be assumed to generalize without comparable investment. The right autonomy level is earned by the environment, not granted because a model can complete an impressive demo.

The goal is not maximum throughput. It is to place human attention where it creates the most value.

Adapted from OpenAI’s Harness engineering: leveraging Codex in an agent-first world.

Harness Engineering series

harness-engineering coding-agents code-review agent-workflow