Blog
Engineering August 24, 2026 7 min read OpenAI

Harness Engineering (1): When engineers start designing the agent's environment

When code generation is no longer scarce, engineering shifts from writing every implementation to designing environments, expressing intent, and building feedback loops.

J

Jonathan

Founder

The default question in AI-assisted software development has been: how much code can the model write?

That question is losing explanatory power. Once an agent can search a repository, edit files, run tests, respond to review, and open a Pull Request, code volume is no longer the only constraint. The harder questions are whether the agent can understand the goal, observe the environment, verify the result, and recover when it drifts.

OpenAI pushed this shift to an unusual extreme. An internal team started with an empty repository and, over five months, grew it to roughly one million lines of code. A three-engineer team drove about 1,500 merged Pull Requests. Application logic, tests, CI, documentation, observability, and internal tools were all generated by Codex; humans did not contribute code directly. The team estimates that the product took about one-tenth of the time required to write the code by hand.

This was not output for its own sake. The product shipped, broke, and was repaired in use, with hundreds of internal users and external alpha testers. The initial repository structure, CI, formatting, package setup, application framework, and even the first AGENTS.md were generated by Codex CLI with GPT‑5, guided only by a small set of templates. There was no existing human-written code to anchor the system.

The important part is not the million lines. It is that the team changed the object of engineering.

Humans steer. Agents execute. Engineers build not only the product, but also the environment in which agents can work reliably.

That is harness engineering.

Human attention becomes the bottleneck when code gets cheaper

Traditional software throughput is constrained by implementation capacity. Even after a requirement is clear, someone still has to design, code, test, review, and fix it. Adding engineers usually adds production capacity.

Coding agents change that constraint. As implementation accelerates, teams encounter a different set of failures:

  • The agent cannot discover why an architectural decision exists.
  • It can change the UI but cannot independently confirm the interaction works.
  • It can read an error log but cannot connect it to the relevant request.
  • It completes the local requirement while violating a repository boundary.
  • It creates more changes than human QA and review can absorb.

“Try again” does not solve these failures reliably. The next run may work, but the team has not created reusable capability.

OpenAI treated each blockage as a missing environmental capability. When the agent lacked a tool, the team added a tool. When it lacked project knowledge, the knowledge moved into the repository. When it could not validate UI behavior, the application and browser became directly inspectable. When a rule was repeatedly violated, the rule became lint or a structural test.

The engineer stopped fixing only the output and started improving the system that produced it.

Harness engineering is not prompt optimization

A prompt can express the intent of one task. A harness determines what the agent can see and do throughout the run, and how it can know the task is complete.

For a coding agent, that environment has at least five layers:

Knowledge: goals, architecture, product rules, and decisions are discoverable
Environment: code, browser state, logs, metrics, and runtime behavior are observable
Tools: search, editing, testing, deployment, and collaboration are executable
Constraints: permissions, dependency boundaries, and quality rules are enforceable
Feedback: failures, reviews, and production bugs become future capability

Without those layers, teams keep writing longer prompts and copying the same context, warnings, and checklists into every run. One task may succeed, but the next one starts from zero.

A good harness turns temporary reminders into stable system behavior. The agent does not need another instruction not to cross an architecture boundary if a structural test blocks it. It does not need a human to relay a Slack decision if the decision and rationale are versioned. It does not need someone to open the page if it can launch and inspect an isolated instance itself.

The goal is not to tell the agent more. It is to make correct action easier and errors visible earlier.

The engineer moves from implementation to system design

This does not make engineers less important. It moves judgment to a higher-leverage layer.

Humans still decide which problem matters, what counts as done, which boundaries must hold, which failures can be repaired automatically, and which risks require judgment. They also decide whether a review comment is local advice or a principle worth encoding into the system.

The difference is that valuable judgment should not remain trapped in one review. It should become documentation, tooling, tests, evals, or constraints that every later agent run can use.

Classify the failure before adding more scaffolding

A team does not need to copy OpenAI’s no-handwritten-code experiment. A more practical starting point is to take one recurring agent failure and identify the missing layer.

FailureDo not onlyAdd to the harness
The agent repeatedly misreads the repositoryRe-explain the foldersA repository map and architecture index
A human must open the page after every fixAdd “please verify”A runnable instance and browser tools
The same review comment keeps returningLeave the same comment againLint, types, or structural tests
Long runs forget early decisionsKeep the entire transcriptVersioned plans and decision logs
PR volume exceeds review capacityAsk people to review fasterRisk tiers, automated checks, and agent review

Model failure does not always mean the model is too weak. Often the environment has failed to expose the information, action, or feedback required for success.

Maturity is measured by the closed loop

A simple harness maturity ladder looks like this:

  1. Generate: the agent writes code; a human supplies context, runs tests, and validates.
  2. Execute: the agent can search, edit, run commands, and finish bounded tasks.
  3. Verify: the agent can launch the application, reproduce issues, inspect runtime signals, and prove the change works.
  4. Improve: failures become documentation, tools, evals, or constraints that prevent repetition.

The dividing line is not how many AI tools the team uses. It is whether the agent can close the loop from understanding to action to evidence, and whether human judgment accumulates inside that loop.

Adapted from OpenAI’s Harness engineering: leveraging Codex in an agent-first world.

Harness Engineering series

harness-engineering coding-agents agent-first codex