Blog
Engineering August 24, 2026 7 min read OpenAI

Harness Engineering (6): Agents copy technical debt too

Agents learn from the patterns already in the repository, including duplication, drift, and temporary patches. High-throughput systems need continuous quality garbage collection.

J

Jonathan

Founder

Agents are good at continuing patterns already present in a repository. That is both a source of speed and a long-term risk.

If the repository has clear layers, stable utilities, and consistent tests, the agent tends to preserve them. If it contains three similar helpers, two logging formats, and several temporary workarounds, those patterns also look like valid precedent.

The faster code is generated, the faster patterns spread. A weak implementation that once took months to reach five call sites can now reach dozens in a week.

OpenAI initially devoted every Friday—20% of the week—to cleaning up accumulated “AI slop.” That did not scale. The team instead encoded opinionated “golden principles” into the repository and ran background Codex jobs to find drift, update quality grades, and open targeted refactoring Pull Requests. OpenAI reports that most of these small cleanup changes can be reviewed in under a minute and merged automatically.

The phrase “AI slop” is not formally defined in the original article. Here it refers to the duplication, temporary patches, and low-quality patterns that accumulate when generated code repeatedly follows imperfect precedent.

This is not a periodic cleanup project. It is engineering garbage collection.

High throughput shortens the debt propagation cycle

Technical debt usually spreads through imitation. A task adds a local helper under time pressure. The next task finds that helper nearby and copies it. Eventually the workaround becomes the apparent standard.

Agents accelerate the process because they search existing code for implementation patterns. A frequent pattern looks legitimate whether or not it began as a good decision.

An agent-first repository therefore needs to watch not only whether each change works, but whether undesirable patterns are spreading:

  • duplicate implementations of the same utility
  • drift in directories and dependency direction
  • temporary compatibility logic becoming a dependency
  • repeated brittle testing patterns
  • divergence between documentation and behavior

These signals may not break today’s build. They raise the failure probability of every future agent decision.

Golden principles describe a sustainable repository state

Golden principles are not an endlessly growing style guide. They are explicit preferences with signals that can be checked repeatedly.

Examples include:

  • Prefer shared utilities over local reimplementations.
  • Validate external data at the boundary instead of probing guessed shapes.
  • Express public behavior through typed interfaces.
  • Keep architecture maps aligned with real code.
  • Promote important quality rules from memory into lint or tests.

A useful principle states the preferred condition, why it matters for future maintenance, and how the system detects drift. Pure aesthetic judgment can remain in review guidance. A pattern with predictable long-term cost should move toward an automated check.

Continuous governance fits agents better than large rewrites

Waiting for debt to accumulate and then scheduling a major refactor creates long branches, broad conflicts, and expensive verification. That is especially difficult while agents keep producing changes.

A better loop breaks governance into frequent, bounded tasks:

Scan for one kind of drift
  ↓
Identify a limited scope
  ↓
Open a small refactoring PR
  ↓
Verify behavior and merge
  ↓
Update the rule or check

Each task handles one recognizable pattern: consolidate duplicate helpers, migrate a deprecated API, add missing log fields, or repair stale documentation links. Small scope makes the change easier for an agent to prove and for a human to review.

The loop behaves like garbage collection. It does not promise that waste will never appear. It prevents waste from accumulating until it changes the shape of the system.

Governance agents need permissions and stopping conditions

A background agent should not receive a vague instruction to “make the repository cleaner.” Governance work needs its own harness:

  • handle one defined drift pattern at a time
  • limit writable directories and change size
  • require behavior tests to remain stable
  • propose rather than auto-merge in high-risk areas
  • stop when behavioral equivalence cannot be shown
  • track findings, fixes, false positives, and rollbacks

These constraints prevent governance from creating new drift. Agents handle repetitive, local, evidence-rich cleanup; humans continue to decide architecture and subjective tradeoffs.

Quality metrics should reflect future agent difficulty

Coverage, complexity, and defect counts still matter. Agent-first systems should also ask whether the next agent can understand and modify the repository efficiently.

Useful signals include first-pass success on repeated task types, escalations caused by undiscoverable rules, repeated review comments, the rate of new duplicate implementations, documentation mismatches, and false-positive or rollback rates from automated cleanup.

No single score has to be perfect. The trend matters. If the same task requires more attempts as the repository grows, legibility or constraints may be decaying.

Humans still decide which patterns deserve to become standards, which exceptions should remain, and which local complexity is justified by the product. The change is that once judgment is formed, it should enter documentation, tools, tests, and governance loops so it does not have to be rediscovered in every Pull Request.

OpenAI is also careful about what remains unknown. The team does not yet know how architectural coherence will evolve over years in a fully agent-generated system, where human judgment creates the most leverage, or how the harness should change as models improve. Continuous governance is therefore not proof that the long-term problem is solved; it is the control loop that makes continued learning possible.

The full harness loop is now visible: knowledge makes facts discoverable; the environment makes outcomes observable; constraints make rules enforceable; review makes throughput scalable; continuous governance keeps the system coherent as it changes.

The model determines how smart one run can be. The harness determines whether the system remains reliable and understandable after thousands of runs.

Adapted from OpenAI’s Harness engineering: leveraging Codex in an agent-first world.

Harness Engineering series

harness-engineering coding-agents technical-debt governance