Long-Running Handoffs
As AI agents become more capable, developers increasingly ask them to take on complex tasks that span hours or days. Getting agents to make consistent progress across multiple context windows remains an open problem.
The core challenge is that long-running agents work in discrete sessions, and each new session begins with no memory of what came before. This resembles a software project staffed by engineers in shifts, where each new engineer arrives without knowing what happened on the previous shift. Context windows are limited, and most complex projects cannot be completed within one window, so agents need a way to bridge coding sessions.
Anthropic’s solution is a two-part harness: an initializer agent sets up the environment on the first run, and a coding agent makes incremental progress in every later session while leaving clear artifacts for the next one.
The Long-Running Agent Problem
The Claude Agent SDK is a general-purpose agent harness for coding and other tasks that require the model to use tools, gather context, plan, and execute. It has context management features such as compaction, which theoretically allow an agent to keep working without exhausting the context window.
But compaction is not sufficient. Even a frontier coding model running in a loop across multiple context windows can fail to build a production-quality web app if it only receives a high-level prompt such as “build a clone of claude.ai.”
Anthropic observed two main failure modes:
| Failure mode | Behavior |
|---|---|
| Over-ambition | The agent tries to one-shot the whole app, runs out of context mid-feature, and leaves half-finished, undocumented work. |
| Premature completion | Later in the project, a fresh agent sees that progress has been made and declares the whole job finished. |
This decomposes the problem into two parts:
- Set up an initial environment that lays the foundation for all features required by the user prompt, so the agent can work step by step and feature by feature.
- Prompt each later agent to make incremental progress while leaving the environment in a clean state at the end of the session.
A “clean state” means code that would be appropriate to merge to a main branch: no major bugs, orderly and documented code, and no unrelated mess for the next developer to clean up.
Two-Part Solution
Anthropic used two roles internally. They are called separate agents here because they use different initial user prompts; the system prompt, tools, and overall harness are otherwise the same.
Initializer Agent
The first session uses a specialized initializer prompt that asks the model to set up the environment future coding agents need. Key artifacts include:
init.sh: a script that starts the development environment, such as installing dependencies and running the dev server.- Feature list: a comprehensive list of end-to-end features expanded from the initial request.
claude-progress.txt: a progress file recording what agents have done.- Initial git commit: a baseline commit showing what files were added during initialization.
The feature list is central. To prevent one-shotting or premature completion, the initializer expands the user’s high-level prompt into a structured feature file. In the claude.ai clone example, this meant more than 200 features, such as a user opening a new chat, typing a query, pressing enter, and seeing an AI response.
All features start as failing, giving later coding agents a clear picture of what complete functionality looks like.
{
"category": "functional",
"description": "New chat button creates a fresh conversation",
"steps": [
"Navigate to main interface",
"Click the 'New Chat' button",
"Verify a new conversation is created",
"Check that chat area shows welcome state",
"Verify conversation appears in sidebar"
],
"passes": false
}
Coding agents may only update the passes field. They should not remove or rewrite feature descriptions or test steps. Anthropic used strong instructions that removing or editing tests is unacceptable because it can lead to missing or buggy functionality.
They eventually chose JSON for this file because the model was less likely to inappropriately change or overwrite JSON files than Markdown files.
Coding Agent
Every later session uses a coding agent. Its job is not to finish everything, but to make incremental progress and leave structured updates.
Incremental progress is critical: the agent works on only one feature at a time, verifies it, and then moves on.
The model must also leave the environment clean after a change. Anthropic found that the best way to elicit this was to ask the model to:
- Commit progress to git with descriptive messages.
- Summarize progress in a progress file.
- Use git to revert bad changes and recover working states when needed.
This prevents future agents from guessing what happened or spending time getting the basic app working again.
Environment Management
The harness does not try to make the model remember more. It externalizes cross-session state into the environment so each new session can quickly orient itself.
Feature List
The feature list turns high-level requirements into verifiable end-to-end behavior. It defines both what remains and what “done” means.
It solves two problems:
- It prevents agents from trying to build the whole app at once.
- It prevents later agents from mistaking partial progress for completion.
It also fixes the success criteria. Later agents can mark a feature as passing, but not change the test itself.
Incremental Progress
Given the initial environment, each coding agent selects one high-priority unfinished feature. After implementation, it tests, updates the progress file, and commits to git.
This may make each session slower, but it helps the whole system converge because each session starts from a known-good state and leaves the next one with a new known-good state.
Testing
Anthropic observed another major failure mode: Claude often marked a feature complete without proper testing. Without explicit prompting, it might change code and run unit tests or curl commands against a dev server, while still failing to notice that the feature did not work end to end.
For web apps, Claude did much better when explicitly prompted to use browser automation tools and test as a human user would.
These tools dramatically improved performance because the agent could find bugs that were not obvious from code alone.
Some issues remained. Claude’s vision and browser automation tools have limits. For example, Claude cannot see browser-native alert modals through Puppeteer MCP, so features depending on those modals tended to be buggier.
Getting Up To Speed
With the environment in place, every coding agent is prompted to run a few orientation steps:
- Run
pwdto see the working directory. The agent can edit only files in that directory. - Read git logs and progress files to understand recent work.
- Read the feature list and choose the highest-priority unfinished feature.
The initializer should also write an init.sh script so later sessions do not need to rediscover how to run the app. Before implementing a new feature, the coding agent should run init.sh, start the development server, and run a basic end-to-end test.
In the claude.ai clone, this meant starting the local server and using Puppeteer MCP to start a new chat, send a message, and receive a response. If the app had been left broken, the agent could identify and fix that first rather than making the problem worse by starting a new feature.
A typical session starts like this:
[Assistant] I'll start by getting my bearings and understanding the current state of the project.
[Tool Use] bash - pwd
[Tool Use] read - claude-progress.txt
[Tool Use] read - feature_list.json
[Assistant] Let me check the git log to see recent work.
[Tool Use] bash - git log --oneline -20
[Assistant] Now let me check if there's an init.sh script to restart the servers.
<Starts the development server>
[Assistant] Now let me navigate to the application and verify that fundamental features still work.
<Tests basic functionality>
[Assistant] Based on verification testing, the core app works. Now I can review the feature list and start the next feature.
<Starts work on a new feature>
Failure Modes And Solutions
| Problem | Initializer agent behavior | Coding agent behavior |
|---|---|---|
| Claude declares victory on the entire project too early. | Create a structured JSON feature list from the input spec, with end-to-end feature descriptions. | Read the feature list at the beginning of the session and choose a single incomplete feature. |
| Claude leaves the environment buggy or with undocumented progress. | Create an initial git repo and progress notes file. | Read progress notes and git logs, run a basic dev-server test, and end with a git commit plus progress update. |
| Claude marks features done prematurely. | Create the feature list. | Self-verify all features and mark passing only after careful testing. |
| Claude wastes time figuring out how to run the app. | Write an init.sh script that starts the dev server. | Start by reading and running init.sh. |
Future Work
This work demonstrates one possible long-running agent harness that lets a model make incremental progress across many context windows. Open questions remain.
One question is whether a single general-purpose coding agent performs best across contexts, or whether a multi-agent architecture would do better. Specialized agents such as testing agents, QA agents, or code cleanup agents may improve subtasks across the software development lifecycle.
The demo was also optimized for full-stack web app development. Future work should generalize these lessons to other domains, such as scientific research or financial modeling, where long-running agentic tasks may need similar patterns.
Key Takeaway
The core lesson is not to make the context window infinite. It is to decompose long-running work into recoverable, verifiable, handoff-friendly engineering actions:
- Use an initializer agent to create the environment and success criteria.
- Use a coding agent to make one feature of progress at a time.
- Use a feature list to fix the definition of done.
- Use
claude-progress.txtand git history as the cross-session bridge. - Use browser automation and user-perspective testing to reduce false completion.
The most effective harnesses often copy practices that make human engineering teams effective: clear handoffs, incremental progress, verified completion, and clean state.
Source
Effective Harnesses for Long-Running Agents — Justin Young, Anthropic, November 2025.