Building Effective Agents
Anthropic’s “Building Effective Agents” distills a simple lesson from production agent work: the most successful agent systems are usually built from simple, composable patterns rather than complex frameworks.
What Are Agents?
“Agent” is used in several ways. Some teams use it for fully autonomous systems that work independently over long periods with many tools. Others use it for more prescriptive systems that follow predefined workflows.
Anthropic groups these implementations under agentic systems, but draws an important architectural distinction:
- Workflows: LLMs and tools are orchestrated through predefined code paths.
- Agents: LLMs dynamically direct their own process and tool usage, maintaining control over how they complete the task.
The rest of the page follows that distinction: when should control flow live in code, and when should more decisions be left to the model?
When To Use Agents, And When Not To
When building LLM applications, start with the simplest solution that can solve the problem. Many applications do not need an agentic system at all; an optimized LLM call with retrieval and in-context examples is often enough.
Agentic systems usually trade latency and cost for better task performance, so the tradeoff needs to be justified:
- Use workflows for well-defined tasks with consistent structure.
- Use agents when the task needs flexibility, model-driven decisions, and an unpredictable number of steps.
The key principle: do not reach for agents when a workflow is enough.
When And How To Use Frameworks
Frameworks can make agentic systems easier to start with. They simplify model calls, tool definitions, tool parsing, and multi-step chaining.
But frameworks also add abstraction layers. They can hide the underlying prompts, model responses, tool inputs, and intermediate state, making systems harder to debug. They can also tempt teams to add complexity before a simpler setup has failed.
Anthropic recommends understanding the underlying APIs and basic components first. Many of these patterns can be implemented in a few lines of code. If you use a framework, make sure you understand what it does under the hood; incorrect assumptions there are a common source of errors.
Building Block: The Augmented LLM
The basic building block of agentic systems is an LLM augmented with retrieval, tools, and memory. Current models can actively use these capabilities: generating search queries, selecting tools, and deciding what information to retain.
Implementation should focus on two things:
- Whether the capabilities are tailored to the specific use case.
- Whether they provide a clear, stable, well-documented interface for the LLM.
The workflow patterns below assume each LLM call can access these augmented capabilities.
Workflow: Prompt Chaining
Prompt chaining decomposes a task into sequential steps. Each LLM call processes the output of the previous one. Programmatic checks, or gates, can be added between steps to keep the process on track.
Structure: LLM -> gate -> LLM -> gate -> LLM
When to use: The task can be cleanly decomposed into fixed subtasks, and the goal is to trade latency for higher accuracy by making each LLM call easier.
Examples:
- Generate marketing copy, then translate it into another language.
- Write a document outline, check that it meets criteria, then write the document from the outline.
Workflow: Routing
Routing classifies an input and sends it to a specialized follow-up task. This separates concerns and lets each category use specialized prompts, tools, or models. Without routing, optimizing for one kind of input can hurt performance on another.
Structure: Input -> classifier -> specialized handler A | B | C
When to use: The task has distinct categories that are better handled separately, and classification can be performed accurately by an LLM or a traditional classifier.
Examples:
- Route customer service queries such as general questions, refunds, and technical support to different downstream processes, prompts, and tools.
- Route easy or common questions to smaller, cheaper models and hard or unusual questions to more capable models.
Workflow: Parallelization
Parallelization runs multiple LLM calls at the same time and aggregates their outputs. It has two common variants:
- Sectioning: break a task into independent subtasks and run them in parallel.
- Voting: run the same task multiple times to get diverse outputs.
When to use: Subtasks can be parallelized for speed, or multiple attempts and perspectives are needed for higher confidence.
Examples:
- Sectioning: one model instance handles a user query while another screens for inappropriate content or requests.
- Sectioning: automated evals where each LLM call evaluates a different aspect of performance.
- Voting: code review where several prompts inspect a change for vulnerabilities.
- Voting: content moderation where different prompts check different risks or use different vote thresholds.
Workflow: Orchestrator-Workers
In the orchestrator-workers workflow, a central LLM dynamically breaks down a task, delegates subtasks to worker LLMs, and synthesizes their results.
Structure: Input -> orchestrator -> [worker1, worker2, ...workerN] -> orchestrator -> output
When to use: The task is complex and the required subtasks cannot be predicted in advance. This differs from parallelization because the subtasks are not predefined; the orchestrator chooses them based on the specific input.
Examples:
- Coding products that make complex changes across multiple files.
- Search tasks that gather and analyze information from multiple sources.
Workflow: Evaluator-Optimizer
In the evaluator-optimizer workflow, one LLM generates a response while another evaluates it and provides feedback in a loop.
Structure: Generator <-> Evaluator (loop until pass)
When to use: Clear evaluation criteria exist, and iterative refinement provides measurable value. Two signs of fit are that human feedback can demonstrably improve the output, and that an LLM can provide useful feedback of that kind.
Examples:
- Literary translation where an evaluator can critique missed nuance.
- Complex search tasks where the evaluator decides whether more searching is warranted.
Agents
Agents are emerging in production as models improve at understanding complex inputs, reasoning and planning, using tools reliably, and recovering from errors.
An agent usually begins with either a user command or an interactive discussion. Once the task is clear, it plans and operates independently, sometimes returning to the user for more information or judgment.
During execution, agents need ground truth from the environment at each step, such as tool results or code execution. They use this feedback to assess progress and plan next actions. They can also pause for human feedback at checkpoints or blockers. Tasks often end at completion, but stopping conditions such as maximum iterations are important for control.
Agents are often straightforward: an LLM using tools in a loop based on environmental feedback.
while not done:
observe environment
decide next action
execute action via tools
evaluate result
When to use: Open-ended problems where the required number of steps is difficult or impossible to predict, and no fixed path can be hardcoded. Agents may run for many turns, so some trust in their decision-making is required.
Risks:
- Higher cost: open-ended loops consume more tokens.
- Compounding errors: mistakes can propagate into later steps.
- Guardrails required: sandboxing, permission boundaries, and human checkpoints matter.
Examples:
- Coding agents that resolve SWE-bench-style tasks across many files.
- Computer-use agents where Claude uses a computer to complete tasks.
Combining And Customizing Patterns
These building blocks are not prescriptions. They are common patterns that developers can shape and combine for specific use cases.
As with any LLM feature, success depends on measuring performance and iterating. Add complexity only when it demonstrably improves outcomes.
Summary
Success in LLM systems is not about building the most sophisticated system. It is about building the right system for the need.
A useful complexity ladder is:
- Start with simple prompts.
- Optimize prompts with evaluations.
- Add multi-step workflows only when single calls fall short.
- Use autonomous agents only when workflows cannot provide the required flexibility.
Anthropic emphasizes three principles when implementing agents:
- Maintain simplicity in the agent design.
- Prioritize transparency by showing the agent’s planning steps.
- Carefully craft the agent-computer interface through tool documentation and testing.
Frameworks can help you start quickly, but production systems should not be afraid to reduce abstraction layers and return to basic components.
Appendix 1: Agents In Practice
Anthropic highlights two promising agent applications. Both show that agents add the most value for tasks that require both conversation and action, have clear success criteria, enable feedback loops, and include meaningful human oversight.
Customer Support
Customer support combines familiar chat interfaces with tool integration:
- Support interactions naturally follow a conversation flow while requiring external information and actions.
- Tools can pull customer data, order history, and knowledge base articles.
- Actions such as issuing refunds or updating tickets can be handled programmatically.
- Success can be measured through user-defined resolutions.
Some companies have adopted usage-based pricing that charges only for successful resolutions, reflecting confidence in agent effectiveness.
Coding Agents
Software development is a strong fit for agents because:
- Code solutions can be verified through automated tests.
- Agents can iterate using test results as feedback.
- The problem space is well-defined and structured.
- Output quality can be measured objectively.
Automated tests help verify functionality, but human review remains important for ensuring solutions fit broader system requirements.
Appendix 2: Prompt Engineering Your Tools
Tools are often central to any agentic system. Tool definitions and specifications deserve the same prompt engineering attention as the overall prompt.
The same action can often be specified in multiple ways. A file edit can be represented as a diff or by rewriting the whole file. Structured output can be returned in Markdown code blocks or inside JSON. These may be cosmetically equivalent for engineers, but some formats are much harder for an LLM to write correctly.
For example, writing a diff requires knowing line counts before writing the new code. Writing code inside JSON adds escaping overhead for newlines and quotes.
Tool format suggestions:
- Give the model enough tokens to think before it writes itself into a corner.
- Keep formats close to what the model naturally saw in internet text.
- Avoid formatting overhead such as counting thousands of lines or escaping large code strings.
Anthropic recommends investing in agent-computer interfaces (ACI) with the same care historically given to human-computer interfaces (HCI).
Practical techniques:
- Put yourself in the model’s shoes: is the tool obvious from its description and parameters alone?
- Use parameter names and descriptions that make the correct usage clear.
- Test tool use on many example inputs and iterate based on the mistakes the model makes.
- Poka-yoke tools so mistakes are harder to make.
In Anthropic’s SWE-bench agent, tool optimization took more time than overall prompt optimization. One concrete fix was requiring absolute file paths after relative paths caused errors when the agent moved away from the repository root.
Source
Building Effective Agents — Anthropic, December 2024.