Agent Team Coordination Model
Start with a capable model and validate the whole system
LLMs decide whether to delegate, how to divide work, whom to contact, and whether results satisfy the goal. The harness provides executable tools, relevant context, reliable messages, cancellation, and explicit outcomes. Model autonomy is the starting assumption; the acceptance criterion is that LLM plus harness can complete the task together.
Compatibility does not mean asking weaker models to infer hidden runtime rules. All supported models need the same clear contract for delivery, waiting, and recovery. Add model-specific reminders or restrictions only after observing a recurring failure, and measure whether they improve completion without unnecessarily limiting stronger models.
Use the smallest effective change: first check missing context, contradictory instructions, and unreliable tools; add new workflow machinery only when those repairs do not address the observed problem.
The Lead owns the goal and the delegation strategy
Use a team when delegation improves quality, elapsed time, or context management enough to justify briefing and integration. The number of specialists or final documents does not determine whether delegation is useful. Members execute their assignments autonomously; the Lead checks results, reconciles disagreements, and owns the final delivery.
team_create declares the work that is clear now. Each member has a unique name, an agent type, a self-contained brief, and optional dependsOn references. Independent members launch after the Lead’s turn ends. Dependent members launch only when all prerequisites are delivered. team_replan can add or revise work as new information emerges; the Lead need not predict the entire workflow before starting.
Five tools provide the coordination contract
| Tool | Caller | Behavior |
|---|---|---|
team_create | Lead | Create members and their current assignments; validate names and cycles. |
team_status | Both | Read current state, at most once per run to prevent polling. Concurrent work may change it. |
team_send_message | Both | Persist a message; active recipients read it at their next step, idle recipients are woken after the sender’s turn. |
team_replan | Lead | Add work, revise delivered work, or cancel a branch. Changes apply sequentially, not transactionally. |
team_dissolve | Lead | Stop remaining runs and close the team. |
A member owns one logical task but can have multiple runs. Revision keeps the member and task identity. A new unrelated assignment uses a new member. Names cannot be reused within the team’s history.
Waiting does not complete the assignment
A member that cannot proceed without an answer calls team_send_message with wait_for_reply: true. The run stops at the tool-step boundary and leaves the task unfinished. A subsequent incoming message can wake the member; post-run reconciliation also catches replies arriving during cleanup. Ordinary messages do not stop the sender. The Lead sends normal messages and ends its turn when it needs to wait.
Before waiting, save progress and include relevant paths in the question. Wake runs rebuild context rather than restoring the full previous conversation. Both Lead and members consume message bodies into reminders retained for the receiving run. Downstream members receive the summaries and paths of their task dependencies.
A clean wait is persisted as finishReason: "waiting" in the existing member run record. Recovery checks that record and the inbox before starting a new run; without new mail, a committed wait stays asleep. No task-state migration is required. A crash before run finalization remains an interrupted-run recovery case, and an incoming message is not guaranteed to answer the question.
Delivery and recovery have explicit effects
Members deliver through complete({ summary, paths }). Short findings can use paths: []; substantive files must exist. Any invalid supplied path blocks team task completion, so downstream work does not start on a missing deliverable. complete must not be used to pause unfinished work.
Failed tasks keep downstream work blocked. Cancelling a failed member clears unfinished downstream branches while preserving the root failure evidence. Replacement members need new names and explicit dependencies; old edges are not redirected automatically. New work can depend on a delivered member. Inspect status after a partial replan failure before retrying.
Revising a delivered task reopens it and marks delivered downstream work stale. Revision requires downstream work to be delivered or cancelled first. All tasks being delivered does not close the team: the Lead accepts the results and dissolves it when no further work remains.
Compatibility starts with a consistent contract
The model chooses assignments, communication, and corrections. The harness supplies reliable execution and explicit tool outcomes. Additional model-specific reminders or restrictions should follow observed failures, rather than fixing every model to one workflow. Contract tests validate transitions and tool behavior; real-model evaluations are still needed to measure delegation quality and completion cost.
Use task for an isolated delegation whose result returns through the tool call. Use team when messaging, coordination, or follow-up is useful. See Agent Team implementation and Subagents.
Evaluate real models with the same work and observable outcomes
The cases below are a test plan, not measured results. Run the same inputs, available tools, workspace snapshot, and budgets for each selected model. Record the exact model identifier and configuration; compare a single-agent baseline where the scenario permits it. Use multiple runs before attributing a failure to model capability.
| Case | Input or intervention | Evidence to inspect |
|---|---|---|
| Direct answer | A small question fully answered by supplied text. | Correct answer without unnecessary delegation. |
| Independent review | Two independent files with distinct review scopes. | Coverage, duplicate work, elapsed time, and integration quality. |
| Dependency handoff | Research produces facts needed by a writer. | Writer starts after delivery and reads the actual upstream paths. |
| Missing decision | Withhold a required region; supply it after the question. | No fabricated assumption; wait preserves the task; reply leads to completion. |
| Revision | Supply corrected evidence after delivery. | Same relevant task is revised and affected downstream results are updated. |
| Failure recovery | Remove a required input or return a tool error. | Failure is surfaced, useful completed work is retained, and replacement work has valid dependencies. |
| Restart while waiting | Restart after the waiting run has finalized, then send a reply. | No run before new mail; the reply triggers useful continuation. |
Measure outcome quality separately from runtime correctness. Record task success, missing or fabricated claims, invalid tool calls, redundant members, duplicate work, waiting loops, elapsed time, and token/cost totals. A cheaper run that produces an incomplete answer is not a successful optimization. Add harness restrictions only for recurring failures whose cause and improvement can be demonstrated.
Teams channels explicitly enable coordination tools
The /teams/:id page sends messages through the chat transport into runChatTurn. The general assistant loadout excludes team, so buildChatToolKeys adds it explicitly for the system Lead in a channel when the host supplies team runtime wiring. The same tool list initializes context.team and reaches the agent loop, enabling team_create, the other coordination tools, and the Lead instructions. Ordinary chats and directly mentioned experts do not receive this orchestration capability. Members use their separate execution path.
Creating a channel makes coordination available; the model still decides whether a request benefits from delegation. A report that team_create is unavailable requires checking runtime tool assembly before adjusting prompts. Restart or reload the backend after changing this code, then send a new message; an already running turn retains its original tools.
One expert definition supports two execution modes
Both standalone subagents and team members use buildDelegatedAgentSetup to assemble model configuration, expert instructions, and the configured step budget. The shared base instructions receive a standalone or member mode; the expert’s professional instructions remain the same. Members no longer receive an automatic threefold step allowance.
The member mode adds team communication and its own successful-delivery and waiting stop conditions. Its team tool factory exposes only team_status and team_send_message; Lead orchestration tools remain unavailable to members, with execution-time role checks retained. Both modes deliver a summary and optional files through complete. Stream persistence, mailbox handling, dependency scheduling, and wake recovery remain in the team lifecycle service rather than the shared configuration builder.