Multi-Agent Collaboration
Two collaboration levels use different protocols
aibuddy does not force every delegation into one multi-agent abstraction. Bounded work that returns one result uses a task subagent. Work with dependencies, member communication, or replanning uses an Agent Team. Both share the real workspace, but they have different lifetimes and coordination protocols.
| Dimension | task subagent | Agent Team |
|---|---|---|
| Lifetime | One run, ending after delivery | Repeated event-driven runs until delivery or dissolution |
| Input | A self-contained delegation brief | Initial task, dependencies, and later inbox messages |
| Coordination | No mid-run question to the lead | Durable messages between the lead and members or peers |
| Scheduling | Independent calls can run concurrently | A member starts only when its dependencies are satisfied |
| Return | Summary, artifact paths, status, and step count | Versioned task deliveries and durable team state |
This distinction keeps simple delegation lightweight without forcing long-running collaboration into a one-shot call.
Subagents isolate execution traces while sharing the real workspace
task creates a fresh runtime context, compaction state, and tool loadout for the child. The subagent receives only the delegation brief, not the lead’s full conversation, and cannot ask the lead for clarification mid-run. The brief must therefore contain the goal, input boundary, allowed mutation scope, expected artifact, and acceptance criteria needed for independent work.
Runtime isolation is not filesystem isolation. The subagent shares the current sandbox and workspace one-to-one with the lead. Its files need no copying or path translation before inspection. Concurrent subagents must receive separate file or directory ownership because the system does not merge competing writes to one target.
Live child steps are forwarded as UI events and persisted as a reopenable execution trace, but they do not enter the lead model’s context. Heartbeats keep the parent tool call active. At completion, only the complete summary, validated artifact paths, skipped items, status, and step count return to the lead. The user retains an inspectable trace while the lead context carries only delivery information.
Agent Team persists coordination as a task graph
An Agent Team declares members and dependencies in one team_create operation. The runtime validates unique member names, dependency targets, and an acyclic graph before it writes the team, members, and tasks. team_create accepts up to eight worker members per call; this is not a lifetime limit enforced on additions, and each member owns one logical task. Models collaborate through stable names rather than internal UUIDs or a teamId parameter.
The repository is authoritative for team state. The server uses PostgreSQL and the desktop host uses SQLite. Graph decisions read fresh member and task records instead of maintaining a second long-lived in-memory graph that can drift. A conditional transition atomically moves a task from claimable to in_progress, preventing concurrent runs from claiming the same work.
Member state and task state are recorded separately. A member can be spawning, working, idle, done, failed, or cancelled; a task can be pending, blocked, stale, claimable, in_progress, completed, failed, or cancelled. An idle process is therefore not confused with a completed task.
Events wake only members that can make progress
A member with unmet dependencies remains idle; the system does not start a model process merely to wait. After a completion event, the router reads a fresh snapshot. It wakes only a member whose task is claimable and whose dependencies independently revalidate as completed. This second check prevents a stale precomputed status from starting downstream work.
Routing follows the graph instead of waking the lead for every update:
| Event | Wake target |
|---|---|
| Direct or broadcast message | The recipients; broadcast excludes the sender |
| Task completed | Downstream members that are now genuinely runnable |
| Task failed, cancelled, or team dissolved | The lead |
| No downstream can run and nobody is executing | The lead, to distinguish completion from a stall |
| Task creation, claim, or ordinary status change | State update only; no model run |
team_status returns immediately and is limited to one call per run to prevent polling. Concurrent work can still change the state. A member that cannot proceed without an answer uses team_send_message with wait_for_reply: true; the run ends while its logical task stays unfinished. A cleanly finalized wait is recorded as finishReason: "waiting"; recovery leaves it asleep without new inbox messages. A crash before finalization still follows interrupted-run recovery.
Before every member step, the runtime injects current identity, unread messages, and claimable work. The identity reminder overwrites a stable key, so role boundaries can be reinforced on every step without accumulating duplicate context.
Messages, deliveries, and files have separate responsibilities
team_send_message carries decisions, blockers, and small amounts of context rather than large artifacts. A message is persisted to the recipient inbox first. Its wake event is deferred until the sender’s current turn commits, so an idle recipient starts after that boundary. An already running recipient can read the message on its next step before the sender finishes; this is not a transaction over the whole turn.
File output belongs in the shared workspace; concise findings may be delivered as a summary with paths: []. The logical task submits a delivery through complete, and any invalid supplied path blocks team task completion. Both Lead and members receive inbox bodies, and downstream members receive dependency summaries and paths. A completion event can hand execution directly to downstream members without involving the lead at every intermediate node. The lead still owns final integration and acceptance; all member runs ending does not by itself prove the combined result.
| Channel | Content | Retention |
|---|---|---|
| Inbox | Short messages, questions, and course corrections | Persisted and consumed on a later step |
| Task delivery | Summary, result, and revision | Append-only record tied to the logical task |
| Shared workspace | Files, code, and inspectable artifacts | Directly readable and writable by every member |
Replanning preserves history instead of overwriting results
team_replan can cancel members, add new members, or revise a completed logical task. A revision keeps the member and task identity, increments a monotonic version, and records the new instruction. Completed downstream work based on the prior version becomes stale and is revised in dependency order. Claims also snapshot the exact dependency revisions used, making the basis of a delivery traceable.
Replan applies sequentially rather than transactionally. Added work may depend on delivered members; failed roots retain failure evidence when unfinished downstream branches are cancelled. Revision requires downstream tasks to be delivered or cancelled first.
Cancellation uses the opposite ordering for safety. The target and its downstream tasks are persisted as cancelled before the member stream is aborted. Stream cleanup therefore cannot overwrite an explicit cancellation with a generic failure. If a member exits unexpectedly or loses its heartbeat, open work becomes failed and the lead is woken to replan. A failed dependency is never silently treated as satisfied.
Run ownership constrains concurrency and recovery
Each member stream is guarded by an in-process active-run table, a cross-node wake lock, and repository ownership keyed by runId. The active run refreshes a heartbeat, and release is conditional on the same runId. Even if two nodes receive the same event, only the run that owns the current claim can proceed.
On the server, a reaper scans stale heartbeats, conditionally reclaims ownership using both runId and the stale cutoff, then follows the normal exit path to mark work failed and notify the lead. Server startup also releases abandoned streams and resumes members from persisted team state. The event bus distributes change; correctness comes from fresh snapshots, conditional claims, and run ownership rather than an assumption of exactly-once delivery.
Every member run records its trigger, finish reason, step count, and token use. Parallelism remains an attributable, auditable resource instead of an opaque set of background calls.
Guarantees and responsibility boundaries
aibuddy can validate task dependencies, recover coordination state, persist messages before wake, converge task claims through conditional updates, and retain revision and run history. It cannot decide whether the decomposition is sensible, prevent two members from being instructed to edit the same file, or perform final acceptance for the lead.
Before introducing multiple agents, establish three boundaries: whether each branch owns distinct information or action, whether its delivery can be checked independently, and who resolves failure or conflicting results. Without those boundaries, parallelism mostly multiplies context, conflicts, and cost.
Implementation anchors
| Responsibility | Module |
|---|---|
| One-shot subagent isolation and result projection | task.tool |
| Team tool protocol and role enforcement | team.tool |
| Durable members, tasks, messages, and run records | TeamRepository |
| DAG validation, conditional claims, and versioned delivery | TeamCoordinator |
| Wake decisions over fresh snapshots | team-graph / team-wake-router |
| Member runs, heartbeats, and recovery | team-execution / team-member-reaper |
Related reading
- Classification and Criteria for deciding whether work benefits from parallelism.
- Subagents for the prompt, run, and return protocol of a bounded delegation.
- Agent Team Coordination for the user-facing member lifecycle and replanning model.