System Overview
The model loop is not the system boundary
One model call can generate text or request a tool, but it cannot by itself carry long-running work. A task may cross several model steps, external actions, user approvals, network interruptions, and background execution. If the system retains only chat text, it cannot reliably answer what is still running, which actions already occurred, or whether the result was delivered.
aibuddy therefore treats the Task as the stable execution unit. The model loop advances the current step; the Task retains the goal, identity, messages, state, workspace, and result. Chat, Team, and Schedule can initiate work through different product surfaces without defining separate agent semantics.
The diagram also defines the boundary between the Agent and System tracks. Agent explains why context, memory, tools, and coordination work. System explains how aibuddy assembles, executes, persists, and recovers those capabilities around one Task.
Interfaces turn change requirements into checkable conditions
A coding agent must answer four questions before changing the system: which layer owns the change, what the input and output are, which capabilities it may depend on, and which existing behaviors must remain unchanged. If those answers exist only in implementation details or team knowledge, the agent has to infer boundaries from a large call graph and is more likely to add platform conditions at the wrong layer.
aibuddy records those answers in shared interfaces and schemas. Interfaces state the permitted inputs, outputs, and dependencies; lifecycle states describe behavior that implementations must preserve. Web and Desktop choose their implementations behind those boundaries. The agent can therefore locate the intended extension point before limiting its change to one capability or platform adapter.
Desktop Task stream recovery is a concrete example. StreamBuffer first defines stream production, read, resume, and status semantics. Web can provide a Redis implementation while Desktop provides a process-local implementation. Adding Desktop recovery does not require copying or rewriting the upper Task flow; it requires implementing the same port and injecting it at the Desktop composition root. Behavioral tests for resume() continue to verify that backfill and live data have no gap or duplication.
Each contract answers a concrete question during a code change:
| What the agent must determine | Evidence supplied by the system | How drift becomes visible |
|---|---|---|
| Whether external data is valid | Types and Zod schemas in @aibuddy/common | Type checking or boundary parsing fails |
| Which layer owns an implementation | TaskClient, repository interfaces, and AppDeps | Dependencies cannot be assembled or interface members are missing |
| How a capability is invoked | ServerToolConfig, tool schemas, and Tool Registry | Definition validation or startup consistency checks fail |
| Which behavior must remain stable | Task, Run, Team, Delivery, and Stream lifecycles | State-transition or behavioral tests fail |
Interfaces do not guarantee that every agent change is correct. They reduce the rules an agent must infer and turn “this change does not follow the architecture” from a subjective review judgment into an explicit compile, parse, startup, or test failure.
A Task extends one call into complete work
A Task is neither one message nor one running process. It is a durable work record that connects at least five classes of fact: goal and execution identity, messages and tool activity, lifecycle state, workspace, and deliverables with usage.
This model has three direct consequences:
- Disconnecting a browser or Desktop window does not have to terminate the task.
- A model can wait for a user, an external event, or an upstream member without representing that wait as failure.
- After a process exits, the host can reassemble the task from persisted facts instead of relying on surviving in-memory objects.
The task runtime preserves these semantics. It enforces one active driver for a work unit, propagates cancellation, and assigns a terminal state only after messages, deliverables, usage, and the final event have been stored.
Every step reassembles the working set
Agent Runtime does not place every available fact and capability into model context. At the beginning of each step, it assembles input from current Task state: stable rules remain early, the goal and progress reflect current facts, and memory, knowledge, Skills, and long-tail tools appear first as indexes. Detailed content loads only after selection.
Progressive disclosure is a cross-module constraint in aibuddy, not only a knowledge-base technique:
| Capability | Disclosed first | Disclosed after selection |
|---|---|---|
| User memory | Manifest of types, names, and descriptions | Matching memory content |
| Knowledge | Document map and section structure | Search evidence and bounded section content |
| Skills | Name and applicability | Full procedure, scripts, and references |
| MCP tools | Server and tool summaries | Tool schema and execution entry point |
The capability catalog can grow while each model call keeps a task-bounded selection surface and token cost. Failure to discover a capability, failure to authorize it, and failure during execution remain separately diagnosable.
Source facts remain separate from runtime representations
Long tasks require compaction, but compaction must not rewrite what occurred. aibuddy retains original messages and tool results while storing compaction snapshots, offload pointers, and task state as derived representations for later execution. The next step can use a smaller context while users and operators can still inspect the source record.
The same rule governs delivery. After the agent produces a file in a sandbox, the delivery service materializes confirmed bytes into platform file storage and creates an Artifact record owned by the Task. Preview and download do not require the original task sandbox to remain alive.
The runtime therefore emits two classes of output. User-facing output includes streamed replies, tool activity, and Deliverables. Recovery and operations use messages, state, usage, events, and file indexes. Both belong to the same Task, but answer different questions; an operational trace cannot replace the record the user actually saw.
Collaboration shares facts without sharing every reasoning trace
Subagents and Agent Team both operate around Tasks. A subagent can share its parent’s workspace while running in an isolated context, then return only its summary, evidence, and deliverable locations. Agent Team adds a task graph of members, dependencies, and messages, and wakes runnable members only after state is committed.
This separation prevents the lead agent from carrying every parallel trace while preserving accountability. A member completing its branch is a local delivery; the coordinator remains responsible for integrating and verifying the whole task.
Platform adapters change infrastructure, not task semantics
Web and Desktop share the domain semantics of Tasks, messages, tool events, Deliverables, and Agent Runtime. Web binds those interfaces to Server, PostgreSQL, Redis, object storage, and workers. Desktop binds them to Electron Main, SQLite, local files, and host execution.
Shared interfaces do not promise platform parity. Web can use Redis for cross-instance locks, cross-instance stream recovery, and queue retries. Desktop provides process-local locks and in-memory stream recovery for the lifetime of its main process, together with local files, device credentials, and stdio MCP. System documentation states the scope of each guarantee instead of hiding different failure boundaries behind one interface name.
Seven constraints hold the system together
| Constraint | Concrete effect |
|---|---|
| Contracts precede implementations | Shared schemas, service ports, tool definitions, and lifecycle transitions establish the boundary before Core and host implementations supply behavior |
| One active driver | Locks and run registration prevent two concurrent trajectories for one work unit |
| Connection is not execution | UI disconnection does not directly change background task state |
| Capabilities disclose on demand | Memory, knowledge, Skills, and MCP content do not remain fully resident in context |
| Projections cannot overwrite source facts | Compaction snapshots and delivery indexes remain separate from original messages and files |
| Non-critical enhancement can degrade | Summary, title, or background memory extraction failure does not block the main task |
| Completion follows durable finalization | Messages, deliverables, usage, and terminal events are stored before completion is exposed |
These constraints define system behavior more strongly than one model or tool list. Models, storage, and transport can change; violating any one of them weakens recoverability or verification.
Current system boundaries
- aibuddy uses an AI Gateway or user-configured providers and does not ship model weights.
- Node is the currently available sandbox backend; Daytona and E2B remain incomplete adapters.
- Memory, knowledge, Prompt, and Skills have explicit update channels, but automated eval, canary rollout, and rollback are not yet one release loop.
- The administrative evaluation screen still includes sample data and is not a unified production evaluation backend.
- Web and Desktop share domain behavior, but queue, recovery, MCP, and identity guarantees remain host-specific.
Related reading
- Task runtime — mutual exclusion, cancellation, waiting, recovery, and finalization around a durable Task.
- Context engineering — how each step assembles information while preserving source facts.
- Memory and knowledge — two-stage memory recall and isolated knowledge retrieval.
- Tools and execution — discovery, authorization, host execution, and result verification.
- Product and platforms — product objects and Web/Desktop adapter boundaries.