System Overview

The model loop is not the system boundary

One model call can generate text or request a tool, but it cannot by itself carry long-running work. A task may cross several model steps, external actions, user approvals, network interruptions, and background execution. If the system retains only chat text, it cannot reliably answer what is still running, which actions already occurred, or whether the result was delivered.

aibuddy therefore treats the Task as the stable execution unit. The model loop advances the current step; the Task retains the goal, identity, messages, state, workspace, and result. Chat, Team, and Schedule can initiate work through different product surfaces without defining separate agent semantics.

aibuddy system runtime model The model loop handles steps; a durable Task carries the complete work PRODUCT ENTRY Task / Chat Team Schedule Durable Task recordgoal · identity · messages · current state · workspace ASSEMBLED FOR EACH AGENT STEP Context assemblystate · memory · knowledgeindexes first, content on demand Agent Runtimeassess → act → observecancel, wait, and finalize share one lifecycle Capability and executionTools · Skills · MCPhost and permission bound access User-visible resultsstream · tool activity · deliverables Recoverable factsmessages · state · usage · events · file index same domain semantics, different host adapters Web · Server / Worker / PostgreSQL Desktop · Electron / SQLite / Local FS

The diagram also defines the boundary between the Agent and System tracks. Agent explains why context, memory, tools, and coordination work. System explains how aibuddy assembles, executes, persists, and recovers those capabilities around one Task.

Interfaces turn change requirements into checkable conditions

A coding agent must answer four questions before changing the system: which layer owns the change, what the input and output are, which capabilities it may depend on, and which existing behaviors must remain unchanged. If those answers exist only in implementation details or team knowledge, the agent has to infer boundaries from a large call graph and is more likely to add platform conditions at the wrong layer.

aibuddy records those answers in shared interfaces and schemas. Interfaces state the permitted inputs, outputs, and dependencies; lifecycle states describe behavior that implementations must preserve. Web and Desktop choose their implementations behind those boundaries. The agent can therefore locate the intended extension point before limiting its change to one capability or platform adapter.

How an interface constrains a code change The agent locates, implements, and verifies against explicit boundaries instead of inferred conventions Request: add Desktop Task stream recovery locate the change before implementing it ① The interface states what must be preserved change location · input and output · allowed dependencies · behavioral invariants StreamBuffer specifies produce · read · resume · getStatus ② The agent changes only the implementation behind it The shared Task flow keeps the same port and gains no Desktop-only branch Web · Redis adapter Shared Task flow unchanged Desktop · memory adapter ③ Contract drift becomes an explicit failure missing method → typecheck invalid input → schema registry drift → startup resume gap → behavior test The agent corrects the identified failure before delivery

Desktop Task stream recovery is a concrete example. StreamBuffer first defines stream production, read, resume, and status semantics. Web can provide a Redis implementation while Desktop provides a process-local implementation. Adding Desktop recovery does not require copying or rewriting the upper Task flow; it requires implementing the same port and injecting it at the Desktop composition root. Behavioral tests for resume() continue to verify that backfill and live data have no gap or duplication.

Each contract answers a concrete question during a code change:

What the agent must determineEvidence supplied by the systemHow drift becomes visible
Whether external data is validTypes and Zod schemas in @aibuddy/commonType checking or boundary parsing fails
Which layer owns an implementationTaskClient, repository interfaces, and AppDepsDependencies cannot be assembled or interface members are missing
How a capability is invokedServerToolConfig, tool schemas, and Tool RegistryDefinition validation or startup consistency checks fail
Which behavior must remain stableTask, Run, Team, Delivery, and Stream lifecyclesState-transition or behavioral tests fail

Interfaces do not guarantee that every agent change is correct. They reduce the rules an agent must infer and turn “this change does not follow the architecture” from a subjective review judgment into an explicit compile, parse, startup, or test failure.

A Task extends one call into complete work

A Task is neither one message nor one running process. It is a durable work record that connects at least five classes of fact: goal and execution identity, messages and tool activity, lifecycle state, workspace, and deliverables with usage.

This model has three direct consequences:

  1. Disconnecting a browser or Desktop window does not have to terminate the task.
  2. A model can wait for a user, an external event, or an upstream member without representing that wait as failure.
  3. After a process exits, the host can reassemble the task from persisted facts instead of relying on surviving in-memory objects.

The task runtime preserves these semantics. It enforces one active driver for a work unit, propagates cancellation, and assigns a terminal state only after messages, deliverables, usage, and the final event have been stored.

Every step reassembles the working set

Agent Runtime does not place every available fact and capability into model context. At the beginning of each step, it assembles input from current Task state: stable rules remain early, the goal and progress reflect current facts, and memory, knowledge, Skills, and long-tail tools appear first as indexes. Detailed content loads only after selection.

Progressive disclosure is a cross-module constraint in aibuddy, not only a knowledge-base technique:

CapabilityDisclosed firstDisclosed after selection
User memoryManifest of types, names, and descriptionsMatching memory content
KnowledgeDocument map and section structureSearch evidence and bounded section content
SkillsName and applicabilityFull procedure, scripts, and references
MCP toolsServer and tool summariesTool schema and execution entry point

The capability catalog can grow while each model call keeps a task-bounded selection surface and token cost. Failure to discover a capability, failure to authorize it, and failure during execution remain separately diagnosable.

Source facts remain separate from runtime representations

Long tasks require compaction, but compaction must not rewrite what occurred. aibuddy retains original messages and tool results while storing compaction snapshots, offload pointers, and task state as derived representations for later execution. The next step can use a smaller context while users and operators can still inspect the source record.

The same rule governs delivery. After the agent produces a file in a sandbox, the delivery service materializes confirmed bytes into platform file storage and creates an Artifact record owned by the Task. Preview and download do not require the original task sandbox to remain alive.

The runtime therefore emits two classes of output. User-facing output includes streamed replies, tool activity, and Deliverables. Recovery and operations use messages, state, usage, events, and file indexes. Both belong to the same Task, but answer different questions; an operational trace cannot replace the record the user actually saw.

Collaboration shares facts without sharing every reasoning trace

Subagents and Agent Team both operate around Tasks. A subagent can share its parent’s workspace while running in an isolated context, then return only its summary, evidence, and deliverable locations. Agent Team adds a task graph of members, dependencies, and messages, and wakes runnable members only after state is committed.

This separation prevents the lead agent from carrying every parallel trace while preserving accountability. A member completing its branch is a local delivery; the coordinator remains responsible for integrating and verifying the whole task.

Platform adapters change infrastructure, not task semantics

Web and Desktop share the domain semantics of Tasks, messages, tool events, Deliverables, and Agent Runtime. Web binds those interfaces to Server, PostgreSQL, Redis, object storage, and workers. Desktop binds them to Electron Main, SQLite, local files, and host execution.

Shared interfaces do not promise platform parity. Web can use Redis for cross-instance locks, cross-instance stream recovery, and queue retries. Desktop provides process-local locks and in-memory stream recovery for the lifetime of its main process, together with local files, device credentials, and stdio MCP. System documentation states the scope of each guarantee instead of hiding different failure boundaries behind one interface name.

Seven constraints hold the system together

ConstraintConcrete effect
Contracts precede implementationsShared schemas, service ports, tool definitions, and lifecycle transitions establish the boundary before Core and host implementations supply behavior
One active driverLocks and run registration prevent two concurrent trajectories for one work unit
Connection is not executionUI disconnection does not directly change background task state
Capabilities disclose on demandMemory, knowledge, Skills, and MCP content do not remain fully resident in context
Projections cannot overwrite source factsCompaction snapshots and delivery indexes remain separate from original messages and files
Non-critical enhancement can degradeSummary, title, or background memory extraction failure does not block the main task
Completion follows durable finalizationMessages, deliverables, usage, and terminal events are stored before completion is exposed

These constraints define system behavior more strongly than one model or tool list. Models, storage, and transport can change; violating any one of them weakens recoverability or verification.

Current system boundaries

  • aibuddy uses an AI Gateway or user-configured providers and does not ship model weights.
  • Node is the currently available sandbox backend; Daytona and E2B remain incomplete adapters.
  • Memory, knowledge, Prompt, and Skills have explicit update channels, but automated eval, canary rollout, and rollback are not yet one release loop.
  • The administrative evaluation screen still includes sample data and is not a unified production evaluation backend.
  • Web and Desktop share domain behavior, but queue, recovery, MCP, and identity guarantees remain host-specific.
Was this page helpful?