Tools & Execution
The registry defines contracts; each run defines capability
aibuddy separates the tool system into a capability plane and an execution plane. The capability plane determines which actions the model can see in the current step. The execution plane determines where those actions may produce effects and under which authority. They intersect only at a concrete call.
Registration does not grant authority. The effective tool set for a run is the intersection of configured toolKeys, injected capability services, connected MCP servers, host sandbox support, and prerequisites such as team or knowledge state. A configuration with missing dependencies can return an empty tool set instead of exposing an action that must fail.
Each ServerToolConfig owns tool creation, runtime instructions, and result hooks. A filesystem group can generate ls, read_file, write_file, edit_file, glob, and grep, then map those names back to one configuration for shared client projection and compaction behavior. defineTool also validates schemas at definition time: JSON Schema must resolve synchronously and its top level must be an object.
At startup, assertConsistent compares registered configurations, TOOL_CONFIG_METAS, and name mappings. A missing built-in metadata entry, orphan declaration, or invalid mapping stops startup rather than allowing the UI and runtime to ship different tool sets silently.
The catalog narrows progressively into the active tool set
Built-in core tools remain directly visible. Skills initially expose only a name and purpose; the agent reads complete instructions and references through view_skill after selecting one. MCP tools are either exposed directly or deferred according to catalog size.
MCP currently uses two count gates:
| Condition | Exposure policy |
|---|---|
| No server exceeds 10 tools and total count does not exceed 30 | Schemas enter the active set directly |
| One server exceeds 10 tools | Defer only that server’s tools |
| Aggregate MCP count exceeds 30 | Defer all MCP tools |
The gates use tool count rather than context-window size. A larger window can contain more schemas, but it does not remove selection ambiguity among similar actions. Deferred entries retain only names and descriptions in the index; tool_search surfaces candidates before their complete schemas become active.
Deferral adapts to the provider. Genuine Anthropic routes with native tool search and OpenAI BYOK routes using the Responses API can keep a stable tool array while the provider expands schemas on demand. Haiku, the platform OpenAI gateway, third-party compatibility endpoints, and other models use aibuddy’s activeTools path. Both strategies consume the same defer decision; only schema presentation differs.
Discovery state on the activeTools path is recreated for each agent run. To keep the protocol valid, aibuddy scans the post-compaction messages being sent. A deferred tool referenced by a retained historical tool-call is reactivated automatically and needs rediscovery only after that reference leaves context.
Host boundaries precede tool execution
A tool contract describes parameters and results; the host determines workspace, paths, credentials, and available backends. The same read_file semantics can operate in a hosted Web workspace or a Desktop project, but the call cannot exceed the scope exposed by the current sandbox implementation.
ISandbox provides one interface for listing, reading, writing, editing, searching, and command execution. Every operation accepts an AbortSignal. The Node adapter forwards it to filesystem work or child processes, and future remote adapters must propagate it into HTTP or WebSocket requests so stopping a task can release a hung operation.
The current createSandbox factory has one runnable adapter: Node. daytona and e2b already appear in SandboxType and configuration structures, but their adapters are placeholders. Selecting either returns an explicit unavailable error. Presence in a type union is not a claim of deployed support.
High-impact actions also require a separate human decision. confirm records the target, scope, and expected consequence before the effect and waits for approval or rejection. That decision does not replace operating-system or sandbox authority, and a similar historical approval must not be treated as permission for a new action.
Long-running execution combines heartbeats, cancellation, and output bounds
A long-running tool cannot merely wait for a Promise to resolve. For execute, aibuddy maintains three control paths:
- The sandbox timeout bounds command execution, with a further 30-second network grace period as a JavaScript-side hard deadline.
- An empty heartbeat is yielded every 30 seconds to re-arm the stream chunk timeout. Heartbeats never enter model context; stdout and stderr continue to reach the client through transient events.
- User stop and parent-agent cancellation propagate through
AbortSignalinto execution and background compaction.
Shell output is capped at 30,000 characters before it enters model context. The projection retains the head, tail, exit code, execution time, and the omitted middle length, while the client still receives the complete live stream. Enforcing this boundary as output is produced is more reliable than waiting for context pressure. When full detail is required, the agent can narrow the command or write output to a file and read bounded slices.
Filesystem tools likewise bound lines, matches, or returned entries at the source. Execution-time limits control one oversized result; later context compaction controls accumulation across steps. The two layers solve different failure modes.
Tool results use projections tailored to each consumer
One call serves at least the client, later model steps, and durable recovery. aibuddy does not require all three to share one representation:
| Consumer | Result form |
|---|---|
| Client | Display-oriented fields, transient progress, and complete streamed output |
| Later model steps | Compact input and output produced by the tool-specific compact hook |
| Persistence and recovery | Raw input and output, stable toolCallId, and an optional recovery pointer |
When a tool result reaches the current 2,500-token precompaction threshold, onToolExecutionEnd starts tool-specific compaction in the background without delaying delivery to the client. A compacted result enters cache and durable storage only when it is smaller than the original. A tool can choose structural trimming, model summarization, or drop. The first two can offload raw input and output to compacted_outputs/<taskId>/<toolCallId>.json and recover it by call ID through view_tool_call. drop is reserved for outputs explicitly declared disposable and creates no recovery pointer.
Completed human interactions use a separate projection. During context reconstruction, the scaffolding for ask_user_question and confirm normalizes into role=user text. The user’s decision remains authoritative without carrying an obsolete tool-call structure.
A successful response still requires effect verification
Tool success proves that a call completed, not that the task objective holds. A file write should be followed by reading the relevant region, a UI action by inspecting resulting state, and a build command by checking its exit code and artifacts. Failed verification should continue into correction rather than being translated into completion.
The lifecycle also gives failures a precise location:
| Stage | Typical failure |
|---|---|
| Discovery | A capability exists, but its description, Skill, or tool_search does not activate it |
| Parameters | A path, identifier, or structured value loses fidelity between intent and schema |
| Authorization | The current user, workspace, or host does not permit the effect |
| Execution | Timeout, cancellation, missing dependency, or a partially applied effect |
| Projection | Raw output exists, but trimming or compaction removed evidence needed later |
| Verification | The call succeeded, but the file, page, or service state still misses the objective |
Stage-specific attribution supports safer recovery. Discovery errors call for a better index, environment errors for a different host, and partial effects for inspection before any retry.
Implementation anchors
| Design responsibility | Module |
|---|---|
| Tool configuration, name mapping, and startup consistency | tool-registry / ServerToolConfig |
| Input schema constraints | defineTool |
| MCP count gates and active tool set | tool-discovery |
Provider-native deferral and activeTools dispatch | tool-defer |
| Sandbox capability and cancellation contract | ISandbox / createSandbox |
| Precompaction, persistence, and recovery by call ID | tool-compactor / view_tool_call |
Related reading
- Tool Design for granularity, parameter fidelity, and error semantics.
- Agent Skills for reusable procedures loaded on demand.
- MCP Integration for external services, namespaces, and credential boundaries.
- Context Engineering for how tool schemas and results consume the model working set.