Tools & Execution

The registry defines contracts; each run defines capability

aibuddy separates the tool system into a capability plane and an execution plane. The capability plane determines which actions the model can see in the current step. The execution plane determines where those actions may produce effects and under which authority. They intersect only at a concrete call.

Tool lifecycle from capability catalog to verified result The capability plane narrows built-in tools, Skills, and MCP into the active tool set. The execution plane validates, authorizes, executes, projects, and verifies each concrete call. A tool call is a controlled lifecycle Capabilities narrow into an active set; calls become verified effects. CAPABILITY PLANE · what the model can see in this step Capability sources built-ins · Skills · MCP Registry creation · mapping · hooks Progressive exposure direct or on-demand search Active tool set schemas sent in this step the model issues a call EXECUTION PLANE · how a call produces and proves an effect Validate schema · path · identifier Authorize host scope · confirmation Execute timeout · heartbeat · cancel AbortSignal Project results client · model · recovery Verify effects inspect state · check artifact Failures remain attributable to discovery, parameters, authority, execution, projection, or verification. registry ≠ authority · tool success ≠ task completion

Registration does not grant authority. The effective tool set for a run is the intersection of configured toolKeys, injected capability services, connected MCP servers, host sandbox support, and prerequisites such as team or knowledge state. A configuration with missing dependencies can return an empty tool set instead of exposing an action that must fail.

Each ServerToolConfig owns tool creation, runtime instructions, and result hooks. A filesystem group can generate ls, read_file, write_file, edit_file, glob, and grep, then map those names back to one configuration for shared client projection and compaction behavior. defineTool also validates schemas at definition time: JSON Schema must resolve synchronously and its top level must be an object.

At startup, assertConsistent compares registered configurations, TOOL_CONFIG_METAS, and name mappings. A missing built-in metadata entry, orphan declaration, or invalid mapping stops startup rather than allowing the UI and runtime to ship different tool sets silently.

The catalog narrows progressively into the active tool set

Built-in core tools remain directly visible. Skills initially expose only a name and purpose; the agent reads complete instructions and references through view_skill after selecting one. MCP tools are either exposed directly or deferred according to catalog size.

MCP currently uses two count gates:

ConditionExposure policy
No server exceeds 10 tools and total count does not exceed 30Schemas enter the active set directly
One server exceeds 10 toolsDefer only that server’s tools
Aggregate MCP count exceeds 30Defer all MCP tools

The gates use tool count rather than context-window size. A larger window can contain more schemas, but it does not remove selection ambiguity among similar actions. Deferred entries retain only names and descriptions in the index; tool_search surfaces candidates before their complete schemas become active.

Deferral adapts to the provider. Genuine Anthropic routes with native tool search and OpenAI BYOK routes using the Responses API can keep a stable tool array while the provider expands schemas on demand. Haiku, the platform OpenAI gateway, third-party compatibility endpoints, and other models use aibuddy’s activeTools path. Both strategies consume the same defer decision; only schema presentation differs.

Discovery state on the activeTools path is recreated for each agent run. To keep the protocol valid, aibuddy scans the post-compaction messages being sent. A deferred tool referenced by a retained historical tool-call is reactivated automatically and needs rediscovery only after that reference leaves context.

Host boundaries precede tool execution

A tool contract describes parameters and results; the host determines workspace, paths, credentials, and available backends. The same read_file semantics can operate in a hosted Web workspace or a Desktop project, but the call cannot exceed the scope exposed by the current sandbox implementation.

ISandbox provides one interface for listing, reading, writing, editing, searching, and command execution. Every operation accepts an AbortSignal. The Node adapter forwards it to filesystem work or child processes, and future remote adapters must propagate it into HTTP or WebSocket requests so stopping a task can release a hung operation.

The current createSandbox factory has one runnable adapter: Node. daytona and e2b already appear in SandboxType and configuration structures, but their adapters are placeholders. Selecting either returns an explicit unavailable error. Presence in a type union is not a claim of deployed support.

High-impact actions also require a separate human decision. confirm records the target, scope, and expected consequence before the effect and waits for approval or rejection. That decision does not replace operating-system or sandbox authority, and a similar historical approval must not be treated as permission for a new action.

Long-running execution combines heartbeats, cancellation, and output bounds

A long-running tool cannot merely wait for a Promise to resolve. For execute, aibuddy maintains three control paths:

  1. The sandbox timeout bounds command execution, with a further 30-second network grace period as a JavaScript-side hard deadline.
  2. An empty heartbeat is yielded every 30 seconds to re-arm the stream chunk timeout. Heartbeats never enter model context; stdout and stderr continue to reach the client through transient events.
  3. User stop and parent-agent cancellation propagate through AbortSignal into execution and background compaction.

Shell output is capped at 30,000 characters before it enters model context. The projection retains the head, tail, exit code, execution time, and the omitted middle length, while the client still receives the complete live stream. Enforcing this boundary as output is produced is more reliable than waiting for context pressure. When full detail is required, the agent can narrow the command or write output to a file and read bounded slices.

Filesystem tools likewise bound lines, matches, or returned entries at the source. Execution-time limits control one oversized result; later context compaction controls accumulation across steps. The two layers solve different failure modes.

Tool results use projections tailored to each consumer

One call serves at least the client, later model steps, and durable recovery. aibuddy does not require all three to share one representation:

ConsumerResult form
ClientDisplay-oriented fields, transient progress, and complete streamed output
Later model stepsCompact input and output produced by the tool-specific compact hook
Persistence and recoveryRaw input and output, stable toolCallId, and an optional recovery pointer

When a tool result reaches the current 2,500-token precompaction threshold, onToolExecutionEnd starts tool-specific compaction in the background without delaying delivery to the client. A compacted result enters cache and durable storage only when it is smaller than the original. A tool can choose structural trimming, model summarization, or drop. The first two can offload raw input and output to compacted_outputs/<taskId>/<toolCallId>.json and recover it by call ID through view_tool_call. drop is reserved for outputs explicitly declared disposable and creates no recovery pointer.

Completed human interactions use a separate projection. During context reconstruction, the scaffolding for ask_user_question and confirm normalizes into role=user text. The user’s decision remains authoritative without carrying an obsolete tool-call structure.

A successful response still requires effect verification

Tool success proves that a call completed, not that the task objective holds. A file write should be followed by reading the relevant region, a UI action by inspecting resulting state, and a build command by checking its exit code and artifacts. Failed verification should continue into correction rather than being translated into completion.

The lifecycle also gives failures a precise location:

StageTypical failure
DiscoveryA capability exists, but its description, Skill, or tool_search does not activate it
ParametersA path, identifier, or structured value loses fidelity between intent and schema
AuthorizationThe current user, workspace, or host does not permit the effect
ExecutionTimeout, cancellation, missing dependency, or a partially applied effect
ProjectionRaw output exists, but trimming or compaction removed evidence needed later
VerificationThe call succeeded, but the file, page, or service state still misses the objective

Stage-specific attribution supports safer recovery. Discovery errors call for a better index, environment errors for a different host, and partial effects for inspection before any retry.

Implementation anchors

Design responsibilityModule
Tool configuration, name mapping, and startup consistencytool-registry / ServerToolConfig
Input schema constraintsdefineTool
MCP count gates and active tool settool-discovery
Provider-native deferral and activeTools dispatchtool-defer
Sandbox capability and cancellation contractISandbox / createSandbox
Precompaction, persistence, and recovery by call IDtool-compactor / view_tool_call
  • Tool Design for granularity, parameter fidelity, and error semantics.
  • Agent Skills for reusable procedures loaded on demand.
  • MCP Integration for external services, namespaces, and credential boundaries.
  • Context Engineering for how tool schemas and results consume the model working set.
Was this page helpful?