Tool Deferral
Why defer at all
A deferred tool is one the model can still call, but whose full JSON schema is kept out of its context until the model
searches for it. This matters once a user connects real MCP servers: as the MCP page explains, a
handful of servers can inject tens of thousands of tokens of tool definitions before the conversation even starts, and a
flat list of hundreds of tools measurably degrades the model’s selection accuracy. Deferral is the “give little by
default, load on demand” answer — the static prefix keeps only a short catalog, and a tool_search surfaces the full
schema of the few tools actually needed.
aibuddy decides whether to defer with one provider-agnostic rule (pickDeferredMcpTools), the same threshold on
every path: a single MCP server with more than 10 tools (DEFAULT_MCP_PER_SERVER_DEFER_COUNT) has its schemas deferred,
and once the connected MCP tools exceed 30 combined (DEFAULT_MCP_AGGREGATE_DEFER_COUNT) everything is. Built-in tools
are never deferred — the loadout is small and bounded. The interesting part is how a deferred tool is then surfaced,
and that is not one mechanism but a choice made per provider.
Pick two of three
Surfacing a deferred tool pulls on three properties that cannot all hold at once:
| Mechanism | Portable | Cache-stable | Native argument constraints |
|---|---|---|---|
activeTools filtering (SDK-level) | Yes — every provider | No — activation grows the tools array | Yes — real schema in the tools param |
Provider tool-search (defer_loading) | No — Anthropic / OpenAI only | Yes — tools array stays constant | Yes — server expands the real schema |
In-context wrapper (a call_tool proxy) | Yes — even weak models | Yes — one fixed tool | No — the model fills free-form args |
The prompt cache is the reason “cache-stable” is a real axis. On every provider the tool definitions sit at the front of
the cached prefix, so changing the tools array invalidates the prefix from that point on. activeTools narrows the
array per step, so the first time the model activates a new deferred tool the whole prefix must be re-written. Native
tool-search avoids this: all deferred tools stay declared with a defer_loading flag (a constant array), and the server
expands a referenced tool’s schema inline in the message history rather than by editing the front of the prefix.
Native argument constraints are why aibuddy does not use the in-context wrapper. When a tool is a real entry in the
tools param, the provider constrains generation to its schema; a wrapper trades that for a free-form arguments blob
that the model fills from a schema it read earlier in context — measurably more malformed calls, especially on weaker
models. aibuddy therefore takes the top two rows only: native tool-search where the provider offers it, activeTools
everywhere else. Both keep native constraints; the fallback simply gives up cache-stability, which is the cheaper thing
to lose on the providers that need it.
Strategy dispatch
resolveDeferralStrategy(modelId, providerKey) returns one of three strategies. It mirrors createModel’s routing
exactly — the deferral strategy must never claim a surface the request will not actually hit — so the distinctions below
are about the genuine provider, not merely the wire protocol.
| Model routing | Strategy | How deferred tools surface |
|---|---|---|
| Genuine Anthropic (first-party BYOK or platform gateway), not Haiku | anthropic | providerOptions.anthropic.deferLoading + server-executed BM25 tool_search |
| Genuine OpenAI BYOK on the Responses API | openai | providerOptions.openai.deferLoading + server-executed tool_search |
| Everything else | activeTools | discovery pool + aibuddy’s own tool_search, filtered via activeTools |
“Everything else” is a deliberately wide net, and two entries in it are easy to get wrong:
- Compat shims are not the genuine provider. DeepSeek exposes an
/anthropic-suffixed endpoint that speaks the Anthropic wire format but does not implementtool_reference; a naive “is this the Anthropic protocol?” check would route it to the native path and break it.usesGenuineAnthropicProtocolexcludes the shim. The same holds for third-party OpenAI-compatible endpoints (Moonshot and the like), which run on Chat Completions, not the Responses API. - OpenAI native requires the Responses API.
tool_searchanddefer_loadingare Responses-only, so aibuddy routes genuine OpenAI (api.openai.com, or an unset base URL) to.responses()and keeps third-party OpenAI-compatible endpoints on.chat(). The platform gateway routes opaquely, so gateway OpenAI stays on theactiveToolsfallback until confirmed. Anthropic has a single surface, so genuine Anthropic — gateway included — takes the native path.
Haiku is excluded from the native path because it does not support tool_reference expansion.
What each strategy does to the toolset
On a native strategy, applyNativeDeferral walks the toolset once: each deferred MCP tool keeps its identity as a
normal client-executed function tool — its execute still runs locally — and only gains a deferLoading flag under the
strategy’s provider namespace; then the provider’s server-executed tool_search tool is injected in place of aibuddy’s
own. The deferred tools stay declared on every request, so the tools array is constant and the prefix cache keeps hitting;
the server hides their schemas from the model’s context and expands the ones the model searches for. This is set up once
at turn setup (recorded on the context as nativeDeferral) and applied when the agent inputs are built.
On the activeTools fallback, a discovery pool is attached instead. The full toolset is registered, but each step’s
prepareStep narrows the visible set to a base list plus whatever the model has discovered via tool_search; a newly
discovered tool’s schema is appended at the end of the context. The two are mutually exclusive — a turn records either
nativeDeferral or a discovery pool, never both.
The activeTools cross-turn trap
The fallback has one correctness hazard worth calling out, because it is invisible until a long conversation hits it. The
discovery pool is created fresh per turn and starts empty. Suppose a turn discovers and calls an MCP tool — the history
now contains a tool_use for it. On the next turn the pool resets, so activeTools no longer lists that tool, yet the
replayed history still references it. The provider then sees a tool call for a tool absent from the current tools
param: at best the model can no longer follow up on it, at worst the request is rejected.
The fix keeps one invariant: activeTools always contains every deferred tool that history references. In prepareStep,
referencedDeferredToolNames scans the messages actually being sent (post-compaction) for tool-call parts naming a
deferred tool and feeds them back through pool.discover(). Because it reads the sent messages, it self-heals across
turns and never re-adds a reference that compaction has already dropped; because discover() is monotonic within a turn,
it only ever finds new names at a turn boundary, so it adds no extra cache churn. The native strategies are immune to this
entirely — their deferred tools stay declared on every request, so a historical reference is always backed by a
declaration.
Where it lives
| Concern | Symbol |
|---|---|
| Which tools to defer (shared policy) | pickDeferredMcpTools |
| Strategy selection (per provider) | resolveDeferralStrategy, usesGenuineAnthropicProtocol, usesOpenAIResponsesProtocol |
| Native marking + search-tool inject | applyNativeDeferral |
Fallback pool + activeTools list | createToolDiscovery |
| Cross-turn correctness seed | referencedDeferredToolNames (called from prepareStep) |
| Setup-time dispatch | tryAttachToolDiscovery |
Related reading
- MCP Integration — why deferral exists, and the tool-search / progressive-disclosure background
- prepareStep — the per-step hook where
activeToolsnarrowing happens - Agent Engine — the loop that assembles the toolset each turn