Tool Interface and Protocol (T)

This page corresponds to §4 of Agent Harness Engineering: A Survey. The tooling layer defines how an agent discovers capabilities, represents callable affordances, and executes actions across heterogeneous runtime boundaries.

The central tension is simple: more tools increase task coverage, but larger action spaces increase prompt footprint, selection error, planning error, and injection surface.

Protocol And Interface Standards

The paper arranges standards by integration boundary:

BoundaryStandardsWhat they standardize
Model to functionFunction callingTyped JSON tool calls inside one model runtime.
Agent to external capabilityMCP, OpenAPITool/resource/prompt access across process or service boundaries.
Agent to agentA2A, ACP, ANPDiscovery, delegation, long-running tasks, and streaming between opaque agents.
Agent to repository/environmentAGENTS.md / AGENT.mdLightweight version-controlled instructions for coding agents.

MCP’s practical value is not only schema interoperability. It is ecosystem mobility: teams can reuse a growing server catalog instead of writing a custom connector for every deployment. A2A addresses a neighboring boundary: agentic applications communicating with other agentic applications. The two are complementary.

Tool Description, Discovery, And Selection

Once protocols define how calls occur, the bottleneck becomes which tools should be surfaced and selected at each step. The survey groups systems such as EasyTool, AnyTool, CRAFT, MetaTool, MCP-Zero, ToolRet, ToolRegistry, SkillRouter, and SkillRet around one observation: registry quality and retrieval-aware orchestration are first-order determinants of downstream success.

Two engineering principles recur:

  • fewer, better-designed tools often outperform a large undifferentiated menu;
  • discovery must be adaptive because static global tool lists do not fit evolving repositories or multi-tenant enterprise deployments.

Tool-Augmented Training

Tool use is not only a runtime problem. Toolformer, Gorilla, ToolLLM/ToolBench, ToolkenGPT, and CREATOR show that models can learn when and how to call tools through self-supervision, instruction tuning, execution-oriented supervision, or controller-style integration.

For coding agents, the relevant tools are often semantic: static analyzers, type checkers, solvers, proof assistants, patch validators, and fault-localization checkers. The survey’s key requirement is that such tools should return evidence artifacts: traces, proof obligations, counterexamples, or structured certificates, not opaque yes/no answers.

Scalability And Session Management

Long-running agents repeatedly hit tool-session problems:

  • stale handles;
  • retries that see inconsistent tool state;
  • parallel calls that mutate shared state;
  • verbose tool traces that saturate the context window;
  • failure attribution that cannot distinguish planner mistakes from interface/protocol failures.

Effective harnesses need explicit tool-session lifecycles, bounded tool-context injection, and observability hooks that make tool failures diagnosable.

Design Implication

Tool capability is jointly determined by model training, interface schema quality, runtime selection strategy, session management, and governance. Improving only one of those pieces rarely fixes the whole layer.

Was this page helpful?