Tool retrieval should optimize capability coverage, not nearest-neighbour similarity

Agentic AI SeedlingPlanted Sep 2026

I want tool retrieval to optimize for whether an agent can complete the task, not whether its available tools resemble the request. Nearest-neighbour similarity ranks individual candidates. Capability coverage asks whether those candidates, taken together, support the operations the task requires. Those are different objectives. A retriever can return an apparently excellent list and still leave the agent without a necessary step. I treat the ranked list as input to context assembly, not as the finished tool set.

Consider a weather request that names a town, while the available forecast endpoint requires coordinates. Several forecast providers might closely match the request. None supplies the missing place-to-coordinate lookup. Loading more forecast schemas makes the list longer without making the workflow executable. A geocoding tool may look less like the original question and still be essential to answering it. That is the distinction I care about — relevance to the wording versus usefulness to the sequence of work.

I would make that sequence explicit enough to guide selection. A capability map should describe the operations a domain supports: looking up an entity, transforming an identifier, retrieving current or historical information, and performing the requested action. The point is not to invent a complete plan before execution. It is to expose dependencies that a single similarity score cannot represent. When I inspect a retrieved set, I want to know which required operations it covers and which remain unavailable, rather than just how highly each description scored.

This makes composition a separate engineering decision. I would first narrow the candidate pool to the relevant functional domains, then select complementary operations within them. A coherent set gives the model tools that belong to the same problem; coverage gives it distinct ways to advance that problem. Diversity is not permission to mix unrelated capabilities into the context. Nor should a category boundary prevent retrieval of a dependency from another domain. The organizing unit is the work to be completed, not the catalog folder that happens to contain the closest match.

Redundancy needs its own treatment. Near-identical tools consume context and ask the model to distinguish overlapping descriptions and schemas. I would remove interchangeable candidates before spending the remaining budget on missing operations. That decision should inspect what the tools actually do — similar descriptions are a signal to investigate, not proof that their capabilities are identical. I would also record what was removed. Otherwise a deduplication step can silently erase the very operation the composition stage was meant to preserve.

Coverage also has to evolve during execution. An early result can reveal a dependency that was not apparent in the initial request. I would start with a focused set, then retrieve against the newly identified subtask when it appears. Searching specifically for the missing operation is a better response than repeating the original broad query and collecting more of the same candidates. Each expansion still shares a context budget with conversation history and tool results. Loading incrementally is how I would keep coverage responsive without treating the entire catalog as permanent prompt material.

I would evaluate this at the task level. My test cases would identify the necessary capabilities, check whether retrieval made them available, and then check whether the agent completed the sequence. That separates a missing-tool failure from a wrong selection among available tools. Ranking quality alone cannot establish that distinction. I want failures to tell me whether to repair the catalog, the composition policy, or the agent’s use of an adequate set.

I concede a narrower case: for a genuinely single-operation request, with its required inputs already available and interchangeable candidate tools, nearest-neighbour selection may be sufficient. There is no missing dependency for a coverage policy to recover.

Beyond that boundary, I would make coverage the acceptance criterion and similarity the candidate-generation mechanism. The question is not whether every tool looks relevant. It is whether the available capabilities let the agent finish the work.