What happens when a model has access to more tools than fit in its head.
Tool calling works well in demonstrations, where a model is given five tools and asked to use one. It degrades badly at the scale real systems reach, where a useful assistant may have access to several thousand tools across dozens of servers.
Loading every tool schema into the context is not an option: the schemas alone can consume more context than the conversation. But a tool the model cannot see is a tool it will never call, so hiding them is not an option either.
Our approach treats tool discovery as a retrieval problem rather than a context problem. The model is given a small permanent set of core tools plus one meta-tool that searches the rest. It describes what it needs, gets back a handful of candidate schemas, and only then calls one.
The measurable effect is a large reduction in context spent on tools that were never going to be relevant, with no loss in the model's ability to reach the tool it needed.
Schemas are loaded in stages: name and one-line purpose first, then the full parameter schema only for the tool the model has actually decided to call. Most tools never need their full schema loaded at all.
An agent that can call tools will call tools. We study explicit budgets — how many calls a task is allowed, what happens when it runs out, and how a system decides that further tool use will not improve the answer. Knowing when to stop turns out to be a harder research problem than knowing what to call.
Real tools time out, rate-limit, return malformed data, and occasionally return confidently wrong data. We study how a model should distinguish those cases, when it should retry, when it should try a different tool, and when it should tell the user it cannot do the thing rather than inventing a result.
This work runs in production inside Ask Zai, where a user with no technical knowledge reaches a large tool library without ever seeing a tool.