Tool Use at Scale: What Breaks When Your Agent Has 40 Tools
Give an agent six tools (search the knowledge base, look up a supplier record, check PO status, create a task, send a notification, escalate to human review) and everything works. It picks the right tool reliably, latency is acceptable, and the permission surface is manageable. Give an agent 43 tools, because it has to integrate with a full ERP stack, and it does not work well. Tool selection errors appear at rates the small set never produced. Latency per request climbs. And one bad decision in a long chain of tool calls can trigger expensive actions that are hard to reverse. The problems are qualitatively different at scale.
The tool selection problem
With six tools, a well-prompted agent makes the right selection nearly every time. The choice space is small, the tools are distinct, and the model's reasoning about which to use is straightforward. With 40+ tools, you are asking the model to navigate a much larger decision space, often with tools that serve overlapping purposes.
Take an ERP integration agent with tools for creating a purchase order, updating one, creating a draft, duplicating one from a template, and creating a blanket PO. These are all distinct operations that serve real purposes. To a human AP manager, the differences are obvious from context. To the model, given an instruction like "set up a recurring order for maintenance supplies," the choice between those five is often wrong.
Two changes fix this. First, tool naming discipline: every tool gets a name that is unambiguous about its effect and scope. Not "create_po" but "create_purchase_order_from_scratch" and "create_recurring_blanket_po": names long enough to describe themselves. Second, tool docstrings that include explicit negative examples: "Use this for one-time purchases. Do NOT use this for recurring or scheduled orders." Negative examples reduced confusion significantly on overlapping tools.
The model uses tool names and descriptions, not just the function signature, to decide which tool to call. Write tool descriptions as if you are writing them for a competent but unfamiliar colleague, not as API documentation. Include what NOT to use the tool for. Those two changes remove most selection errors.
Latency compounds with tool chain depth
A single LLM call takes 1–3 seconds. A single tool call (API roundtrip to an ERP) takes 0.5–2 seconds. With a simple agent that makes one or two tool calls per request, total latency is tolerable. A complex agent might make six to ten tool calls to complete a task: retrieving context, verifying preconditions, executing the action, confirming the result. That adds up to 15–25 seconds per request before any retry logic.
Three optimizations bring that down to something a user will tolerate:
- →Parallel tool execution. When the agent needs to fetch context from multiple sources before acting, those fetches can run concurrently. LangGraph's parallel node execution made this straightforward to implement. Most of our context fetches went from sequential to concurrent.
- →Tool result caching. Supplier records, PO templates and approval thresholds change rarely. Cache tool results with sensible TTLs (supplier record: 10 minutes, approval threshold: 1 hour, PO template: 24 hours).
- →Streaming for user-facing responses. For tasks where the user is waiting, stream the agent's reasoning while tool calls execute in the background. The user sees progress instead of a blank screen.
Permission surface grows faster than you expect
Each tool an agent can call is a potential action it might take. With six tools, the action surface is limited and auditable. With 40, the combinations of actions the agent might chain together create an audit surface that is genuinely hard to reason about.
Here is the kind of problem that appears. An ERP agent can create a PO, look up and apply a vendor discount schedule, and submit the PO for approval. Each of those is correct and authorized on its own. Chain them together (create a PO, apply the maximum term discount regardless of order size, auto-submit before anyone reviews the line items) and you get POs with incorrect pricing that sail through approval, because they are correctly formatted and look like normal output.
Our take
The tool registry pattern
Above about 20 tools, stop including every tool in every agent invocation. Use a tool registry instead: all available tools are registered with metadata (category, required permissions, risk level, typical use cases). At the start of each task, a lightweight routing step selects the relevant subset, usually 8–12 tools, based on the task type.
This has three benefits. First, it reduces the tool selection error rate by removing irrelevant tools from the choice space. An agent working on a supplier payment query does not need the tools for HR data access or report generation. Second, it reduces context window consumption, since 40 tool definitions add meaningful tokens to every call. Third, it allows per-task permission scoping: the tool subset for a read-only lookup task does not include write tools, even if the agent's service key technically has write access.
The tool registry is the architectural decision that does most for reliability in a high-tool-count agent. It is a few days of work in a framework like LangGraph, and it pays for itself in better tool selection and fewer accidental high-impact actions.When to split an agent instead of adding more tools
There is a threshold where adding more tools to a single agent stops being the right architecture. We put it at around 25–30 tools for a single agent without a tool registry, or 40–50 with one. Beyond that, the load on the model (navigating the tool space, keeping the task coherent, handling errors from a wider range of tool call patterns) starts to degrade reliability in ways that tool descriptions and testing cannot fully fix.
The alternative is a multi-agent architecture: a coordinator agent that understands the task and delegates to specialized sub-agents, each with a smaller, coherent tool set. A procurement coordinator delegates to a supplier lookup agent (5 tools), a PO management agent (8 tools), and an approval workflow agent (4 tools). Each is testable independently. Each has a clear permission scope. The coordinator does not need to understand all 17 tools. It only needs to route the task to the right specialist.
Multi-agent architectures add orchestration complexity. They are worth the tradeoff when a single agent's tool count is producing reliability problems that you have already tried to solve with naming, documentation, and tool registries. Do not reach for multi-agent because it sounds more sophisticated. Reach for it when you have hit the reliability ceiling of the simpler approach.