kimsia
Budding · Updated Aug 2026 ·6 min

Deferred Tools in MCP: How Lazy Loading Actually Works

Your agent doesn't need 50 tool schemas in context — it needs one search tool and an index. Here's how deferred loading works, where it lives (hint: not in the MCP protocol), and how each host implements it differently.

Note

v0.1 — first public version. This guide follows kaizen rules: shipped incomplete, improved on a visible cadence. The changelog at the bottom tracks every revision, and the verification status section is honest about which claims I’ve tested firsthand versus compiled from documentation. Last checked against the MCP spec of 2026-07-28.

The problem: context bloat

Connect a few MCP servers to an agent and every tool definition — name, description, full JSON schema for every parameter — lands in the model’s context window on every single message. Real numbers people have reported: 66K tokens consumed before the first keystroke (Scott Spence), and Simon Willison naming context pollution as the reason he “rarely used MCP” at all.

Three fixes exist, and they compose:

  1. Tool search / lazy loading — don’t load schemas until needed. Anthropic reports ~85% token reduction. This guide’s subject.
  2. Code execution (“code mode”) — expose tools as code APIs in a sandbox; intermediate data never touches the model. 98.7% reduction in Anthropic’s demo (150K → 2K tokens).
  3. Skill-wraps-MCP — a skill’s frontmatter (~50 tokens) describes when to use a server; bundled scripts call it programmatically. Near-zero context until triggered.

Rule of thumb as of mid-2026: lazy loading is table stakes, code execution wins when you chain many calls or move large payloads, and skills win when the “tool” is really a procedure.

The two boundaries

The single most clarifying frame, and the one most explainers skip:

  1. Client ↔ server: the MCP client always fetches the full catalog via tools/list. The server is completely unchanged by deferred loading.
  2. Client ↔ model: the client decides which schemas the model’s context actually receives.

Deferred loading lives entirely at boundary 2. It is not an MCP protocol feature. If you’re a server author, there is nothing to implement. If you’re a client or framework author, this is entirely your problem.

The flow

  1. The client connects to all MCP servers and pulls full schemas — held outside model context.
  2. The model’s context gets a lightweight index (tool names + one-line descriptions) plus one meta-tool: tool search. Full schemas are absent — the model literally cannot call an unloaded tool, because it doesn’t know the parameter names.
  3. When the model needs a capability, it calls tool_search(query=...).
  4. The client injects the matching full definition(s) into context.
  5. The tool stays loaded for the rest of the conversation and is called normally.
Warning

The characteristic failure mode: the model guesses parameter names from a one-line description without searching first. The call fails — an unloaded tool isn’t wired up. Search first, then call.

API mechanics (Anthropic platform)

On the Claude API, all tool definitions go in the tools array as usual; deferred ones are marked defer_loading: true. At least one tool stays non-deferred — normally the search tool itself. Two built-in search variants exist, tool_search_tool_regex_20251119 and tool_search_tool_bm25_20251119, and you can supply a custom searcher (e.g. embeddings-based) that returns tool_reference blocks, which the API expands the same way.

The subtle, load-bearing detail: discovered definitions expand inline at that point in the conversation body, not in the prefix. Your system prompt and always-loaded core tools remain a stable prefix — so prompt caching survives tool discovery.

For MCP servers specifically: mcp_toolset with default_config: {defer_loading: true}, plus per-tool overrides to keep hot tools always loaded.

The part no explainer states plainly

Deferred loading requires the API/client to support mid-conversation schema injection. Anthropic’s API has first-class machinery for this. Most frameworks don’t — they pass tools once, at the start of context. That’s why pydantic-ai (issue #3590) and the MCP go-sdk (issue #762) can’t trivially replicate it, and why every host reimplements deferral its own way.

Which is exactly why a cross-host comparison is the useful thing to maintain:

HostDeferred loadingNotes
Claude APIdefer_loading + tool search toolInline expansion preserves prompt caching
Claude CodeAutomaticActivates when MCP tool descriptions exceed ~10% of context (~10K tokens); zero server-author effort
Claude Desktop / CoworkBuilt inSkills + code execution are first-class; skill-wraps-MCP most natural here
LibreChatdeferred_tools (default on)Furthest along in open source; also has Programmatic Tool Calling (MCP tools executed inside the code sandbox)

(This table is the section I most intend to deepen — per-host mechanics, versions, and config examples are coming in v0.2.)

What stateless MCP (2026-07-28) changes — and doesn’t

The 2026-07-28 spec made MCP stateless at the protocol layer: the initialize/initialized handshake and Mcp-Session-Id header are gone, protocol version and capabilities travel inline in _meta per request, and any server instance can serve any request — plain round-robin load balancing works. A new server/discover method provides stateless, cacheable capability discovery, and tools/list results now carry ttlMs and cacheScope so clients and gateways can cache catalogs HTTP-style.

Here’s the thing: none of that touches the context-window boundary. Stateless MCP changes the fetch/cache boundary — catalogs are identical from any replica and cacheable like HTTP resources — but context bloat is still fought entirely in the client layer, with the three fixes above.

Insight

If you maintain an MCP server, the concrete post-migration action item is small: emit ttlMs/cacheScope on your list responses so gateways and clients can cache your catalog.

Verification status

Kaizen honesty. As of v0.1:

  • Compiled from vendor docs and community sources, checked 2026-08-01: the API mechanics, the Claude Code threshold, the LibreChat feature set, the stateless spec summary.
  • Not yet independently tested by me: defer_loading behavior against the API (including the inline-expansion/caching claim), and the exact Claude Code activation threshold. Firsthand tests are the v0.2 work, with a runnable demo repo.

Sources

Changelog

  • v0.1.1 · 2026-08-01 — added the provenance seal: this guide is human-directed and AI-drafted, with the Pangram receipt linked in the margin. The stated goal is to rewrite until it scores as fully human-written — progress will show in this changelog.
  • v0.1 · 2026-08-01 — first public version. Two-boundaries framing, the flow, API mechanics, host comparison (shallow), stateless-MCP section. Known gaps: no diagrams, host table needs per-host depth, key claims not yet tested firsthand.