Deferred Tools in MCP: How Lazy Loading Actually Works
Your agent doesn't need 50 tool schemas in context — it needs one search tool and an index. Here's how deferred loading works, where it lives (hint: not in the MCP protocol), and how each host implements it differently.
v0.1 — first public version. This guide follows kaizen rules: shipped incomplete, improved on a visible cadence. The changelog at the bottom tracks every revision, and the verification status section is honest about which claims I’ve tested firsthand versus compiled from documentation. Last checked against the MCP spec of 2026-07-28.
The problem: context bloat
Connect a few MCP servers to an agent and every tool definition — name, description, full JSON schema for every parameter — lands in the model’s context window on every single message. Real numbers people have reported: 66K tokens consumed before the first keystroke (Scott Spence), and Simon Willison naming context pollution as the reason he “rarely used MCP” at all.
Three fixes exist, and they compose:
- Tool search / lazy loading — don’t load schemas until needed. Anthropic reports ~85% token reduction. This guide’s subject.
- Code execution (“code mode”) — expose tools as code APIs in a sandbox; intermediate data never touches the model. 98.7% reduction in Anthropic’s demo (150K → 2K tokens).
- Skill-wraps-MCP — a skill’s frontmatter (~50 tokens) describes when to use a server; bundled scripts call it programmatically. Near-zero context until triggered.
Rule of thumb as of mid-2026: lazy loading is table stakes, code execution wins when you chain many calls or move large payloads, and skills win when the “tool” is really a procedure.
The two boundaries
The single most clarifying frame, and the one most explainers skip:
- Client ↔ server: the MCP client always fetches the full catalog via
tools/list. The server is completely unchanged by deferred loading. - Client ↔ model: the client decides which schemas the model’s context actually receives.
Deferred loading lives entirely at boundary 2. It is not an MCP protocol feature. If you’re a server author, there is nothing to implement. If you’re a client or framework author, this is entirely your problem.
The flow
- The client connects to all MCP servers and pulls full schemas — held outside model context.
- The model’s context gets a lightweight index (tool names + one-line descriptions) plus one meta-tool: tool search. Full schemas are absent — the model literally cannot call an unloaded tool, because it doesn’t know the parameter names.
- When the model needs a capability, it calls
tool_search(query=...). - The client injects the matching full definition(s) into context.
- The tool stays loaded for the rest of the conversation and is called normally.
The characteristic failure mode: the model guesses parameter names from a one-line description without searching first. The call fails — an unloaded tool isn’t wired up. Search first, then call.
API mechanics (Anthropic platform)
On the Claude API, all tool definitions go in the tools array as usual; deferred
ones are marked defer_loading: true. At least one tool stays non-deferred —
normally the search tool itself. Two built-in search variants exist,
tool_search_tool_regex_20251119 and tool_search_tool_bm25_20251119, and you can
supply a custom searcher (e.g. embeddings-based) that returns tool_reference
blocks, which the API expands the same way.
The subtle, load-bearing detail: discovered definitions expand inline at that point in the conversation body, not in the prefix. Your system prompt and always-loaded core tools remain a stable prefix — so prompt caching survives tool discovery.
For MCP servers specifically: mcp_toolset with
default_config: {defer_loading: true}, plus per-tool overrides to keep hot tools
always loaded.
The part no explainer states plainly
Deferred loading requires the API/client to support mid-conversation schema injection. Anthropic’s API has first-class machinery for this. Most frameworks don’t — they pass tools once, at the start of context. That’s why pydantic-ai (issue #3590) and the MCP go-sdk (issue #762) can’t trivially replicate it, and why every host reimplements deferral its own way.
Which is exactly why a cross-host comparison is the useful thing to maintain:
| Host | Deferred loading | Notes |
|---|---|---|
| Claude API | defer_loading + tool search tool | Inline expansion preserves prompt caching |
| Claude Code | Automatic | Activates when MCP tool descriptions exceed ~10% of context (~10K tokens); zero server-author effort |
| Claude Desktop / Cowork | Built in | Skills + code execution are first-class; skill-wraps-MCP most natural here |
| LibreChat | deferred_tools (default on) | Furthest along in open source; also has Programmatic Tool Calling (MCP tools executed inside the code sandbox) |
(This table is the section I most intend to deepen — per-host mechanics, versions, and config examples are coming in v0.2.)
What stateless MCP (2026-07-28) changes — and doesn’t
The 2026-07-28 spec made MCP stateless at the protocol layer: the
initialize/initialized handshake and Mcp-Session-Id header are gone, protocol
version and capabilities travel inline in _meta per request, and any server
instance can serve any request — plain round-robin load balancing works. A new
server/discover method provides stateless, cacheable capability discovery, and
tools/list results now carry ttlMs and cacheScope so clients and gateways can
cache catalogs HTTP-style.
Here’s the thing: none of that touches the context-window boundary. Stateless MCP changes the fetch/cache boundary — catalogs are identical from any replica and cacheable like HTTP resources — but context bloat is still fought entirely in the client layer, with the three fixes above.
If you maintain an MCP server, the concrete post-migration action item is small:
emit ttlMs/cacheScope on your list responses so gateways and clients can cache
your catalog.
Verification status
Kaizen honesty. As of v0.1:
- Compiled from vendor docs and community sources, checked 2026-08-01: the API mechanics, the Claude Code threshold, the LibreChat feature set, the stateless spec summary.
- Not yet independently tested by me:
defer_loadingbehavior against the API (including the inline-expansion/caching claim), and the exact Claude Code activation threshold. Firsthand tests are the v0.2 work, with a runnable demo repo.
Sources
- MCP specification & changelog — the 2026-07-28 stateless revision (checked 2026-08-01)
- Anthropic Engineering: Advanced tool use — the why, and the 85% / 98.7% reduction numbers (checked 2026-08-01)
- Claude platform docs: tool search tool — closest
thing to a canonical reference on
defer_loading(checked 2026-08-01) - LibreChat agents docs — the open-source implementation (checked 2026-08-01)
- pydantic-ai #3590 and go-sdk #762 — the framework-gap discussions
Changelog
- v0.1.1 · 2026-08-01 — added the provenance seal: this guide is human-directed and AI-drafted, with the Pangram receipt linked in the margin. The stated goal is to rewrite until it scores as fully human-written — progress will show in this changelog.
- v0.1 · 2026-08-01 — first public version. Two-boundaries framing, the flow, API mechanics, host comparison (shallow), stateless-MCP section. Known gaps: no diagrams, host table needs per-host depth, key claims not yet tested firsthand.