Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Read by thought-leaders and decision-makers around the world. Phone Number: +1-650-246-9381 Email: pub@towardsai.net
228 Park Avenue South New York, NY 10003 United States
Website: Publisher: https://towardsai.net/#publisher Diversity Policy: https://towardsai.net/about Ethics Policy: https://towardsai.net/about Masthead: https://towardsai.net/about
Name: Towards AI Legal Name: Towards AI, Inc. Description: Towards AI is the world's leading artificial intelligence (AI) and technology publication. Founders: Roberto Iriondo, , Job Title: Co-founder and Advisor Works for: Towards AI, Inc. Follow Roberto: X, LinkedIn, GitHub, Google Scholar, Towards AI Profile, Medium, ML@CMU, FreeCodeCamp, Crunchbase, Bloomberg, Roberto Iriondo, Generative AI Lab, Generative AI Lab VeloxTrend Ultrarix Capital Partners Denis Piffaretti, Job Title: Co-founder Works for: Towards AI, Inc. Louie Peters, Job Title: Co-founder Works for: Towards AI, Inc. Louis-François Bouchard, Job Title: Co-founder Works for: Towards AI, Inc. Cover:
Towards AI Cover
Logo:
Towards AI Logo
Areas Served: Worldwide Alternate Name: Towards AI, Inc. Alternate Name: Towards AI Co. Alternate Name: towards ai Alternate Name: towardsai Alternate Name: towards.ai Alternate Name: tai Alternate Name: toward ai Alternate Name: toward.ai Alternate Name: Towards AI, Inc. Alternate Name: towardsai.net Alternate Name: pub.towardsai.net
5 stars – based on 497 reviews

Frequently Used, Contextual References

TODO: Remember to copy unique IDs whenever it needs used. i.e., URL: 304b2e42315e

Resources

Free: 6-day Agentic AI Engineering Email Guide.
Learnings from Towards AI's hands-on work with real clients.
Claude Code Tool Search Nearly Halves Your Context Bill
Latest   Machine Learning

Claude Code Tool Search Nearly Halves Your Context Bill

Last Updated on October 6, 2026 by Editorial Team

Author(s): Decoding AI by Nueravi

Originally published on Towards AI.

Claude Code Tool Search Nearly Halves Your Context Bill
Claude Code 2.1.285 with 65 MCP tools and tool search on, captured 30 September 2026. Decoding AI.

Key takeaways

  • Claude Code tool search cut the first request from 27,184 tokens to 9,722 (-64%) with five MCP servers and 65 MCP tools connected. We captured both requests on 30 September 2026.
  • With no MCP servers at all, it still halves the bill: 17,546 tokens down to 8,756, because Claude Code 2.1.285 also defers 13 of its own 23 built-in tools.
  • A whole task costs 46% less, not 64%. Lazy loading needs one extra round trip to fetch the tool, and the task still came out 29,543 tokens against 54,422.
  • Loading one MCP tool on demand added 140 tokens. Keeping all 65 loaded cost 9,638.
  • Behind a custom ANTHROPIC_BASE_URL, tool search is off by default. Our proxied run sent every tool eagerly, byte for byte the same as forcing it off.

TL;DR

Does Claude Code tool search save tokens? Yes: with 65 MCP tools connected, it cut what Claude Code sends before your prompt from 27,184 tokens to 9,722, a 64% drop, measured on the wire. Tool search is Claude Code’s lazy tool loading: instead of sending every tool definition on every request, it sends a short list of tool names and a ToolSearch tool, and pulls in a full definition only when the model asks for it.

Status, 30 September 2026: Claude Code 2.1.285, official MCP reference servers 2026.8.31, @playwright/mcp 0.0.83. Re-check your own setup with the 18-line script below.

This is part 2 of The Context Tax, measured, a four-part series that publishes one part a day. Part 1 found that Claude Code sends about 18,000 tokens before you type, and that 82% of it is tool definitions. Today asks whether you can stop paying for tools you never call.

Place your bet before you scroll: over a whole task, does lazy loading save more than half? Answer in the comments. The result is in the middle of this post.

Readers’ leaderboard from part 1

Part 1 asked readers to run a capture script and post their number. No one has posted a number in the comments yet, so the board is still our six baselines: Aider 2,283, OpenCode 6,713, Gemini CLI 9,583, Codex CLI 10,020, Qwen Code 15,079 and Claude Code 18,006 tokens. Post yours under this part and it goes on tomorrow’s board.

What is Claude Code tool search?

Claude Code tool search is lazy tool loading: the agent keeps most tool definitions out of the request and loads each one only when the model searches for it. In our capture, Claude Code 2.1.285 with tool search on sent 11 tools in full, including ToolSearch itself, and replaced everything else with a one-line list of names. When the model needed a tool, it called ToolSearch, and the next request carried a reference that Anthropic's API expands into the full definition. The Anthropic tool search documentation says deferred tools do not enter the context window until discovered.

How much does lazy tool loading save on the first request?

With five MCP servers connected, Claude Code tool search cut the first request by 64%, from 27,184 tokens to 9,722. The five servers were the official filesystem, memory, everything and sequential-thinking reference servers plus Microsoft’s Playwright server, 65 MCP tools in total. Here is what each configuration sent before the word “hi”:

  • No MCP, eager: 17,546 tokens. 23 tools, 14,930 of them tool definitions.
  • No MCP, lazy: 8,756 tokens. 11 tools loaded, 6,052 tokens of definitions.
  • 65 MCP tools, eager: 27,184 tokens. 91 tools, 24,234 tokens of definitions.
  • 65 MCP tools, lazy: 9,722 tokens. The same 11 tools, plus a 966-token list of the names it can load.

Your MCP servers cost 9,638 tokens when loaded eagerly and 966 when deferred. That is a 90% cut on the line you configured yourself.

Same setup, tool search off and on. Decoding AI, 30 September 2026.

Does lazy loading save more than half on a whole task?

No, but close: over a whole task, tool search saved 46%, not the 64% of the first request. We gave Claude Code one job, “Remember that the deploy is on Friday”, which needs one MCP tool, the memory server’s create_entities. Eager mode took two requests and 54,422 tokens. Lazy mode took three, because the model first had to search for the tool, and still used only 29,543 tokens.

How did your bet do? The extra round trip cost about 9,700 tokens; every request after it saved about 17,300. Lazy loading was ahead from the first request.

What does loading a tool on demand cost?

Fetching create_entities on demand added 140 tokens to every later request in the session. That is its definition, now part of the context. Five loads cost about 700 tokens. Each request after the search was 9,933 tokens in lazy mode against 27,235 in eager mode, 64% less, so the saving compounds with every turn you take.

Why is Claude Code tool search not working with my proxy?

Because Claude Code 2.1.285 turns tool search off by default when ANTHROPIC_BASE_URL points anywhere other than Anthropic's own API. Our default run went through a local server, and it sent all 91 tools, byte for byte the same request as forcing tool search off. The binary's own log message explains the choice: it disables tool search for non-first-party hosts and tells you to set ENABLE_TOOL_SEARCH=true if your proxy forwards tool_reference blocks. If you route Claude Code through a gateway, you may be paying the full tool bill without knowing it.

Download the Medium app

Status line: checked in Claude Code 2.1.285 on 30 September 2026. Re-check with ENABLE_TOOL_SEARCH=true claude -p hi against the script below.

How did we measure this?

On 30 September 2026 we installed Claude Code 2.1.285 from npm in a clean sandbox and pointed it at a local stand-in server through ANTHROPIC_BASE_URL, in an empty folder with a fresh home directory. For the first-request census, the server recorded the request and returned an error, so no model answered. For the task, we played the model: the server replied with a fixed script (search for the tool if it is not loaded, call it, then say "Saved"), so eager and lazy differed only in what Claude Code sent. The memory server really wrote the note. Each mode ran three times; totals varied by at most 3 tokens. We counted tokens with OpenAI's o200k_base tokenizer, the same method as part 1, excluding the user's own prompt.

The captured requests, split and counted. Scripts are kept with the draft.

Your turn: run it both ways

Your turn. Tomorrow’s part opens with readers’ numbers.

Run this and post both numbers in the comments; tomorrow’s part opens with the leaderboard. Save it as context_tax.py, run python3 context_tax.py, then in your own project run ANTHROPIC_BASE_URL=http://127.0.0.1:8787 ENABLE_TOOL_SEARCH=false claude -p hi, then the same with ENABLE_TOOL_SEARCH=true. It counts the raw request JSON, so its totals run a little higher than ours; the gap between the two lines is what matters. Nothing leaves your machine.

# context_tax.py - what your agent sends before you type, eager vs lazy (stdlib only)
import http.server, json
def count(s):
try:
import tiktoken; return len(tiktoken.get_encoding("o200k_base").encode(s))
except Exception: return len(s) // 4 # rough estimate without tiktoken
class H(http.server.BaseHTTPRequestHandler):
def do_POST(self):
j = json.loads(self.rfile.read(int(self.headers["content-length"])))
tools = j.get("tools", [])
live = [t for t in tools if not t.get("defer_loading")]
rest = {k: j.get(k) for k in ("system", "messages")}
s, t = count(json.dumps(rest)), count(json.dumps(live))
print(f"tools loaded {len(live)} ({t:,} tokens) | deferred {len(tools)-len(live)} | TOTAL {s+t:,}", flush=True)
self.send_response(500); self.end_headers()
def log_message(self, *a): pass
print("listening on 127.0.0.1:8787", flush=True)
http.server.HTTPServer(("127.0.0.1", 8787), H).serve_forever()

On our five-server setup it printed 31,233 eager and 11,079 lazy. Smallest lazy total wins.

What this does not show

  • We played the model. A real model may search with keywords instead of an exact name, load more tools than it needs, or search twice. Each extra load costs its definition; each extra search costs a round trip.
  • Token counts, not bills. Prompt caching makes repeated prefixes cheaper after the first turn, and our tokenizer is not Anthropic’s. The ratios are what compare.
  • One task, one tool. A task that needs ten MCP tools narrows the gap.
  • We did not measure quality. Anthropic reports that tool search improved tool-selection accuracy on its own MCP evaluations with large tool libraries; we did not test that.
  • Our proxy finding comes from the default path and the binary’s own message, not from a trace of every gateway.

FAQ

How do I enable tool search in Claude Code? Set ENABLE_TOOL_SEARCH=true in your environment or in the env block of ~/.claude/settings.json. Setting it to false forces every tool to load up front.

Does tool search work with MCP servers? Yes. In our capture, all 65 MCP tools were deferred and replaced by a 966-token list of their names.

Does tool search slow Claude Code down? It adds one round trip the first time a tool is needed. In our task that was one extra request out of three.

Is tool search on by default? In 2.1.285 it is not when ANTHROPIC_BASE_URL points to a non-Anthropic host. The binary's message says it is otherwise the default; we could only capture the proxied path.

What is in part 3? Subagents: what each one pays before it starts working, and whether three of them cost three times one. It lands tomorrow.

Which MCP server in your setup would you never miss if it loaded only on demand?

This is The Context Tax, measured: four parts, one a day, each measuring one line on your agent’s bill with the method attached. Part 3 lands tomorrow and asks what a subagent pays before it does any work — follow Decoding AI to get it.

If this saved you a few thousand tokens a turn, clap so more developers see it, tell us in the comments which MCP server you would load only on demand and what your two numbers were, and follow Decoding AI — then hit subscribe to get part 3 by email tomorrow.

Sources & References

Join thousands of data leaders on the AI newsletter. Join over 80,000 subscribers and keep up to date with the latest developments in AI. From research to projects and ideas. If you are building an AI startup, an AI-related product, or a service, we invite you to consider becoming a sponsor.

Published via Towards AI


Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

→ 6-Day Agentic AI Engineering Email Guide — one practical lesson per day

→ Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

→ AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

→ Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

→ AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.