webfetch

MCP.Pizza Chef: firish

This tool lets you run web searches locally for your AI assistant, cutting costs and token use drastically compared to hosted search services. It works with Claude Desktop and uses multiple search engines like DuckDuckGo without requiring keys, though keys can improve results. It caches queries intelligently to speed up repeated searches and lets your AI fetch fresh or cached web content as needed. Setup requires developer steps like installing Python packages and configuring environment variables.

Other
Web/Research

Use This MCP server To

Search the web from Claude Desktop without extra hosted fees Get faster web search results with fewer tokens used Cache repeated search queries to save cost and speed up responses Combine results from multiple search engines for better answers Fetch and compress web pages to show only relevant sentences Force fresh search results when needed Save and reuse model-learned findings from searches

README

webfetch

Web search for LLM agents that you run yourself - up to 8x fewer input tokens and 3x lower cost than hosted web_search, at the same accuracy.

Hosted web-search tools charge $10 per thousand searches and then bill you again for every token of retrieved content they push into your context window. webfetch replaces them with a local pipeline - multi-engine search, page fetching and extraction, semantic reranking, sentence-level compression - exposed as a web_search tool your model calls like any other. And unlike every hosted tool and search API we surveyed, repeated and paraphrased queries are served from a semantic cache for free.

(Install with pip install webfetch-llm; the import name is webfetch.)

Jump to: The headline · What you get · Getting started · Check your setup · Full benchmark results · Claude Code · Agent loop · Savings report · How it works · Caveats

Same accuracy. A third of the cost. An eighth of the tokens.

One agent loop, one model, one judge, 50 SimpleQA questions. The only thing that changes between rows is the search tool:

search tool accuracy input tok/query cost/query
Anthropic hosted web_search 96% 17,408 $0.108
webfetch (4-engine fusion) 92% 3,467 $0.035
webfetch (DDG only, $0 in fees) 84% 3,623 $0.026

Swap Opus for gpt-5.6-sol and the same webfetch tool hits 96% - hosted parity - at $0.040/query and 2,156 tokens: an eighth of what the hosted tool pushes into your context. Full results cover every arm we ran.

These numbers are the WORST case for webfetch - measured on an empty cache. In real use the gap widens on its own: repeats and rewords serve from cache for free, and the token advantage is paid again on every later turn that keeps search results in context. It adds up to receipts like this one, from an ordinary Claude Code session:

savings report rendered in Claude Code

Every claim in this README is generated by an eval harness that ships in this repo - the question sets, per-question records, judging protocol, and the negative results are all in evals/, and every table can be regenerated with one command. Don't take our word for the grading: evals/results/README.md maps every table row to the raw result file that produced it, down to per-question judge verdicts.

What you get

A search pipeline you own (4-engine RRF fusion, local extraction, sentence-level compression). Results come from reciprocal-rank fusion across DuckDuckGo, Brave, Serper, and Tavily - whichever of them you have keys for. DDG needs no key, so the tool works at literally zero cost out of the box; every key you add joins the fusion automatically. Pages are fetched and extracted locally (trafilatura, readability, newspaper4k, Playwright rendering for JS pages and 403 walls), chunked, ranked by a hybrid BM25 + bi-encoder cascade with a cross-encoder on top, then compressed to the sentences that answer the query - measured 50% fewer tokens at zero recall loss.

Caching nobody else has (exact + semantic matching, volatility-aware TTLs). Two layers in one sqlite file: page text by URL, ranked results by query. Identical queries hit an exact cache. Paraphrased queries hit a semantic cache - an embedding shortlist verified by an NLI cross-encoder, tuned eval-first for precision (zero wrong-target matches across every live run we have done). Cache lifetimes adapt to the query: prices and scores expire in 15 minutes, current-ish topics in 7 days, release notes and specs in 90 - classified by the calling model's hint or a local classifier. The model sees provenance on every cached result ([cache: semantic match to "...", 2h old, recent]) and can send force_fresh when it disagrees. No hosted tool or search API we surveyed offers any client-visible caching at all.

webfetch FAQ

Can I use webfetch to perform web searches within Claude Desktop?
Yes — webfetch integrates with Claude Desktop to provide local web search capabilities without extra hosted service costs.
Can I use webfetch to reduce the cost and token usage of web searches?
Yes — it cuts costs by up to two-thirds and uses up to eight times fewer tokens than hosted search tools.
Do I need API keys to use webfetch?
No — DuckDuckGo works without any key, but adding keys for Brave, Serper, or Tavily can improve search results.
How difficult is it to set up webfetch?
It requires developer-level setup involving installing packages and configuring environment variables.
Does webfetch cache search results?
Yes — it has a unique semantic caching system that saves repeated and paraphrased queries for free.
Can I force webfetch to get fresh search results?
Yes — the tool supports a 'force_fresh' option to bypass cached results.
Is webfetch free to use?
The core DuckDuckGo search is free; other engines may require API keys but often have free tiers.
Which apps support webfetch?
Currently, it is supported in Claude Desktop and can be added to other MCP-compatible clients with developer setup.