PyScrappy is an AI-native web scraping toolkit that turns websites into structured, LLM-ready data. Use it as a Python library or expose it as an MCP server for AI agents.
📖 Documentation: pyscrappy.vercel.app
- Generic scraper — give it any URL, get back structured text, links, images, tables, and metadata
- LLM-ready output —
.to_markdown()turns any result into clean Markdown; also.to_json()and.to_dataframe() - MCP server — expose the scrapers as tools for AI agents (Claude, Cursor, local LLMs, …)
- JS rendering — optional Playwright backend for JavaScript-heavy sites
- Custom selectors — pass CSS selectors to extract exactly what you need
- Concurrent scraping —
scrape_many/scrape_allrun scrapes in parallel - Proxy & scraping-API support — route through a proxy or ScraperAPI/ScrapeOps for blocked sites
- Retry & rate-limiting — built-in exponential backoff and per-domain rate limiting
- Type-safe — full type hints,
py.typedmarker - 20+ built-in scrapers — Wikipedia, IMDB, stocks, news, GitHub, Amazon/IKEA, YouTube, and more
pip install pyscrappyOptional extras: