webscraping-ai-mcp-server

webscraping-ai-mcp-server

MCP.Pizza Chef: webscraping-ai

The webscraping-ai-mcp-server is a Model Context Protocol server that integrates with WebScraping.AI to provide powerful web data extraction capabilities. It supports structured data extraction, HTML content retrieval with JavaScript rendering, CSS selector-based extraction, and plain text extraction. Features include proxy support, device emulation, concurrent request management, and custom JavaScript execution, enabling robust and scalable web scraping workflows within MCP environments.

Use This MCP server To

Extract structured data from dynamic web pages with JavaScript rendering Retrieve and parse HTML content for analysis or processing Perform question answering based on live web page content Use CSS selectors to target and extract specific web page elements Manage concurrent web scraping requests with rate limiting Emulate different devices (desktop, mobile, tablet) for responsive scraping Execute custom JavaScript on target pages to manipulate or extract data Select proxies by type and country for anonymous or localized scraping Monitor API account usage to manage scraping quotas and limits

README

WebScraping.AI MCP Server

A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.

Features

  • Question answering about web page content
  • Structured data extraction from web pages
  • HTML content retrieval with JavaScript rendering
  • Plain text extraction from web pages
  • CSS selector-based content extraction
  • Multiple proxy types (datacenter, residential) with country selection
  • JavaScript rendering using headless Chrome/Chromium
  • Concurrent request management with rate limiting
  • Custom JavaScript execution on target pages
  • Device emulation (desktop, mobile, tablet)
  • Account usage monitoring

Installation

Running with npx

env WEBSCRAPING_AI_API_KEY=your_api_key npx -y webscraping-ai-mcp

Manual Installation

# Clone the repository
git clone https://github.com/webscraping-ai/webscraping-ai-mcp-server.git
cd webscraping-ai-mcp-server

# Install dependencies
npm install

# Run
npm start

Configuring in Cursor

Note: Requires Cursor version 0.45.6+

The WebScraping.AI MCP server can be configured in two ways in Cursor:

webscraping-ai-mcp-server FAQ

How do I install the webscraping-ai-mcp-server?
You can install it manually by cloning the GitHub repo and running npm install, or quickly run it with npx using your WebScraping.AI API key.
Can this server handle JavaScript-rendered web pages?
Yes, it uses headless Chrome/Chromium to render JavaScript content for accurate data extraction.
What types of proxies are supported?
It supports datacenter and residential proxies with country selection to optimize scraping reliability and anonymity.
How does the server manage multiple concurrent requests?
It includes built-in rate limiting and concurrent request management to ensure stable and efficient scraping.
Can I extract data using CSS selectors?
Yes, the server supports CSS selector-based content extraction for precise targeting of web elements.
Is device emulation available?
Yes, it can emulate desktop, mobile, and tablet devices to scrape responsive web pages accurately.
Can I run custom JavaScript on scraped pages?
Yes, you can execute custom JavaScript on target pages to manipulate or extract data as needed.
How do I monitor my API usage?
The server provides account usage monitoring to help you track and manage your WebScraping.AI API quotas.
Which LLM providers can I use with this MCP server?
This MCP server is model-agnostic and can be integrated with LLMs like OpenAI, Anthropic Claude, and Google Gemini for enhanced workflows.