A Model Context Protocol (MCP) server for web scraping using Crawl4AI
Wellknown found it in public sources; nobody has proven control of it yet. Claiming takes one click if the repository is under your GitHub account, or a small file on your domain otherwise. Verified owners get the badge, 15-minute checks, status alerts, edits that outrank crawled data, and a ranking boost.
Agents can do it too: POST https://wellknown.network/api/v1/claims with {"agent":"realtimex-web-scraping-mcp-server","method":"well_known_file"} — machine-readable steps at claim.json, guide at /docs/claim.
Everything here was measured by our prober or read from a registry. Nothing is self-reported.
Attributed to the source that supplied each field. Treated as claims, not facts.
# Web Scraping MCP Server A Model Context Protocol (MCP) server that provides utility tools for web scraping using Crawl4AI as the core engine. Supports multiple extraction modes including markdown generation, HTML processing, structured data extraction via CSS selectors, and AI-powered semantic analysis. ## Features - **Multiple Extraction Modes**: Support for markdown, HTML, schema-based, and LLM-powered extraction - **Advanced Browser Automation**: Connects to a browser via CDP for full control, with proxy support and stealth features - **Structured Output**: Returns structured JSON data for programmatic consumption - **Session Management**: Persistent sessions for multi-step workflows - **Comprehensive Error Handling**: Robust error handling with detailed error information - **Type Safety**: Full type hints and Pydantic validation - **Cost-Aware LLM Usage**: Smart LLM usage with budget controls and caching ## Supported Extraction Modes ### 1. Markdown Extraction (`markdown`) - **Use Case**: Clean content extraction for analysis and documentation - **Output**: Clean markdown with citations and structured formatting - **Best For**: Articles, blog posts, documentation ### 2. HTML Extraction (`html`) - **Variants**: Clean HTML or raw HTML - **Use Case**: HTML processing and structure analysis - **Best For**: Template extraction, HTML analysis ### 3. Schema-Based Extraction (`schema`) - **Use Case**: Structured data extraction using CSS selectors - **Output**: JSON array of extracted objects - **Best For**: Product catalogs, data tables, listings ### 4. LLM-Powered Extraction (`llm`) - **Use Case**: AI-powered semantic extraction and analysis - **Output**: Structured JSON or text based on instructions - **Best For**: Content analysis, sentiment extraction, complex reasoning ## Quick Start ### Installation ```bash # Install via pip pip install web-scraping-mcp-server # Or use with uvx (no installation required) uvx web-scraping-mcp-server ``` ### Bootstr…
Mapped onto the structured taxonomy from declared text and observed tool names. Confidence shown for derived entries.
Every source is kept verbatim. Field changes are logged as events.