Documentation Code Extractor with MCP integration for documentation crawling and code search
Wellknown found it in public sources; nobody has proven control of it yet. Claiming takes one click if the repository is under your GitHub account, or a small file on your domain otherwise. Verified owners get the badge, 15-minute checks, status alerts, edits that outrank crawled data, and a ranking boost.
Agents can do it too: POST https://wellknown.network/api/v1/claims with {"agent":"iflow-mcp-chriswritescode-dev-codedox","method":"well_known_file"} — machine-readable steps at claim.json, guide at /docs/claim.
Everything here was measured by our prober or read from a registry. Nothing is self-reported.
Attributed to the source that supplied each field. Treated as claims, not facts.
# CodeDox - AI-Powered Documentation Search & Code Extraction **Transform any documentation site into a searchable code database** - CodeDox crawls documentation websites, intelligently extracts code snippets with context, and provides lightning-fast search via PostgreSQL full-text search and MCP (Model Context Protocol) integration for AI assistants. ## 📚 Documentation For full documentation, installation guides, API reference, and more, visit: ### **[https://chriswritescode-dev.github.io/codedox/](https://chriswritescode-dev.github.io/codedox/)** ## Quick Start ### Docker Setup (Recommended) ```bash # Clone the repository git clone https://github.com/chriswritescode-dev/codedox.git cd codedox # Configure environment cp .env.example .env # Edit .env to add your CODE_LLM_API_KEY (optional for AI-enhanced extraction) # Run the automated setup ./docker-setup.sh # Access the web UI at http://localhost:5173 # MCP tools available at http://localhost:8000/mcp ``` ### Manual Installation See the [full installation guide](https://chriswritescode-dev.github.io/codedox/getting-started/installation/) for detailed instructions. ## Key Features - **Intelligent Web Crawling**: Depth-controlled crawling with URL pattern filtering and domain restrictions - **Smart Code Extraction**: Dual-mode extraction (Automatic Title / Description or LLM Generated Titles and Descriptions) - **Enhanced Search Modes**: Standard code search with intelligent markdown fallback for comprehensive results - **Lightning-Fast Search**: PostgreSQL full-text search with fuzzy matching - **GitHub Repository Processing**: Clone and extract documentation from GitHub repositories with full path support (e.g., `/tree/main/docs`) - **HTTP-First MCP Integration**: MCP tools via HTTP endpoints with Streamable HTTP transport support (MCP 2025-03-26 spec) - **Full Documentation Access**: Get complete markdown content from any documentation page for full context - **Modern Web Dashboard**: React + Type…
Mapped onto the structured taxonomy from declared text and observed tool names. Confidence shown for derived entries.
Every source is kept verbatim. Field changes are logged as events.