{"$schema":"https://wellknown.network/schemas/agent-record-v1.json","schemaVersion":"1","id":"ag_ma37hs6238sv","handle":"scrapedatshi-mcp","url":"https://wellknown.network/agents/scrapedatshi-mcp","links":{"self":"https://wellknown.network/agents/scrapedatshi-mcp/record.json","html":"https://wellknown.network/agents/scrapedatshi-mcp","markdown":"https://wellknown.network/agents/scrapedatshi-mcp/record.md","api":"https://wellknown.network/api/v1/agents/scrapedatshi-mcp","status":"https://wellknown.network/api/v1/agents/scrapedatshi-mcp/status","claim":"https://wellknown.network/agents/scrapedatshi-mcp/claim","claimApi":"https://wellknown.network/api/v1/claims","claimDescriptor":"https://wellknown.network/agents/scrapedatshi-mcp/claim.json","badge":"https://wellknown.network/agents/scrapedatshi-mcp/badge.svg","openapi":"https://wellknown.network/openapi.json"},"ard":{"identifier":"urn:air::server:scrapedatshi-mcp","type":"application/mcp-server-card+json"},"kind":"mcp_server","declared":{"name":"scrapedatshi-mcp","summary":"MCP server for the scrapedatshi RAG pipeline API — use scrapedatshi tools directly from Claude Desktop","description":"# scrapedatshi-mcp\n\nMCP (Model Context Protocol) server for the [scrapedatshi](https://scrapedatshi.com) RAG pipeline API.\n\nUse scrapedatshi's scraping, crawling, extraction, and vector DB sync tools directly from **Claude Desktop** — no code required.\n\n---\n\n## What you can do\n\nJust talk to Claude naturally:\n\n- *\"Scrape https://docs.example.com and give me the chunks\"*\n- *\"Extract the text from this PDF: https://example.com/annual-report.pdf\"*\n- *\"Extract all tables from this local PDF: C:/Users/me/Documents/financials.pdf\"*\n- *\"Chunk this PDF URL: https://my-bucket.s3.amazonaws.com/report.pdf\"* — PDF URLs are automatically detected and extracted\n- *\"Crawl https://example.com/products and extract the title and price from every page\"*\n- *\"Sync https://docs.example.com to my Pinecone index using OpenAI embeddings\"*\n- *\"Crawl the entire docs.stripe.com site (all 800 pages) and inject it into my Pinecone index\"* — large sites are auto-batched server-side, no manual pagination needed\n- *\"What embedding providers does scrapedatshi support?\"*\n- *\"Inspect my Pinecone index and tell me what embedding model was used\"*\n- *\"Query my Pinecone index for information about API authentication\"*\n- *\"Query my LanceDB with hybrid search — I need to find exact IDs and names, not just semantic matches\"*\n- *\"Chunk https://docs.example.com using hierarchical chunking so the LLM gets full context on retrieval\"*\n- *\"Ingest all the JSON files in my ./scrapy_output/ folder into my Pinecone index\"*\n\n---\n\n## Tools exposed\n\n| Tool | What it does |\n|---|---|\n| `verify_provider_key` | Verify an LLM or embedding API key + get live model list |\n| `get_usage_guide` | Returns the guided wizard flow and tool selection reference |\n| `scrape_url` | Scrape a URL and return clean Markdown — no chunking, just the raw text |\n| `pdf_extract` | Extract text or tables from a PDF (URL or local file) — no chunking, no embedding needed |\n| `chunk_url` | Scrape & chunk a single URL into RAG-ready text segments |\n| …","publisher":null,"homepage":"https://docs.scrapedatshi.com/sdk/mcp","repository":"https://github.com/scrapedatshi/scrapedatshi-mcp/issues","version":"0.6.11","license":"MIT","protocols":["mcp"],"tags":["ai","chunking","claude","embeddings","llm","mcp","model-context-protocol","rag","scraping","vector-database"],"pricing":null,"endpoints":[{"url":"pypi:scrapedatshi-mcp","type":"package_pypi","auth":null,"probeable":false}],"skills":null,"tools":null,"extra":null,"attribution":{"kind":"pypi","name":"pypi","license":"pypi","repoUrl":"pypi","summary":"pypi","version":"pypi","description":"pypi","homepageUrl":"pypi"}},"derived":{"capabilities":[{"slug":"data.web-scraping","name":"Web Scraping","confidence":1,"provenance":"declared"},{"slug":"data.database","name":"Databases","confidence":1,"provenance":"derived"},{"slug":"data.vector-search","name":"Vector Search","confidence":1,"provenance":"declared"},{"slug":"dev.docs-lookup","name":"Documentation Lookup","confidence":0.791,"provenance":"derived"}],"categories":["data","dev"],"language":"en"},"observed":{"status":"unknown","statusReason":"Distributed as a package to run locally; no network endpoint to check.","lastOkAt":null,"lastProbedAt":null,"statusComputedAt":null,"reliability30d":null,"latestObservations":[],"tools":null,"package":{"name":"scrapedatshi-mcp","registry":"pypi","observedAt":"2026-09-10T12:23:12.075Z","publishedAt":"2026-07-22T02:58:29.239568Z","latestVersion":"0.6.11"}},"verification":{"claimed":false,"claimedAt":null,"proofs":[]},"provenance":{"sources":[{"source":"pypi","key":"scrapedatshi-mcp","url":"https://pypi.org/project/scrapedatshi-mcp/","firstSeenAt":"2026-09-10T12:21:00.522Z","fetchedAt":"2026-09-10T12:21:00.522Z","normalizedAt":"2026-09-10T12:21:00.522Z"}]},"firstSeenAt":"2026-09-10T12:21:00.522Z","updatedAt":"2026-09-10T12:23:12.075Z"}