{"$schema":"https://wellknown.network/schemas/agent-record-v1.json","schemaVersion":"1","id":"ag_vhhx674mhhze","handle":"pdf-indexer-mcp","url":"https://wellknown.network/agents/pdf-indexer-mcp","links":{"self":"https://wellknown.network/agents/pdf-indexer-mcp/record.json","html":"https://wellknown.network/agents/pdf-indexer-mcp","markdown":"https://wellknown.network/agents/pdf-indexer-mcp/record.md","api":"https://wellknown.network/api/v1/agents/pdf-indexer-mcp","status":"https://wellknown.network/api/v1/agents/pdf-indexer-mcp/status","claim":"https://wellknown.network/agents/pdf-indexer-mcp/claim","claimApi":"https://wellknown.network/api/v1/claims","claimDescriptor":"https://wellknown.network/agents/pdf-indexer-mcp/claim.json","badge":"https://wellknown.network/agents/pdf-indexer-mcp/badge.svg","openapi":"https://wellknown.network/openapi.json"},"ard":{"identifier":"urn:air::server:pdf-indexer-mcp","type":"application/mcp-server-card+json"},"kind":"mcp_server","declared":{"name":"pdf-indexer-mcp","summary":"A Model Context Protocol (MCP) server for indexing and semantically searching PDF research papers","description":"# PDF Indexer MCP Server\n\nA **Model Context Protocol (MCP) server** that enables AI agents to download, index, and semantically search PDF research papers. This server provides 8 tools that AI agents can discover and use autonomously to build research paper knowledge bases and answer questions.\n\n## What is MCP?\n\n**Model Context Protocol (MCP)** is a standardized protocol that allows AI agents to discover and use tools. Instead of being limited to text generation, agents become **action-capable systems** that can:\n\n1. **Discover Tools**: Agents automatically discover available tools from connected MCP servers\n2. **Understand Capabilities**: Agents read tool descriptions and parameters to understand what each tool can do\n3. **Execute Tasks**: Agents call tools with appropriate parameters to accomplish goals\n4. **Compose Workflows**: Agents can combine multiple tools from different servers to solve complex problems\n\n### How MCP Works\n\n```\nAI Agent → MCP Protocol → Tool Server → Execution → Results → Agent\n```\n\nWhen you connect this MCP server to an AI agent (like in Cursor, Claude Desktop, or via OpenAI Agents framework), the agent automatically:\n\n- Discovers all 8 tools available from this server\n- Understands what each tool does from their descriptions\n- Uses the tools when they're needed to complete tasks\n- Can combine tools in complex workflows\n\n## Features\n\n- 📥 **PDF Download**: Download research papers from URLs\n- 📄 **Intelligent Chunking**: Two chunking strategies:\n  - **Header-based**: Preserves document structure (ideal for academic papers)\n  - **S2 chunking**: Spatial-semantic hybrid approach for optimal semantic chunks\n- 🗄️ **Database Indexing**: Store papers and chunks in SQLite with navigation indices\n- 🔍 **Semantic Search**: Search papers using MLX-optimized embeddings (Qwen3-Embedding-0.6B)\n- ⚡ **FAISS Vector Index**: Fast similarity search even for thousands of chunks\n- 🧠 **Context-Aware**: Retrieve surrounding chunks for better context\n\n## Quick …","publisher":null,"homepage":"https://github.com/lizTheDeveloper/pdf-indexer-mcp#readme","repository":"https://github.com/lizTheDeveloper/pdf-indexer-mcp#readme","version":"1.0.0","license":"GPL-3.0","protocols":["mcp"],"tags":["mcp","model-context-protocol","pdf","research-papers","semantic-search","rag","embeddings","faiss","mlx"],"pricing":null,"endpoints":[{"url":"pypi:pdf-indexer-mcp","type":"package_pypi","auth":null,"probeable":false}],"skills":null,"tools":null,"extra":null,"attribution":{"kind":"pypi","name":"pypi","license":"pypi","repoUrl":"pypi","summary":"pypi","version":"pypi","description":"pypi","homepageUrl":"pypi"}},"derived":{"capabilities":[{"slug":"data.vector-search","name":"Vector Search","confidence":1,"provenance":"declared"},{"slug":"research.academic","name":"Academic Research","confidence":1,"provenance":"derived"},{"slug":"data.database","name":"Databases","confidence":0.802,"provenance":"derived"}],"categories":["data","research"],"language":"en"},"observed":{"status":"unknown","statusReason":"Distributed as a package to run locally; no network endpoint to check.","lastOkAt":null,"lastProbedAt":null,"statusComputedAt":null,"reliability30d":null,"latestObservations":[],"tools":null,"package":{"name":"pdf-indexer-mcp","registry":"pypi","observedAt":"2026-09-10T09:26:57.132Z","publishedAt":"2025-11-02T23:37:09.139220Z","latestVersion":"1.0.0"}},"verification":{"claimed":false,"claimedAt":null,"proofs":[]},"provenance":{"sources":[{"source":"pypi","key":"pdf-indexer-mcp","url":"https://pypi.org/project/pdf-indexer-mcp/","firstSeenAt":"2026-09-10T09:25:50.879Z","fetchedAt":"2026-09-10T09:25:50.879Z","normalizedAt":"2026-09-10T09:25:50.879Z"}]},"firstSeenAt":"2026-09-10T09:25:50.879Z","updatedAt":"2026-09-10T09:26:57.132Z"}