Media intelligence analysis for audio, video, and images via the Echosaw MCP server.
Describe, classify and OCR images.
Media intelligence analysis for audio, video, and images via the Echosaw MCP server.
Image header probing, bbox conversion, resize plans and colour maths. 4 of 6 free.
63 pay-per-call tools for agents: vision, text, data, web, blockchain. USDC on Base via x402.
Web content extraction API for AI agents. Pay per call with x402 (USDC on Base) - no account, no API key.
AI-powered codebase analysis — call graphs, security, dead code, complexity. 150+ tools.
Understand your videos with Reka AI — search, ask questions, and extract insights.
Local OCR and on-device screen reading for MCP clients (Claude, Cursor): read any window, locate elements to click, assert what's on screen. Apple Vision, no API keys, nothing uploaded.
Persistent visual cache for LLM-driven software development. Caches screenshots using perceptual hashing, vector search, and AX trees to prevent token overhead and visual hallucination loops.
MCP server for Gemini 3.1 Flash with vision, chat, and image generation capabilities
Vision analysis CLI + MCP server backed by Seed 2.0 via Volcano Ark or any OpenAI-compatible endpoint
A TypeScript MCP server that gives text-only LLMs image understanding through StepFun vision models.
MCP server that gives Claude Code the ability to watch and understand videos — extracts frames via ffmpeg and processes audio via multiple backends
MCP server for Dynamsoft SDKs - Capture Vision, Barcode Reader (Mobile/Python/Web), Dynamic Web TWAIN, and Document Viewer. Provides documentation, code snippets, and API guidance.
MCP server for delegating executor-only coding, vision, and command tasks to a Mimo-compatible model route.
MCP server for integrating Omnisearch with LLMs
MCP server for secure, bounded OpenAI-compatible vision analysis
Fast screenshot capture tool for web pages - optimized for Claude Vision API
VisionPower: a portable image understanding MCP server for Codex, Claude Desktop, Cursor, and other agents.
MCP server exposing Mistral AI capabilities over MCP: chat, embeddings, FIM, vision, OCR, audio, agents, moderation, classification, files, batch, workflows, sampling, prompts, resources, and Streamable HTTP.
Multi-model vision MCP server (GLM-4.6V, Kimi, Qwen-VL, GPT-4o) that lets text-only coding models in Claude Code see and understand images.