Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.
Evaluate model or agent outputs.
Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.
Investigate fraud from Claude, Cursor or any MCP client. Explain any verdict with the evidence behind it, pivot from one signup to every account sharing its device, IP or inbox, check an entity against a cross-operator abuse network, and work the review q
Eval-backed tool discovery for AI agents on the auxiliar.ai web-access gateway. recommend_tools picks the best search, scraping, browser-automation or voice provider for a job from measured benchmarks (quality, latency, cost, error rate), returning routes
Model Context Protocol server for ZeroBounce email validation, verification, and intelligence. Provides comprehensive email validation, AI scoring, email finder, domain search, and bulk processing capabilities for AI coding assistants.
MCP server for LLM/VLM model selection — compare 300+ models with real-time benchmarks, pricing, and personalized recommendations. No API key required.
AINative ZeroDB Memory MCP Server - 18 tools for agent memory: 9 memory + 5 write-back (Slack, Gmail, Calendar, GitHub, Notion) + 4 plan artifacts (persistent plans/PRDs across sessions with diff history). Auto-context middleware. Production-ready for Cla
Multi-model consensus: 2-6 frontier LLMs answer, an independent judge synthesises one answer.
MCP server that exposes the Langfuse REST API as tools — query traces, observations, sessions, scores, prompts, datasets, and metrics from any MCP client.
Stop shipping agents on vibes. Score every agent output for quality, safety, and cost.
Guarded CLI and MCP client for JustHandled's x402 utility gateway
Autonomous MCP server for Omniology — real-USDC AI agent skill contests on Solana mainnet. A live benchmark against real agents, every 88 seconds, 24/7.
Agent-native MCP server for TokRepo - discover, verify, plan, safely install, hand off, and push AI assets from MCP clients.
MCP server exposing paid (x402) and free tools for AI agents — content provenance, quality scoring, and action audit — each backed by Ed25519-signed receipts on Cloudflare Workers.
MCP server with 22 tools for marketplace search, price comparison, license verification, childcare costs, estate sales, storage pricing, GSA auctions, home services, FCC lookup, and AI skill safety — TCGPlayer, Reverb, Grailed, Redfin, Poshmark, Craigslis
Parallel video rendering with live dashboard, GPU auto-detection, checkpoint system, and stream-copy concat. Includes an MCP server, a Claude Code skill, and a CLI.
MCP server for StressZero Intelligence API — burnout prevention self-assessment (non-medical) for AI agents, coaches, and HR platforms
MCP quality scanner and tool discovery for autonomous agents. Search 27,843+ indexed tools, check AEO (answer-engine-optimization) readiness scores, validate MCP server quality, find alternatives — the standard tool-validation layer before any agent calls
Word search generator, crossword generator, and sudoku generator + solver in one local-first CLI — printable PDF puzzle worksheets, themed word banks, verifiable LLM evals, and bundled agent skills.
DecisionMatrix MCP — deterministic multi-criteria decision analysis for AI agents: score, rank & explain options with weighted-sum, weighted-product, or TOPSIS. Runs over stdio (npx) or as a remote server.
A local-first workbench for developing, inspecting, and benchmarking MCP servers against local (LM Studio, Ollama) or remote (OpenRouter) models — Web UI, CLI, and MCP interface over one shared store.