wellknown
SearchCapabilitiesReliabilityJournalSourcesDocs
/mcpAdd to your agentSign in
wellknown

A live, machine-queryable index of AI agents, MCP servers and tools. Every record separates what publishers declare from what we observed.

Product
  • Search
  • Capabilities
  • Reliability report
  • Journal
  • Sources & freshness
  • Network status
Developers
  • Documentation
  • HTTP API
  • MCP server
  • Claim an agent
  • Dashboard
For machines
  • llms.txt
  • openapi.json
  • agent-card.json
  • sitemap.xml
  • About our crawler
Records are indexed from public sources and attributed to them. Observed data is ours; declared data is theirs.v0.1 · early network
Capabilities/AI Services/ai.evaluation

Evaluation & Benchmarks

Evaluate model or agent outputs.

134 records56 liveRefine in search →
@kalisky/skarMCP serverMCP
by kalisky
Unknown

Skar turns a captured AI agent trace into a committed pytest regression test. MCP server + CLI. Use when a tool-using agent run fails and you want to lock the failure as an executable test.

Testing~66%Evaluation & Benchmarks~62%
npm:@kalisky/skar · local package9/wkpublished 4 mo ago
larkxMCP serverMCP
by sumit_lk_patel
Unknown

AI codebase indexer and MCP server for Claude Code, Cursor, and Copilot. Pre-index your project into a compact graph and measure real token savings with `larkx bench`.

Evaluation & BenchmarksKnowledge Graphs~100%Security Review~75%
npm:larkx · local package6/wkpublished 4 mo ago
vulcnMCP serverMCP
by open-cipher
Unknown

Security evals for the AI era. Probes · Targets · Graders · Proof. Confirmed XSS / SQLi / BOLA / prompt-injection / MCP-RCE with reproducible proof attached to every finding.

Evaluation & BenchmarksPrompt Management~100%Security Review~75%Security Scanning~66%
npm:vulcn · local package2/wkpublished 4 mo ago
AgentEinsteinAgentA2A
Unknown

Autonomous crypto and market intelligence from Agent Einstein. Skills cover whale and smart-money flow, token and contract security scanning, MEV detection, cross-chain and wallet analytics, prediction-market signals, DeFi yield, backtesting, and ML price forecasting (TimesFM + Kronos dual-model co…

Reasoning & PlanningVideo Generation~100%Contracts & LegalPaymentsMarkets & Trading~100%
https://emc2ai.io
Engineering Leaders CommunityAgentA2A
by Engineering Leaders Community
Unknown

Agent-callable services from the Engineering Leaders Community: meetup topic evaluation, speaker stage placement, and community reach — grounded in 3,300+ CEE engineering leaders and 12 meetups a year at 120+ attendees since 2019.

Evaluation & Benchmarks~100%MarketingPricing & QuotesProduct Analytics3D & Design~64%
https://www.engineeringleaders.io
UK Break Clause AnalyzerMCP serverMCP
by Ankur-stockheads
Unknown

Assess UK lease tenant break clauses with grounded citations and a measured hallucination rate.

Contracts & Legal~100%Evaluation & Benchmarks~71%
pypi:uk-break-clause-analyzer-mcp · local packagepublished 2 mo ago
revettr-mcpMCP serverMCP
by AlexanderLawson17
Unknown

Counterparty risk scoring for agentic commerce via x402 micropayments.

PaymentsEvaluation & Benchmarks~100%
https://revettr.com/mcppublished 5 mo ago
HlidoAgentA2AMCP
by Hlido s.r.o.
Unknown

The independent trust layer for AI agents — each agent hands-on tested for one evidence-backed Laddoo Score (0-100). Query trust checks, scorecards, incidents, and recommendations over MCP.

Evaluation & BenchmarksDatabases~86%
https://hlido.eu/mcp
SirenicAgentA2A
by Sirenic
Unknown

Pay-per-call French & European company data (official registers: INSEE Sirene, INPI RNE, NBB, Zefix and 8 more, plus worldwide entities via LEI/GLEIF). Every skill is a paid HTTP resource: send a JSON data part {"path": "/v1/..."} or {"skill": "<id>", "params": {...}}; the agent replies with an x40…

Contracts & Legal~100%Monitoring & ObservabilityNotificationsPayments~77%3D & Design~64%
https://api.sirenic.eu/a2a
True Value RankingsAgentA2A
by True Value Rankings LLC
Unknown

Cryptocurrency research and scoring platform. TVR's scoring model evaluates cryptocurrencies using 8 fundamental metrics rooted in Sound Value principles to produce TVR Scores (0-100). TVR's valuation formula produces TVR Estimated Values (fair value estimates) and Valuation Categories (Undervalued…

Evaluation & BenchmarksCode DocumentationFilesystemAgent Discovery~69%
https://truevaluerankings.com
StratalizeAgentA2A
Unknown

Stratalize — attested finance, legal, healthcare, and compliance intelligence. Signed, independently verifiable receipt on every call (trust.stratalize.com/verify).

Evaluation & Benchmarks~69%Travel & Booking~64%Deep Web Research~54%
https://www.stratalize.com/api/a2a
DiezXSuiteAgentA2A
by diezX
Unknown

A Business Observability platform for mid-market companies: it shows how your company operates in real time and lets you act from there. diezX connects once to a company's existing systems (ERP, CRM, HR, Google Workspace) and builds a live Graph Rail of five operational domains: collaboration, sale…

3D & DesignCI/CD & Deploy~100%CRM & Sales~100%Workflow AutomationCode Review
https://agents.diezx.ai/api/a2a/suite
orq.aiAgentA2A
by orq.ai
Unknown

LLM operations and agent platform: model gateway/router, prompt management, evaluators, datasets, experiments, deployments, and tracing.

CI/CD & Deploy~100%Evaluation & BenchmarksPrompt Management
https://api.orq.ai/v3/router
The Aggregate — LLM benchmark aggregateMCP serverMCP
by ai.theaggregate
Unknown

Fused LLM rankings: one IRT/Elo scale across ~5,000 public benchmark leaderboards, updated daily.

Evaluation & Benchmarks~100%
https://theaggregate.ai/mcp
← Previouspage 7 of 7Next →
Machine query
GET /api/v1/search?capability=ai.evaluation&status=live
Aliases

evaluation, evals, benchmark, scoring, grading, judge

Also in AI Services
  • Model Access
  • Prompt Management
  • Agent Discovery