OSWorld-MCP: A comprehensive MCP server for computer-use agents with 158 validated tools
Wellknown found it in public sources; nobody has proven control of it yet. Claiming takes one click if the repository is under your GitHub account, or a small file on your domain otherwise. Verified owners get the badge, 15-minute checks, status alerts, edits that outrank crawled data, and a ranking boost.
Agents can do it too: POST https://wellknown.network/api/v1/claims with {"agent":"iflow-mcp-x-plug-osworld-mcp","method":"well_known_file"} — machine-readable steps at claim.json, guide at /docs/claim.
Everything here was measured by our prober or read from a registry. Nothing is self-reported.
Attributed to the source that supplied each field. Treated as claims, not facts.
# OSWorld-MCP: Benchmarking MCP Tool Invocation in Computer-Use Agents ## 🔔 Updates **2025-10-28:** We released our paper and project page! 🎉 📄 [Read the Paper](https://arxiv.org/abs/2510.24563) | 🌐 [Visit the Project Page](https://osworld-mcp.github.io) --- ## 📑 Overview & Key Highlights OSWorld-MCP is a comprehensive and fair benchmark for evaluating computer-use agents in real-world scenarios. It jointly measures **Model Context Protocol (MCP)** tool invocation capabilities, **graphical user interface (GUI)** operation skills, and **decision-making** performance. Designed as an extension of **OSWorld**, it significantly improves realism, balance, and comparability in evaluation. **Key Features & Findings** - **158 validated MCP tools**, spanning **7 common applications** (LibreOffice Writer, Calc, Impress, VS Code, Google Chrome, VLC, OS utilities). Among them, **25 distractor tools** for robustness testing - **250 tool-beneficial tasks** → 69% of benchmark tasks benefit from MCP tools - Multi-round tool invocation possible, posing real decision-making challenges - **MCP tools boost model accuracy & efficiency** — e.g., OpenAI o3: 8.3% → 20.4% (15 steps) - Highest observed Tool Invocation Rate (**TIR**) = 36.3% (Claude-4-Sonnet, 50 steps) → indicating ample room for improvement - MCP tools improve agent metrics - Higher tool invocation correlates with higher accuracy - Combining tools introduces significant challenges **Architecture Overview**  *Figure: OSWorld-MCP evaluation framework integrating GUI actions and MCP tool invocations.* --- ## ⚙️ Installation & Usage ### 1️⃣ Preparation: Code Setup ```bash # Clone OSWorld base repo git clone https://github.com/xlang-ai/OSWorld.git # Clone OSWorld-MCP git clone https://github.com/X-PLUG/OSWorld-MCP.git ``` Integrate **OSWorld-MCP** files into OSWorld to enable MCP support. --- ### 2️⃣ Preparation: Docker Environment 1. Copy…
Mapped onto the structured taxonomy from declared text and observed tool names. Confidence shown for derived entries.
Every source is kept verbatim. Field changes are logged as events.