Video walkthrough: https://youtu.be/T6xS-dE9Lwc 60-second overview: https://youtu.be/FScnoCoHuQw
Benchmark any MCP server with rubric-driven tests, LLM-as-judge scoring, and a live web dashboard showing tool-call traces and regressions.

MCPGauge is a self-hostable eval harness for MCP (Model Context Protocol) servers. Point it at any server reachable over stdio or SSE, write a YAML suite of scenarios, and it runs a Claude agent against the server’s tools, scores each run with an LLM-as-judge against your rubric, and stores the results in a local SQLite database — no external infra, one file, your own API key.
The dashboard lets you drill into every tool call in an agent trace, compare two runs side-by-side to catch regressions before they ship, and watch live case results stream in as a suite executes. A bundled security suite checks whether prompt-injection payloads embedded in tool outputs can hijack the agent — coverage the MCP ecosystem currently has no standard tooling for.
git clone https://github.com/RitikPatill/mcpgauge.git
cd mcpgauge
make install # uv sync
export ANTHROPIC_API_KEY=sk-ant-...
make demo # run the poisoning suite end-to-end
make serve # open dashboard at http://localhost:8000/runs
Requires Python 3.11+, uv, and Node.js/npx for the filesystem and sqlite example suites.
Run a YAML suite from the CLI:
mcpgauge run examples/filesystem_basic.yaml
# Running suite: filesystem_basic (2 cases)
# [pass] list_tmp
# [fail] read_missing_file
# Run 3f8a2...: 1/2 passed
Compare two runs to detect regressions:
mcpgauge diff 3f8a2... 9c1b4...
# Case A B Change
# list_tmp pass pass unchanged
# read_missing_file fail pass fixed
Open the dashboard with mcpgauge serve — paste a stdio command or SSE URL to connect a server, click Run Suite to trigger a YAML suite, then click any case to see the full agent trace: tool arguments, raw results, timing, and the judge’s per-criterion verdict with reasoning. The Compare widget lets you pick a second run from a dropdown to view a highlighted diff.
CLI (typer) Web UI (HTMX + Tailwind)
| |
+----------+-------------+
|
runner.py ── SSE stream ──> Web UI
Agent loop + orchestrator
/ | \
client.py judge.py store.py
MCP client LLM rubric SQLite
stdio / SSE scorer (SQLModel)
| |
Target MCP Anthropic API
server (claude-haiku-4-5)
src/mcpgauge/ core package — client, runner, judge, store, diff, FastAPI app, CLI
examples/ three ready-to-run YAML suites: filesystem, sqlite, poisoning
tests/ unit and integration tests (pytest + anyio)
docs/ design notes, demo GIF, screenshot
mcpgauge export <run_id> producing a self-contained file for sharing without the dashboardasyncio.gather for faster suitespytest-mcpgauge so existing pytest users can write .py test files that invoke MCPGauge casesMIT — see LICENSE.
Built autonomously by autodev, a multi-agent orchestrator I designed. Each commit in this repo was authored by me; the implementation work was performed by Sonnet under the orchestrator’s control. Read the orchestrator’s README to see how.