Video walkthrough: https://youtu.be/Wv9Vbx1k2NA 60-second overview: https://youtu.be/zP4OpZfykcY
A web-based eval harness that points a Claude agent at any MCP server, runs rubric-based test scenarios, captures tool-call traces, and scores with LLM-as-judge.

MCP Eval Studio is a FastAPI backend and Streamlit dashboard for testing Model Context Protocol servers the way a QA engineer tests an API. You point it at an MCP server (stdio or SSE), write a scenario in plain English, attach a rubric of weighted checks — must_call, must_not_call, output_regex, judge_criterion — and hit Run. A Claude agent executes the scenario against the live server while every tool call, argument, result, token count, and latency is recorded.
After the run an LLM-as-judge scores each rubric criterion and writes a per-check verdict with rationale to SQLite. Results appear in a dashboard where you can inspect the trace waterfall, read the scorecard, diff two runs side by side, and export a Markdown or JSON report. A prompt-injection canary preset surfaces cases where tool outputs embed instruction-override text — the attack vector documented in the “Poison everywhere” research.
git clone https://github.com/RitikPatill/mcp-eval-studio.git
cd mcp-eval-studio
# Install package + dev dependencies
pip install -e ".[dev]"
# Provide your Anthropic key
export ANTHROPIC_API_KEY=sk-ant-...
# Run the test suite (integration tests auto-skip if npx is absent)
pytest -v
# Terminal 1 — API server
uvicorn mcp_eval_studio.api:app --reload
# Terminal 2 — Streamlit dashboard
streamlit run mcp_eval_studio/app.py
# Optional: run the starter scenario pack headlessly
mcp-eval run scenarios/starter_pack.yaml
mcp-eval run scenarios/starter_pack.yaml --output results.json
To boot the full demo with a reference filesystem MCP server in one command:
bash record_demo.sh
Requirements for the demo script: node, python, uvicorn, and mcp-eval on PATH. Results land in /tmp/mcp-demo/results.json; the dashboard is at http://localhost:8501.
Open the dashboard at http://localhost:8501. On the Servers page, add an MCP server by entering its stdio command (e.g. npx -y @modelcontextprotocol/server-filesystem /tmp) or an SSE URL. The tool, resource, and prompt list auto-populates on connect.
On the Scenarios page, load scenarios/starter_pack.yaml or write a new scenario: a natural-language prompt and a rubric YAML block. Hit Run All to execute every scenario in sequence. Open any run to see the tool-call waterfall, rubric scorecard with per-check pass/fail and judge rationale, and token/cost totals. Use the Compare toggle to diff the current run against a previous one. Export buttons on the Run Detail page produce JSON or Markdown reports.
From the CLI, mcp-eval run <pack.yaml> exits 0 when all checks pass and 1 on any failure — suitable for CI.
flowchart LR
UI[Streamlit UI] -->|REST| API[FastAPI]
API --> Runner[Agent Runner]
API --> Store[(SQLite)]
Runner --> Claude[Anthropic API]
Runner --> MCP[MCP Client]
MCP -->|stdio/SSE| Server[Target MCP Server]
Runner --> Tracer[Trace Collector]
Tracer --> Store
API --> Judge[LLM Judge]
Judge --> Claude
Judge --> Store
The Agent Runner is a tool-use loop: send messages to Claude, forward any tool_use blocks to the MCP client, feed results back as tool_result blocks, repeat until stop or max_iterations. Every hop is written to the Trace Collector with monotonic timestamps and accumulated token costs. The Judge is a second Claude call whose system prompt contains the rubric and whose input is the serialized trace; it returns structured JSON per-check verdicts via tool-use for structured output.
mcp-eval-studio/
├── mcp_eval_studio/
│ ├── app.py # Streamlit dashboard — waterfall, scorecard, diff, export
│ ├── api.py # FastAPI REST backend
│ ├── cli.py # `mcp-eval run <pack.yaml>` entry point
│ ├── db.py # SQLite persistence
│ ├── report.py # Markdown report renderer
│ ├── client/ # MCP client wrapper (stdio + SSE)
│ ├── runner/ # Agent runner + trace models
│ ├── rubric/ # Rubric schema, Pydantic models, deterministic evaluators
│ └── judge/ # LLM-as-judge scorer + injection canary preset
├── scenarios/
│ └── starter_pack.yaml # 3 scenarios: happy-path, refusal-path, injection-canary
├── tests/ # 9 test modules, ~60 tests, integration tests auto-skip
├── docs/ # screenshot.png and demo.gif
├── record_demo.sh # end-to-end demo script
├── pyproject.toml # dependencies, entry points, ruff + pytest config
└── Makefile # dev / test / lint / typecheck / run targets
uses: RitikPatill/mcp-eval-studio@v1 that runs mcp-eval run in CI and posts the rubric scorecard as a PR commentMIT — see LICENSE.
Built autonomously by autodev, a multi-agent orchestrator I designed. Each commit in this repo was authored by me; the implementation work was performed by Sonnet under the orchestrator’s control. Read the orchestrator’s README to see how.