mcp-eval-studio

MCP Eval Studio

Video walkthrough: https://youtu.be/Wv9Vbx1k2NA 60-second overview: https://youtu.be/zP4OpZfykcY

A web-based eval harness that points a Claude agent at any MCP server, runs rubric-based test scenarios, captures tool-call traces, and scores with LLM-as-judge.

demo

What it is

MCP Eval Studio is a FastAPI backend and Streamlit dashboard for testing Model Context Protocol servers the way a QA engineer tests an API. You point it at an MCP server (stdio or SSE), write a scenario in plain English, attach a rubric of weighted checks — must_call, must_not_call, output_regex, judge_criterion — and hit Run. A Claude agent executes the scenario against the live server while every tool call, argument, result, token count, and latency is recorded.

After the run an LLM-as-judge scores each rubric criterion and writes a per-check verdict with rationale to SQLite. Results appear in a dashboard where you can inspect the trace waterfall, read the scorecard, diff two runs side by side, and export a Markdown or JSON report. A prompt-injection canary preset surfaces cases where tool outputs embed instruction-override text — the attack vector documented in the “Poison everywhere” research.

Quickstart

git clone https://github.com/RitikPatill/mcp-eval-studio.git
cd mcp-eval-studio

# Install package + dev dependencies
pip install -e ".[dev]"

# Provide your Anthropic key
export ANTHROPIC_API_KEY=sk-ant-...

# Run the test suite (integration tests auto-skip if npx is absent)
pytest -v

# Terminal 1 — API server
uvicorn mcp_eval_studio.api:app --reload

# Terminal 2 — Streamlit dashboard
streamlit run mcp_eval_studio/app.py

# Optional: run the starter scenario pack headlessly
mcp-eval run scenarios/starter_pack.yaml
mcp-eval run scenarios/starter_pack.yaml --output results.json

To boot the full demo with a reference filesystem MCP server in one command:

bash record_demo.sh

Requirements for the demo script: node, python, uvicorn, and mcp-eval on PATH. Results land in /tmp/mcp-demo/results.json; the dashboard is at http://localhost:8501.

Usage

Open the dashboard at http://localhost:8501. On the Servers page, add an MCP server by entering its stdio command (e.g. npx -y @modelcontextprotocol/server-filesystem /tmp) or an SSE URL. The tool, resource, and prompt list auto-populates on connect.

On the Scenarios page, load scenarios/starter_pack.yaml or write a new scenario: a natural-language prompt and a rubric YAML block. Hit Run All to execute every scenario in sequence. Open any run to see the tool-call waterfall, rubric scorecard with per-check pass/fail and judge rationale, and token/cost totals. Use the Compare toggle to diff the current run against a previous one. Export buttons on the Run Detail page produce JSON or Markdown reports.

From the CLI, mcp-eval run <pack.yaml> exits 0 when all checks pass and 1 on any failure — suitable for CI.

Architecture

flowchart LR
    UI[Streamlit UI] -->|REST| API[FastAPI]
    API --> Runner[Agent Runner]
    API --> Store[(SQLite)]
    Runner --> Claude[Anthropic API]
    Runner --> MCP[MCP Client]
    MCP -->|stdio/SSE| Server[Target MCP Server]
    Runner --> Tracer[Trace Collector]
    Tracer --> Store
    API --> Judge[LLM Judge]
    Judge --> Claude
    Judge --> Store

The Agent Runner is a tool-use loop: send messages to Claude, forward any tool_use blocks to the MCP client, feed results back as tool_result blocks, repeat until stop or max_iterations. Every hop is written to the Trace Collector with monotonic timestamps and accumulated token costs. The Judge is a second Claude call whose system prompt contains the rubric and whose input is the serialized trace; it returns structured JSON per-check verdicts via tool-use for structured output.

Project structure

mcp-eval-studio/
├── mcp_eval_studio/
│   ├── app.py          # Streamlit dashboard — waterfall, scorecard, diff, export
│   ├── api.py          # FastAPI REST backend
│   ├── cli.py          # `mcp-eval run <pack.yaml>` entry point
│   ├── db.py           # SQLite persistence
│   ├── report.py       # Markdown report renderer
│   ├── client/         # MCP client wrapper (stdio + SSE)
│   ├── runner/         # Agent runner + trace models
│   ├── rubric/         # Rubric schema, Pydantic models, deterministic evaluators
│   └── judge/          # LLM-as-judge scorer + injection canary preset
├── scenarios/
│   └── starter_pack.yaml   # 3 scenarios: happy-path, refusal-path, injection-canary
├── tests/              # 9 test modules, ~60 tests, integration tests auto-skip
├── docs/               # screenshot.png and demo.gif
├── record_demo.sh      # end-to-end demo script
├── pyproject.toml      # dependencies, entry points, ruff + pytest config
└── Makefile            # dev / test / lint / typecheck / run targets

Roadmap

License

MIT — see LICENSE.


Built autonomously by autodev, a multi-agent orchestrator I designed. Each commit in this repo was authored by me; the implementation work was performed by Sonnet under the orchestrator’s control. Read the orchestrator’s README to see how.