create-mcp-eval
Generate comprehensive eval tests for any MCP server using @mcpjam/sdk. Supports Jest and…
Drive MCPJam's hosted eval tools end to end — check what a run will cost and disclose, launch it, poll it to a verdict, and triage a failure down to the step that failed. Use when connected to MCPJam's MCP server and asked to run, re-run, investigate, or compare eval results,
$ npx -y skills add MCPJam/inspector --skill run-mcpjam-evals --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/run-mcpjam-evalsContext preview
The summary Claude sees to decide when to auto-load this skill.
Drive MCPJam's hosted eval tools end to end — check what a run will cost and disclose, launch it, poll it to a verdict, and triage a failure down to the step that failed. Use when connected to MCPJam's MCP server and asked to run, re-run, investigate, or compare eval results,
name: run-mcpjam-evals description: Drive MCPJam's hosted eval tools end to end — check what a run will cost and disclose, launch it, poll it to a verdict, and triage a failure down to the step that failed. Use when connected to MCPJam's MCP server and asked to run, re-run, investigate, or compare eval results, rather than to author eval files locally.
You have MCPJam's hosted eval tools. This is the order to call them in, and the two places calling them wrong costs real money.
**This skill is about running evals that already exist.** To *write* eval files, use `create-mcp-eval` (SDK tests) or `mcpjam-eval-import` (convert an existing corpus). To author suites through these tools instead, see `references/authoring.md`.
| You are about to… | Read | |---|---| | Work out why a run did not pass | `references/triage.md` | | Create a suite or add cases through the tools | `references/authoring.md` |
list_eval_suites → get_eval_run_disclosure → run_eval_suite → get_eval_run (poll)
↓
list_eval_run_iterations → get_eval_run_steps
↓
compare_eval_run**1. Orient.** `list_eval_suites` returns suites with latest-run summaries and pass-rate trends. With no `project` it uses the most recently updated accessible one — fine for a quick look, but pass `project` explicitly the moment you are about to change or spend anything, or you may act on a project the user did not mean.
**2. Disclose before you spend.** `get_eval_run_disclosure` answers what a run does *before* it happens: which models it calls and where they route, which judges can fire, and what is captured or retained. Call it when a user has not run this suite before, when the suite has changed, or whenever they ask what a run will do. It is read-only and free.
**3. Launch.** `run_eval_suite` is **asynchronous** — it returns a `runId` immediately and the work continues server-side.
**4. Poll.** `get_eval_run` with `project` and `runId` until `status` is terminal: `completed`, `failed`, or `cancelled`. Anything else means still running; wait and ask again. Do not re-launch because a result is not ready.
**5. Read the verdict.** On anything other than a pass, **start at `get_eval_run`'s `decisionSummary`** — it carries the verdict and `verdictSource`, which tells you what actually decided the outcome. Do not jump straight to iterations; you will read a lot of rows without knowing what you are looking for. `references/triage.md` is the decision tree from here.
**6. Compare, when there is a baseline.** `compare_eval_run` with `baseRunId` (or `baseCommitSha`) classifies each case as `regressed`, `fixed`, `new_case`, `removed_case`, `changed`, `unchanged_passed`, or `unchanged_failed`. This is the tool that answers "did my change break anything", which a single run's pass rate cannot.
Two tools spend the organization's model budget. Both are marked `COSTS MONEY` in their descriptions, and neither is reversible once started.
Rules:
`get_eval_run` is free, but it is not free of latency, and a tight loop reads as a hung agent.
Open the hosted app. No install needed. 👉 app.mcpjam.com ... or run MCPJam locally for HTTP/S and local STDIO servers:
Repo: MCPJam/inspector
Generate comprehensive eval tests for any MCP server using @mcpjam/sdk. Supports Jest and…
Convert MCPJam Explore-generated test cases into @mcpjam/sdk eval tests. Produces one test…
Drive an API Playground session through MCPJam's remote MCP tools or cloud CLI, inspect…
Interpret and use `mcpjam` probe, doctor, OAuth, XAA (Cross-App Access / ID-JAG), apps…
Convert an existing test corpus (promptfoo YAML, pytest, Jest, CSV) into an MCPJam eval suite…
Defines every member of MCPJam's user-value chain vocabulary — the six stages, the five stage…