explore-to-sdk-evals
Convert MCPJam Explore-generated test cases into @mcpjam/sdk eval tests. Produces one test…
Generate comprehensive eval tests for any MCP server using @mcpjam/sdk. Supports Jest and Vitest with deterministic and LLM-driven test patterns.
$ npx -y skills add MCPJam/inspector --skill create-mcp-eval --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/create-mcp-evalContext preview
The summary Claude sees to decide when to auto-load this skill.
Generate comprehensive eval tests for any MCP server using @mcpjam/sdk. Supports Jest and Vitest with deterministic and LLM-driven test patterns.
name: create-mcp-eval description: Generate comprehensive eval tests for any MCP server using @mcpjam/sdk. Supports Jest and Vitest with deterministic and LLM-driven test patterns.
Generate eval tests for MCP servers using **@mcpjam/sdk**.
Read this file first. It carries the two things you need before writing any code — what to ask the user, and the rules the generated tests must follow — and routes you to the rest only when you actually need it.
Load a reference when you reach the step that needs it, not before.
| You are about to… | Read | |---|---| | Scaffold `package.json`, `tsconfig.json`, `.env.example`, `.gitignore` | `references/project-setup.md` | | Call `MCPClientManager`, `HostRunner`, `PromptResult`, `EvalTest`, `EvalSuite`, validators, or the MCPJam reporter | `references/sdk-api.md` | | Choose a shape — config block, toggled suites, shared reporter, parameterized agents, save modes, multi-turn, validator coverage | `references/patterns.md` | | Write the file out | `references/template.md` | | Debug a test that runs but behaves oddly | `references/common-mistakes.md` | | Turn an MCPJam **Agent Brief** into tests | `references/agent-brief.md` |
Before generating any code, collect the following from the user:
| Question | Options | Default | |----------|---------|---------| | **Connection type** | `stdio` (local binary) or `http` (SSE/Streamable HTTP URL) | `http` | | **Test framework** | `jest`, `vitest`, or `none` (SDK-only) | _(detect from repo; fall back to `vitest`)_ | | **LLM provider** | See Supported Providers table below. Format: `provider/model` | _(must ask user)_ | | **Save results to MCPJam** | `none`, `auto` (saves when MCPJAM_API_KEY is set), or `reporter` (shared EvalRunReporter). Use an MCPJam API key (`sk_…`) from **Settings → API keys**; optionally set `MCPJAM_PROJECT_ID` to file results under a specific project (defaults to the org’s Default project). | _(must ask user)_ | | **Tool list** | Ask user to paste their tool names or an **Agent Brief** (`references/agent-brief.md`) | — |
If the user provides an **Agent Brief** (markdown with `## Tools` table), parse it to auto-populate tool names, descriptions, parameters, and suggested eval scenarios. See `references/agent-brief.md`.
You MUST ask the developer which LLM provider they want before generating any code. Do not default to any provider.
**Supported Providers:**
| Provider | Model format | Env var | Example model | |----------|-------------|---------|---------------| | `openai` | `openai/<model>` | `OPENAI_API_KEY` | `openai/gpt-4o-mini` | | `anthropic` | `anthropic/<model>` | `ANTHROPIC_API_KEY` | `anthropic/claude-sonnet-4-20250514` | | `google` | `google/<model>` | `GOOGLE_API_KEY` | `google/gemini-2.0-flash` | | `mistral` | `mistral/<model>` | `MISTRAL_API_KEY` | `mistral/mistral-small-latest` | | `deepseek` | `deepseek/<model>` | `DEEPSEEK_API_KEY` | `deepseek/deepseek-chat` | | `xai` | `xai/<model>` | `XAI_API_KEY` | `xai/grok-2` | | `openrouter` | `openrouter/<model>` | `OPENROUTER_API_KEY` | `openrouter/openai/gpt-4o-mini` | | `azure` | `azure/<deployment>` | `AZURE_API_KEY` | `azure/gpt-4o` | | `ollama` | `ollama/<model>` | _(none, local)_ | `ollama/llama3` | | Custom | `<name>/<model>` | _(configurable)_ | `litellm/gpt-4` |
Once the user selects a provider, use the corresponding env var name and model format in all generated code:
Before generating tests, check what the codebase already uses:
Then:
In all cases, use `@mcpjam/sdk` for the eval harness (`HostRunner`, `EvalTest`, `EvalSuite`, validators).
---
Follow these rules when generating eval test files:
1. **Deterministic suite first** — always include a deterministic test section using `HostRunner.mock()` that validates the test structure itself without requiring LLM calls or server connections.
2. **One EvalTest per tool** — create a separate `EvalTest` for each tool you want to evaluate. Each test should prompt the runner with a natural-language request and assert the correct tool was selected.
3. **Single-shot LLM tests are non-deterministic** — a single `runner.run()` may not select the expected tool every time. For single-shot tests, prefer saving results to MCPJam without hard-asserting (`expect(...).toBe(true)`). Use `EvalTest` with `iterations >= 3` and assert on `accuracy()` for reliable pass/fail gates. Reserve hard asserts for high-confidence cases (negative tests, multi-turn with clear context).
4. **Write unambiguous prompts for similar tools** — when a server has tools with overlapping descriptions (e.g., `create_view` vs `export_to_excalidraw`), prompts must reference the tool's *unique* action. Mention specific verbs, targets, or outcomes. Bad: "Share my diagram". Good: "Export and upload my diagram to excalidraw.com so I can open it in a browser".
5. **Multi-turn for related tools** — when tools logically chain together (e.g., `get_user` then `list_workspaces`), create a multi-turn test using `{ context: previousResult }`.
6. **Negative test** — always include at least one test that verifies the runner does NOT call tools when given an irrelevant prompt (e.g., "What is the capital of France?"). Use `matchNoToolCalls()`.
7. **Reasonable defaul
Open the hosted app. No install needed. 👉 app.mcpjam.com ... or run MCPJam locally for HTTP/S and local STDIO servers:
Repo: MCPJam/inspector
Convert MCPJam Explore-generated test cases into @mcpjam/sdk eval tests. Produces one test…
Drive an API Playground session through MCPJam's remote MCP tools or cloud CLI, inspect…
Interpret and use `mcpjam` probe, doctor, OAuth, XAA (Cross-App Access / ID-JAG), apps…
Convert an existing test corpus (promptfoo YAML, pytest, Jest, CSV) into an MCPJam eval suite…
Drive MCPJam's hosted eval tools end to end — check what a run will cost and disclose, launch…
Defines every member of MCPJam's user-value chain vocabulary — the six stages, the five stage…