/mcp-builder
Builds production MCP servers via 4-phase methodology: research, implement, test, evaluate. Triggers: build MCP, new MCP, MCP integration, MCP server scaffold.
$ npx -y skills add softspark/ai-toolkit --skill mcp-builder --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/mcp-builder
Context preview
The summary Claude sees to decide when to auto-load this skill.
Builds production MCP servers via 4-phase methodology: research, implement, test, evaluate. Triggers: build MCP, new MCP, MCP integration, MCP server scaffold.
SKILL.md
mcp-builder.SKILL.mdname: mcp-builder
description: "Builds production MCP servers via 4-phase methodology: research, implement, test, evaluate. Triggers: build MCP, new MCP, MCP integration, MCP server scaffold."
effort: high
disable-model-invocation: true
argument-hint: "[service name or API description]"
allowed-tools: Read, Write, Edit, Bash, Grep, Glob
MCP Builder
$ARGUMENTS
Build a production-grade MCP server following Anthropic's 4-phase methodology.
When to Use
- Wrapping a third-party REST API as MCP tools
- Exposing an internal database or service to Claude
- Creating reusable integrations for the team
- Migrating a custom tool into the MCP ecosystem
For MCP protocol theory, see `mcp-patterns` knowledge skill (auto-loaded).
4-Phase Workflow
Phase 1 — Research & Planning
1. Read the target API's documentation (OpenAPI spec, README, changelog). 2. Identify the 5-15 most useful operations. Prefer workflow-oriented tools over 1:1 API mirror. 3. Decide transport: `stdio` for local dev tools, `streamable-http` for remote/shared. 4. Decide language: **TypeScript recommended** (best SDK), Python acceptable (`mcp` package). 5. List required secrets (API keys, tokens) and their env var names.
Output: `PLAN.md` with tool list, transport choice, auth model.
Phase 2 — Implementation
Scaffold:
my-mcp/
├── package.json # or pyproject.toml
├── src/
│ ├── server.ts # entry point
│ ├── client.ts # API client (axios/httpx)
│ ├── tools/ # one file per tool
│ ├── schemas.ts # Zod/Pydantic schemas
│ └── errors.ts # typed errors
├── .env.example
└── README.md
Per tool:
- Input/output schemas (Zod for TS, Pydantic for Python)
- Clear `name` with service prefix (e.g. `github_create_issue`)
- Description starts with a verb, mentions trigger keywords
- Annotations: `readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`
- Pagination support via `cursor` or `page` parameters
- Focused responses — filter noise, don't dump raw API payloads
Phase 3 — Review & Testing
- TypeScript: `npm run typecheck && npm run lint && npm test`
- Python: `ruff check . && mypy --strict src/ && pytest`
- MCP Inspector dry-run:
npx @modelcontextprotocol/inspector node dist/server.js
- Verify each tool's schema validates a real request and rejects malformed input.
Phase 4 — Evaluation
Write 10 realistic end-user questions that an LLM should be able to answer using your server. Run them through Claude with the server attached. Grade: did the model call the right tool? Did the response give enough to answer? Fix the description, schema, or response format of any tool that failed.
Example eval questions for a `github-mcp`: 1. "What issues are open on repo X with label `bug`?" 2. "Create an issue titled Y in repo Z" 3. "Who has the most commits this month in repo X?"
When a tool fails an eval, the cause is almost always the description, not the schema. Score each tool against the description rubric in `mcp-patterns` (one-line purpose, WHEN TO USE, WHEN NOT TO USE, CRITICAL, self-test). A tool with an empty **WHEN NOT TO USE** is under-specified — it will misfire the moment a second tool in the same server overlaps with it, so add the boundary before re-running the eval. See `mcp-patterns` → "How to Write a Tool Description" for the full rubric and worked example.
Tool Design Checklist
- [ ] Name has service prefix and is verb-led
- [ ] Description mentions when to use it and includes trigger keywords
- [ ] Description carries a non-empty **WHEN NOT TO USE** that names overlapping tools (see `mcp-patterns` rubric)
- [ ] Input schema is strict, no free-form `object` with `additionalProperties: true`
- [ ] Output is focused — essential fields only, with pagination cursor if applicable
- [ ] Error responses are actionable ("API returned 403 — check `GITHUB_TOKEN` env var")
- [ ] Annotations set correctly (readonly/destructive/idempotent)
- [ ] No secrets logged or echoed in errors
- [ ] Rate limiting respects the upstream API
Transport Cheat Sheet
| Scenario | Transport | |----------|-----------| | Local dev tool, 1 user | `stdio` | | Remote server, multiple users | `streamable-http` with SSE | | Internal company tool, auth required | `streamable-http` + OAuth proxy | | Embedded in IDE/editor | `stdio` spawned by editor |
Registration Cheat Sheet
Local Claude Code (`.mcp.json`):
{
"mcpServers": {
"my-mcp": {
"command": "node",
"args": ["dist/server.js"],
"env": { "API_KEY": "$MY_API_KEY" }
}
}
}Global Claude Code (user-scope):
claude mcp add my-mcp --scope user -- node /path/to/server.js
Claude Desktop: same JSON, placed in `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS).
Common Pitfalls
| Mistake | Fix | |---------|-----| | 1:1 API mirror with 80 tools | Pick 10 workflow-oriented tools | | `description: "wrapper for /users endpoint"` | `description: "Find users by email, role, or team. Use when the user mentions employees, staff, or access"` | | Dumping raw JSON responses | Filter to 3-5 fields the agent actually needs | | Logging API keys on error | Redact all env vars in error formatters | | `exit 1` on transient errors | Retry with exponential backoff, surface final error | | Stdout pollution (MCP stdio) | All logs go to **stderr**, stdout is JSON-RPC only |
Rules
- **MUST** pick 5-15 workflow-oriented tools, not a 1:1 API mirror. The model routes by task, not by endpoint.
- **MUST** use strict input schemas (Zod for TS, Pydantic for Python). `additionalProperties: true` lets the model invent fields and drift.
- **MUST** set correct tool annotations: `readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint` — the host uses these for safety UIs and auto-approval policies
- **NEVER** expose an MCP server on a public network without auth. MCP clients default to trusting the transport — attackers r
Read more
name: mcp-builder description: "Builds production MCP servers via 4-phase methodology: research, implement, test, evaluate. Triggers: build MCP, new MCP, MCP integration, MCP server scaffold." effort: high disable-model-invocation: true argument-hint: "[service name or API description]" allowed-tools: Read, Write, Edit, Bash, Grep, Glob
MCP Builder
$ARGUMENTS
Build a production-grade MCP server following Anthropic's 4-phase methodology.
When to Use
- Wrapping a third-party REST API as MCP tools
- Exposing an internal database or service to Claude
- Creating reusable integrations for the team
- Migrating a custom tool into the MCP ecosystem
For MCP protocol theory, see `mcp-patterns` knowledge skill (auto-loaded).
4-Phase Workflow
Phase 1 — Research & Planning
1. Read the target API's documentation (OpenAPI spec, README, changelog). 2. Identify the 5-15 most useful operations. Prefer workflow-oriented tools over 1:1 API mirror. 3. Decide transport: `stdio` for local dev tools, `streamable-http` for remote/shared. 4. Decide language: **TypeScript recommended** (best SDK), Python acceptable (`mcp` package). 5. List required secrets (API keys, tokens) and their env var names.
Output: `PLAN.md` with tool list, transport choice, auth model.
Phase 2 — Implementation
Scaffold:
my-mcp/ ├── package.json # or pyproject.toml ├── src/ │ ├── server.ts # entry point │ ├── client.ts # API client (axios/httpx) │ ├── tools/ # one file per tool │ ├── schemas.ts # Zod/Pydantic schemas │ └── errors.ts # typed errors ├── .env.example └── README.md
Per tool:
- Input/output schemas (Zod for TS, Pydantic for Python)
- Clear `name` with service prefix (e.g. `github_create_issue`)
- Description starts with a verb, mentions trigger keywords
- Annotations: `readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`
- Pagination support via `cursor` or `page` parameters
- Focused responses — filter noise, don't dump raw API payloads
Phase 3 — Review & Testing
- TypeScript: `npm run typecheck && npm run lint && npm test`
- Python: `ruff check . && mypy --strict src/ && pytest`
- MCP Inspector dry-run:
npx @modelcontextprotocol/inspector node dist/server.js
- Verify each tool's schema validates a real request and rejects malformed input.
Phase 4 — Evaluation
Write 10 realistic end-user questions that an LLM should be able to answer using your server. Run them through Claude with the server attached. Grade: did the model call the right tool? Did the response give enough to answer? Fix the description, schema, or response format of any tool that failed.
Example eval questions for a `github-mcp`: 1. "What issues are open on repo X with label `bug`?" 2. "Create an issue titled Y in repo Z" 3. "Who has the most commits this month in repo X?"
When a tool fails an eval, the cause is almost always the description, not the schema. Score each tool against the description rubric in `mcp-patterns` (one-line purpose, WHEN TO USE, WHEN NOT TO USE, CRITICAL, self-test). A tool with an empty **WHEN NOT TO USE** is under-specified — it will misfire the moment a second tool in the same server overlaps with it, so add the boundary before re-running the eval. See `mcp-patterns` → "How to Write a Tool Description" for the full rubric and worked example.
Tool Design Checklist
- [ ] Name has service prefix and is verb-led
- [ ] Description mentions when to use it and includes trigger keywords
- [ ] Description carries a non-empty **WHEN NOT TO USE** that names overlapping tools (see `mcp-patterns` rubric)
- [ ] Input schema is strict, no free-form `object` with `additionalProperties: true`
- [ ] Output is focused — essential fields only, with pagination cursor if applicable
- [ ] Error responses are actionable ("API returned 403 — check `GITHUB_TOKEN` env var")
- [ ] Annotations set correctly (readonly/destructive/idempotent)
- [ ] No secrets logged or echoed in errors
- [ ] Rate limiting respects the upstream API
Transport Cheat Sheet
| Scenario | Transport | |----------|-----------| | Local dev tool, 1 user | `stdio` | | Remote server, multiple users | `streamable-http` with SSE | | Internal company tool, auth required | `streamable-http` + OAuth proxy | | Embedded in IDE/editor | `stdio` spawned by editor |
Registration Cheat Sheet
Local Claude Code (`.mcp.json`):
{
"mcpServers": {
"my-mcp": {
"command": "node",
"args": ["dist/server.js"],
"env": { "API_KEY": "$MY_API_KEY" }
}
}
}Global Claude Code (user-scope):
claude mcp add my-mcp --scope user -- node /path/to/server.js
Claude Desktop: same JSON, placed in `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS).
Common Pitfalls
| Mistake | Fix | |---------|-----| | 1:1 API mirror with 80 tools | Pick 10 workflow-oriented tools | | `description: "wrapper for /users endpoint"` | `description: "Find users by email, role, or team. Use when the user mentions employees, staff, or access"` | | Dumping raw JSON responses | Filter to 3-5 fields the agent actually needs | | Logging API keys on error | Redact all env vars in error formatters | | `exit 1` on transient errors | Retry with exponential backoff, surface final error | | Stdout pollution (MCP stdio) | All logs go to **stderr**, stdout is JSON-RPC only |
Rules
- **MUST** pick 5-15 workflow-oriented tools, not a 1:1 API mirror. The model routes by task, not by endpoint.
- **MUST** use strict input schemas (Zod for TS, Pydantic for Python). `additionalProperties: true` lets the model invent fields and drift.
- **MUST** set correct tool annotations: `readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint` — the host uses these for safety UIs and auto-approval policies
- **NEVER** expose an MCP server on a public network without auth. MCP clients default to trusting the transport — attackers r
Professional-grade AI coding toolkit with multi-platform support. Machine-enforced safety, 109 skills, 44 agents, expanded lifecycle hooks, persona presets, experimental opt-in plugin packs, and benchmark tooling — works with Claude Code, Claude Chat/Cowork,
Repo: softspark/ai-toolkit
Other skills on ai-toolkit.
- /ai-toolkit-rules
Mandatory engineering, security, testing, git, performance, quality, and response rules. Claude MUST load this skill for every technical, coding, debugging, review, architecture, DevOps, data, or file-editing task in Chat or Cowork.
Open skill - /mem-search
Search past coding sessions using natural language. Finds relevant observations, decisions, and context from previous work.
Open skill - /a11y-validate
Accessibility validator: WCAG 2.1 AA, EN 301 549, EAA. Triggers: a11y, accessibility, WCAG, EAA, ARIA, contrast, keyboard, screen reader.
Open skill - /agent-creator
Creates new specialized agents with frontmatter, tools, delegation. Triggers: new agent, create agent, agent scaffold, specialized agent.
Open skill - /analyze
Analyzes code quality, complexity, patterns across codebase. Triggers: quality report, hotspot scan, code analysis, architecture signal.
Open skill - /api-patterns
REST/GraphQL API design: naming, versioning, pagination, idempotency, OpenAPI. Triggers: API design, REST, GraphQL, OpenAPI, Swagger, idempotency, rate limit.
Open skill

