gsd-headless
Orchestrate GSD (Git Ship Done) projects programmatically via headless CLI. Use when an agent…
Build, iterate, and evaluate Model Context Protocol (MCP) servers that expose external services as tools an LLM can call. Covers schema/tool design, error handling, pagination, MCP Inspector testing, and an eval set. Use when asked to "build an MCP server", "create an MCP tool",
$ npx -y skills add open-gsd/gsd-pi --skill create-mcp-server --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/create-mcp-serverContext preview
The summary Claude sees to decide when to auto-load this skill.
Build, iterate, and evaluate Model Context Protocol (MCP) servers that expose external services as tools an LLM can call. Covers schema/tool design, error handling, pagination, MCP Inspector testing, and an eval set. Use when asked to "build an MCP server", "create an MCP tool",
name: create-mcp-server description: Build, iterate, and evaluate Model Context Protocol (MCP) servers that expose external services as tools an LLM can call. Covers schema/tool design, error handling, pagination, MCP Inspector testing, and an eval set. Use when asked to "build an MCP server", "create an MCP tool", "wrap this API as MCP", "expose X to Claude", or when extending GSD with custom tool integrations.
<objective> Produce a high-quality MCP server that an LLM can actually use — not one that merely parses spec-compliant. Quality is measured by how well the server enables real-world task completion, which means the tool descriptions, error messages, and pagination behave under model reasoning, not just at the wire level. </objective>
<context> gsd-pi consumes MCP heavily — see `src/resources/extensions/mcp-client/`, `src/resources/extensions/gsd/mcp-project-config.ts`, `src/resources/extensions/gsd/workflow-mcp.ts`, and `/gsd mcp` commands. Users frequently want to extend GSD with project-specific MCP servers (internal APIs, data sources, domain tools). This skill fills the authoring gap between "MCP exists" and "I have a working server."
Invocation points:
</context>
<core_principle> **THE QUALITY METRIC IS TASK COMPLETION, NOT SCHEMA VALIDITY.** A server that lists 30 tools with cryptic names and empty descriptions passes the protocol but fails the point. The tool description is the only thing an LLM has to decide whether to call it — write it like documentation for a stranger under time pressure.
**DESIGN FOR THE MODEL, NOT THE API.** A raw REST endpoint is rarely the right tool. Group, filter, and pre-shape responses so the model gets what it needs to reason, not a 40KB JSON blob it has to summarize. Fewer, deeper tools beat many, shallow ones. </core_principle>
<process>
1. **Study modern MCP design.** Read the latest MCP protocol docs (not training data — fetch them). Read 2–3 reference implementations to see current patterns. 2. **Pick a framework.** TypeScript is the default — the reference SDK is the most mature. Python is fine for data-heavy or ML adjacencies. 3. **Analyze the target API.** Map the external service's endpoints, auth, rate limits, pagination, error shapes. Identify what a human workflow on top of it actually looks like — that's the cut line for tool design. 4. **Produce a brief.** One page: what the server does, who calls it, the 5–10 tools you plan to expose, and the top 3 design trade-offs. Confirm with the user.
Skeleton:
server/
src/
index.ts # MCP entry point — stdio or sse transport
client.ts # API client with auth, retries, typed errors
tools/ # one file per tool, or grouped by domain
pagination.ts # shared cursor handling
errors.ts # MCP-friendly error formatting
package.json # @modelcontextprotocol/sdk as dep
tsconfig.json
README.md # how to run, env vars, rate-limit notes
evals.xml # 10 eval questions (Phase 4)Core infrastructure goes first: API client with typed errors, pagination helpers, consistent retry/timeout behavior. Do not inline these per tool.
For each tool:
1. **Name:** verb-noun, lowercase, snake_case. `search_issues`, `get_customer`, `create_deployment`. Not `do_thing` or `api_v2_post`. 2. **Description (frontmatter):** 2–4 sentences. State what the tool does, when to use it, when NOT to use it, and any required fields or quirks. This is the model's entire interface to the tool — write it carefully. 3. **Input schema (JSON Schema):** required fields marked, every field has a description, enums enumerated, examples included for free-form strings. 4. **Output shape:** typed, minimal, decision-ready. If the raw API returns 40 fields and only 6 matter for follow-up calls, return 6. 5. **Error handling:** never return raw HTTP errors. Translate to human-readable messages: "Rate limit exceeded (retry in 30s)", "Authorization expired", "No record found for ID X". Include the action the caller should take next. 6. **Pagination:** expose cursors explicitly. Do not leak "page N of M" into the model — leak "more results available, pass `cursor: abc123` to continue."
1. Run the server under MCP Inspector. Verify it registers, every tool lists with its description, inputs schema-validate, outputs shape correctly. 2. Call every tool at least once manually through the Inspector UI. Check error paths. 3. Fix any "looks fine in isolation, breaks under the Inspector's framing" issues.
Write 10 evaluation questions in `evals.xml` that exercise the server end-to-end. Each question should require 2+ tool calls and at least one decision the model has to make based on earlier output. Cover:
Format:
<evals>
<eval id="1">
<question>...user request...</question>
<expected>...concrete observable answer or tool-call sequence...</expected>
</eval>
</evals>Run the evals. If the model can't complete them, the server — not the model — needs work. Iterate on descriptions, error messages, and tool granularity.
Write the project's `.mcp.json` entry using `/gsd mcp init` as a starting point. Document env vars and startup in README.md. If the server is globally useful, suggest the user file it as a durable skill via `spike-wrap-up` or publish it.
</process>
<anti_patterns>
GSD Pi is a local-first coding agent for planning, implementing, verifying, and tracking project work from the command line.
Repo: open-gsd/gsd-pi
Orchestrate GSD (Git Ship Done) projects programmatically via headless CLI. Use when an agent…
Audit and improve web accessibility following WCAG 2.1 guidelines. Use when asked to "improve…
Browser automation CLI for AI agents. Use when interacting with websites — navigating pages,…
Design or review an HTTP/REST/GraphQL API for versioning, pagination, error shapes,…
Apply modern web development best practices for security, compatibility, and code quality.…
Ask a quick side question about your current work without derailing the main task. Answers…