/vercel-plugin-eval
Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection, dedup correctness, and coverage. Launches real Claude Code sessions via WezTerm, monitors debug logs, and produces a structured coverage report.
$ npx -y skills add vercel-labs/vercel-plugin --skill vercel-plugin-eval --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/vercel-plugin-eval
Context preview
The summary Claude sees to decide when to auto-load this skill.
Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection, dedup correctness, and coverage. Launches real Claude Code sessions via WezTerm, monitors debug logs, and produces a structured coverage report.
SKILL.md
vercel-plugin-eval.SKILL.mdname: vercel-plugin-eval
description: Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection, dedup correctness, and coverage. Launches real Claude Code sessions via WezTerm, monitors debug logs, and produces a structured coverage report.
Plugin Eval
Launch real Claude Code sessions with the plugin installed, monitor debug logs in real-time, and verify every hook fires correctly with proper dedup.
DO NOT (Hard Rules)
- **DO NOT** use `claude --print` or `-p` — hooks don't fire, no files created
- **DO NOT** use `--dangerously-skip-permissions`
- **DO NOT** create projects in `/tmp/` — always use `~/dev/vercel-plugin-testing/`
- **DO NOT** manually wire hooks or create `settings.local.json` — use `npx add-plugin`
- **DO NOT** set `CLAUDE_PLUGIN_ROOT` manually
- **DO NOT** use `bash -c` in WezTerm — use `/bin/zsh -ic`
- **DO NOT** use full path to claude — use the `x` alias
- **DO NOT** write eval scripts — do everything as Bash tool calls in the conversation
**Copy the exact commands below. Do not improvise.**
Quick Start
**Always append a timestamp** to directory names so reruns don't overwrite old projects:
# 1. Create test dir & install plugin (with timestamp)
TS=$(date +%Y%m%d-%H%M)
SLUG="my-eval-$TS"
mkdir -p ~/dev/vercel-plugin-testing/$SLUG
cd ~/dev/vercel-plugin-testing/$SLUG
npx add-plugin https://github.com/vercel/vercel-plugin -s project -y
# 2. Launch session via WezTerm
wezterm cli spawn --cwd /Users/johnlindquist/dev/vercel-plugin-testing/$SLUG -- /bin/zsh -ic \
"unset CLAUDECODE; VERCEL_PLUGIN_LOG_LEVEL=debug x '<PROMPT>' --settings .claude/settings.json; exec zsh"
# 3. Find debug log (wait ~25s for session start)
find ~/.claude/debug -name "*.txt" -mmin -2 -exec grep -l "$SLUG" {} +What to Monitor
Hook firing (all 8 registered hooks)
LOG=~/.claude/debug/<session-id>.txt
# SessionStart (3 hooks)
grep "SessionStart.*success" "$LOG"
# PreToolUse skill injection
grep -c "executePreToolHooks" "$LOG" # total calls
grep -c "provided additionalContext" "$LOG" # injections
# UserPromptSubmit
grep "UserPromptSubmit.*success" "$LOG"
# PostToolUse validate + shadcn font-fix
grep "posttooluse-validate.*provided" "$LOG"
grep "PostToolUse:Bash.*success" "$LOG"
# SessionEnd cleanup
grep "SessionEnd" "$LOG"
Dedup correctness (the key metric)
TMPDIR=$(node -e "import {tmpdir} from 'os'; console.log(tmpdir())" --input-type=module)
CLAIMDIR="$TMPDIR/vercel-plugin-<session-id>-seen-skills.d"
# Claim files = one per skill, atomic O_EXCL
ls "$CLAIMDIR"
# Compare: injections should equal claims
inject_meta=$(grep -c "skillInjection:" "$LOG")
claims=$(ls "$CLAIMDIR" 2>/dev/null | wc -l | tr -d ' ')
echo "Injections: $((inject_meta / 3)) | Claims: $claims"`skillInjection:` appears 3x per actual injection in the debug log (initial check, parsed, success). Divide by 3.
PostToolUse validate quality
Look for real catches — API key bypass, outdated models, wrong patterns:
grep "VALIDATION" "$LOG" | head -10
Scenario Design
Describe **products and features**, never name specific technologies. Let the plugin infer which skills to inject. Always end prompts with: "Link the project to my vercel-labs team so we can deploy it later. Skip any planning and just build it. Get the dev server running."
Coverage targets by scenario type
| Scenario Type | Skills Exercised | |--------------|-----------------| | AI chat app | ai-sdk, ai-gateway, nextjs, ai-elements | | Durable workflow | workflow, ai-sdk, vercel-queues | | Monorepo | turborepo, turbopack, nextjs | | Edge auth + routing | routing-middleware, auth, sign-in-with-vercel | | Chat bot (multi-platform) | chat-sdk, ai-sdk, vercel-storage | | Feature flags + CRM | vercel-flags, vercel-queues, ai-sdk | | Email pipeline | email, satori, ai-sdk, vercel-storage | | Marketplace/payments | payments, marketplace, cms | | Kitchen sink | micro, ncc, all niche skills |
Hard-to-trigger skills (8 of 44)
These need explicit technology references in the prompt because agents don't naturally reach for them:
- `ai-elements` — say "use the AI Elements component registry"
- `v0-dev` — say "generate components with v0"
- `vercel-firewall` — say "use Vercel Firewall for rate limiting"
- `marketplace` — say "publish to the Vercel Marketplace"
- `geist` — say "install the geist font package"
- `json-render` — name files `components/chat-*.tsx`
Coverage Report
Write results to `.notes/COVERAGE.md` with:
1. **Session index** — slug, session ID, unique skills, dedup status 2. **Hook coverage matrix** — which hooks fired in which sessions 3. **Skill injection table** — which of the 44 skills triggered 4. **Dedup stats** — injections vs claims per session 5. **Issues found** — bugs, pattern gaps, validation findings
Cleanup
rm -rf ~/dev/vercel-plugin-testing
Read more
name: vercel-plugin-eval description: Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection, dedup correctness, and coverage. Launches real Claude Code sessions via WezTerm, monitors debug logs, and produces a structured coverage report.
Plugin Eval
Launch real Claude Code sessions with the plugin installed, monitor debug logs in real-time, and verify every hook fires correctly with proper dedup.
DO NOT (Hard Rules)
- **DO NOT** use `claude --print` or `-p` — hooks don't fire, no files created
- **DO NOT** use `--dangerously-skip-permissions`
- **DO NOT** create projects in `/tmp/` — always use `~/dev/vercel-plugin-testing/`
- **DO NOT** manually wire hooks or create `settings.local.json` — use `npx add-plugin`
- **DO NOT** set `CLAUDE_PLUGIN_ROOT` manually
- **DO NOT** use `bash -c` in WezTerm — use `/bin/zsh -ic`
- **DO NOT** use full path to claude — use the `x` alias
- **DO NOT** write eval scripts — do everything as Bash tool calls in the conversation
**Copy the exact commands below. Do not improvise.**
Quick Start
**Always append a timestamp** to directory names so reruns don't overwrite old projects:
# 1. Create test dir & install plugin (with timestamp)
TS=$(date +%Y%m%d-%H%M)
SLUG="my-eval-$TS"
mkdir -p ~/dev/vercel-plugin-testing/$SLUG
cd ~/dev/vercel-plugin-testing/$SLUG
npx add-plugin https://github.com/vercel/vercel-plugin -s project -y
# 2. Launch session via WezTerm
wezterm cli spawn --cwd /Users/johnlindquist/dev/vercel-plugin-testing/$SLUG -- /bin/zsh -ic \
"unset CLAUDECODE; VERCEL_PLUGIN_LOG_LEVEL=debug x '<PROMPT>' --settings .claude/settings.json; exec zsh"
# 3. Find debug log (wait ~25s for session start)
find ~/.claude/debug -name "*.txt" -mmin -2 -exec grep -l "$SLUG" {} +What to Monitor
Hook firing (all 8 registered hooks)
LOG=~/.claude/debug/<session-id>.txt # SessionStart (3 hooks) grep "SessionStart.*success" "$LOG" # PreToolUse skill injection grep -c "executePreToolHooks" "$LOG" # total calls grep -c "provided additionalContext" "$LOG" # injections # UserPromptSubmit grep "UserPromptSubmit.*success" "$LOG" # PostToolUse validate + shadcn font-fix grep "posttooluse-validate.*provided" "$LOG" grep "PostToolUse:Bash.*success" "$LOG" # SessionEnd cleanup grep "SessionEnd" "$LOG"
Dedup correctness (the key metric)
TMPDIR=$(node -e "import {tmpdir} from 'os'; console.log(tmpdir())" --input-type=module)
CLAIMDIR="$TMPDIR/vercel-plugin-<session-id>-seen-skills.d"
# Claim files = one per skill, atomic O_EXCL
ls "$CLAIMDIR"
# Compare: injections should equal claims
inject_meta=$(grep -c "skillInjection:" "$LOG")
claims=$(ls "$CLAIMDIR" 2>/dev/null | wc -l | tr -d ' ')
echo "Injections: $((inject_meta / 3)) | Claims: $claims"`skillInjection:` appears 3x per actual injection in the debug log (initial check, parsed, success). Divide by 3.
PostToolUse validate quality
Look for real catches — API key bypass, outdated models, wrong patterns:
grep "VALIDATION" "$LOG" | head -10
Scenario Design
Describe **products and features**, never name specific technologies. Let the plugin infer which skills to inject. Always end prompts with: "Link the project to my vercel-labs team so we can deploy it later. Skip any planning and just build it. Get the dev server running."
Coverage targets by scenario type
| Scenario Type | Skills Exercised | |--------------|-----------------| | AI chat app | ai-sdk, ai-gateway, nextjs, ai-elements | | Durable workflow | workflow, ai-sdk, vercel-queues | | Monorepo | turborepo, turbopack, nextjs | | Edge auth + routing | routing-middleware, auth, sign-in-with-vercel | | Chat bot (multi-platform) | chat-sdk, ai-sdk, vercel-storage | | Feature flags + CRM | vercel-flags, vercel-queues, ai-sdk | | Email pipeline | email, satori, ai-sdk, vercel-storage | | Marketplace/payments | payments, marketplace, cms | | Kitchen sink | micro, ncc, all niche skills |
Hard-to-trigger skills (8 of 44)
These need explicit technology references in the prompt because agents don't naturally reach for them:
- `ai-elements` — say "use the AI Elements component registry"
- `v0-dev` — say "generate components with v0"
- `vercel-firewall` — say "use Vercel Firewall for rate limiting"
- `marketplace` — say "publish to the Vercel Marketplace"
- `geist` — say "install the geist font package"
- `json-render` — name files `components/chat-*.tsx`
Coverage Report
Write results to `.notes/COVERAGE.md` with:
1. **Session index** — slug, session ID, unique skills, dedup status 2. **Hook coverage matrix** — which hooks fired in which sessions 3. **Skill injection table** — which of the 44 skills triggered 4. **Dedup stats** — injections vs claims per session 5. **Issues found** — bugs, pattern gaps, validation findings
Cleanup
rm -rf ~/dev/vercel-plugin-testing
Comprehensive Vercel ecosystem plugin — relational knowledge graph, skills for every major product, specialized agents, and Vercel conventions. Turns any AI agent into a Vercel expert.
Repo: vercel-labs/vercel-plugin
Other skills on vercel.
- /benchmark-agents
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow DevKit, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. Designed to stress-test skill injection for complex, multi-system builds.
Open skill - /benchmark-e2e
End-to-end benchmark suite for vercel-plugin. Runs realistic projects through skill injection, launches dev servers, verifies everything works, analyzes conversation logs, and produces an improvement report for overnight self-improvement loops.
Open skill - /benchmark-sandbox
Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels. Provisions ephemeral microVMs with Claude Code + plugin pre-installed, runs benchmark prompts, extracts hook artifacts, and produces coverage reports.
Open skill - /benchmark-testing
Create and launch benchmark test projects to exercise vercel-plugin skill injection across realistic scenarios. Sets up isolated directories, installs the plugin, and spawns WezTerm panes running Claude Code with crafted prompts.
Open skill - /plugin-audit
Audit vercel-plugin performance on real-world projects. Extracts tool calls from Claude Code conversation logs, tests hook matching against actual inputs, identifies pattern coverage gaps, and checks plugin cache staleness. Use when asked to audit, test, or investigate plugin
Open skill - /release
Release vercel-plugin — run gates, bump version, generate artifacts, commit, and push. Use when asked to "release", "ship", "bump and push", or "cut a release".
Open skill

