benchmark-agents
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features —…
Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection, dedup correctness, and coverage. Launches real Claude Code sessions via WezTerm, monitors debug logs, and produces a structured coverage report.
$ npx -y skills add vercel/vercel-plugin --skill vercel-plugin-eval --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/vercel-plugin-evalContext preview
The summary Claude sees to decide when to auto-load this skill.
Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection, dedup correctness, and coverage. Launches real Claude Code sessions via WezTerm, monitors debug logs, and produces a structured coverage report.
name: vercel-plugin-eval description: Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection, dedup correctness, and coverage. Launches real Claude Code sessions via WezTerm, monitors debug logs, and produces a structured coverage report.
Launch real Claude Code sessions with the plugin installed, monitor debug logs in real-time, and verify every hook fires correctly with proper dedup.
**Copy the exact commands below. Do not improvise.**
**Always append a timestamp** to directory names so reruns don't overwrite old projects:
# 1. Create test dir & install plugin (with timestamp)
TS=$(date +%Y%m%d-%H%M)
SLUG="my-eval-$TS"
mkdir -p ~/dev/vercel-plugin-testing/$SLUG
cd ~/dev/vercel-plugin-testing/$SLUG
npx add-plugin https://github.com/vercel/vercel-plugin -s project -y
# 2. Launch session via WezTerm
wezterm cli spawn --cwd /Users/johnlindquist/dev/vercel-plugin-testing/$SLUG -- /bin/zsh -ic \
"unset CLAUDECODE; VERCEL_PLUGIN_LOG_LEVEL=debug x '<PROMPT>' --settings .claude/settings.json; exec zsh"
# 3. Find debug log (wait ~25s for session start)
find ~/.claude/debug -name "*.txt" -mmin -2 -exec grep -l "$SLUG" {} +LOG=~/.claude/debug/<session-id>.txt # SessionStart (3 hooks) grep "SessionStart.*success" "$LOG" # PreToolUse skill injection grep -c "executePreToolHooks" "$LOG" # total calls grep -c "provided additionalContext" "$LOG" # injections # UserPromptSubmit grep "UserPromptSubmit.*success" "$LOG" # PostToolUse validate + shadcn font-fix grep "posttooluse-validate.*provided" "$LOG" grep "PostToolUse:Bash.*success" "$LOG" # SessionEnd cleanup grep "SessionEnd" "$LOG"
TMPDIR=$(node -e "import {tmpdir} from 'os'; console.log(tmpdir())" --input-type=module)
CLAIMDIR="$TMPDIR/vercel-plugin-<session-id>-seen-skills.d"
# Claim files = one per skill, atomic O_EXCL
ls "$CLAIMDIR"
# Compare: injections should equal claims
inject_meta=$(grep -c "skillInjection:" "$LOG")
claims=$(ls "$CLAIMDIR" 2>/dev/null | wc -l | tr -d ' ')
echo "Injections: $((inject_meta / 3)) | Claims: $claims"`skillInjection:` appears 3x per actual injection in the debug log (initial check, parsed, success). Divide by 3.
Look for real catches — API key bypass, outdated models, wrong patterns:
grep "VALIDATION" "$LOG" | head -10
Describe **products and features**, never name specific technologies. Let the plugin infer which skills to inject. Always end prompts with: "Link the project to my vercel-labs team so we can deploy it later. Skip any planning and just build it. Get the dev server running."
| Scenario Type | Skills Exercised | |--------------|-----------------| | AI chat app | ai-sdk, ai-gateway, nextjs, ai-elements | | Durable workflow | workflow, ai-sdk, vercel-queues | | Monorepo | turborepo, turbopack, nextjs | | Edge auth + routing | routing-middleware, auth, sign-in-with-vercel | | Chat bot (multi-platform) | chat-sdk, ai-sdk, vercel-storage | | Feature flags + CRM | vercel-flags, vercel-queues, ai-sdk | | Email pipeline | email, satori, ai-sdk, vercel-storage | | Marketplace/payments | payments, marketplace, cms | | Kitchen sink | micro, ncc, all niche skills |
These need explicit technology references in the prompt because agents don't naturally reach for them:
Write results to `.notes/COVERAGE.md` with:
1. **Session index** — slug, session ID, unique skills, dedup status 2. **Hook coverage matrix** — which hooks fired in which sessions 3. **Skill injection table** — which of the 44 skills triggered 4. **Dedup stats** — injections vs claims per session 5. **Issues found** — bugs, pattern gaps, validation findings
rm -rf ~/dev/vercel-plugin-testing
Comprehensive Vercel ecosystem plugin — relational knowledge graph, skills for every major product, specialized agents, and Vercel conventions. Turns any AI agent into a Vercel expert.
Repo: vercel-labs/vercel-plugin
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features —…
End-to-end benchmark suite for vercel-plugin. Runs realistic projects through skill…
Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels.…
Create and launch benchmark test projects to exercise vercel-plugin skill injection across…
Audit vercel-plugin performance on real-world projects. Extracts tool calls from Claude Code…
Release vercel-plugin — run gates, bump version, generate artifacts, commit, and push. Use…