benchmark-e2e
End-to-end benchmark suite for vercel-plugin. Runs realistic projects through skill…
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. Designed to stress-test skill injection for complex, multi-system builds.
$ npx -y skills add vercel/vercel-plugin --skill benchmark-agents --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/benchmark-agentsContext preview
The summary Claude sees to decide when to auto-load this skill.
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. Designed to stress-test skill injection for complex, multi-system builds.
name: benchmark-agents description: Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features — Workflow SDK, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. Designed to stress-test skill injection for complex, multi-system builds.
Launch real Claude Code sessions with the plugin installed, verify skill injection, monitor PostToolUse validation catches, and produce a coverage report. This skill covers the full eval loop: setup → launch → monitor → verify → fix → release → repeat.
Evals are run by **you, in this conversation**, not by scripts. The process is:
1. You create directories and install the plugin via Bash tool calls 2. You spawn WezTerm panes with `wezterm cli spawn` — each pane runs an independent Claude Code interactive session 3. You wait, then check debug logs and claim dirs to see what the plugin injected 4. You inspect the generated source code for correctness 5. You read conversation logs to find what the user had to correct 6. You update skills/hooks, run `/release`, and spawn more evals
**Never use `claude --print`, eval scripts, or `Bun.spawn(["claude", ...])`**. These do not work because:
**The WezTerm interactive approach is the only method that exercises the plugin correctly.** Every eval in our history (60+ sessions) used this approach.
These are **absolute prohibitions**. Violating any of them wastes the entire eval run:
**Copy the exact commands below. Do not improvise.**
**Always append a timestamp** to directory names so reruns don't overwrite old projects:
<slug>-<yyyymmdd>-<hhmm>
Example: `tarot-card-deck-20260309-1227`, `interior-designer-20260309-1227`
Generate the timestamp with: `date +%Y%m%d-%H%M`
TS=$(date +%Y%m%d-%H%M) SLUG="my-app-$TS" mkdir -p ~/dev/vercel-plugin-testing/$SLUG cd ~/dev/vercel-plugin-testing/$SLUG npx add-plugin https://github.com/vercel/vercel-plugin -s project -y
wezterm cli spawn --cwd /Users/johnlindquist/dev/vercel-plugin-testing/$SLUG -- /bin/zsh -ic \ "unset CLAUDECODE; VERCEL_PLUGIN_LOG_LEVEL=debug x '<PROMPT>' --settings .claude/settings.json; exec zsh"
Key flags:
find ~/.claude/debug -name "*.txt" -mmin -2 -exec grep -l "$SLUG" {} +Create dirs and install plugin in a loop, then spawn each WezTerm pane:
TS=$(date +%Y%m%d-%H%M)
cd ~/dev/vercel-plugin-testing
for name in tarot-deck interior-designer superhero-origin; do
d="${name}-${TS}"
mkdir -p "$d" && (cd "$d" && npx add-plugin https://github.com/vercel/vercel-plugin -s project -y)
done
# Then spawn each (these run in separate terminal panes)
wezterm cli spawn --cwd .../tarot-deck-$TS -- /bin/zsh -ic "unset CLAUDECODE; VERCEL_PLUGIN_LOG_LEVEL=debug x '...' --settings .claude/settings.json; exec zsh"
wezterm cli spawn --cwd .../interior-designer-$TS -- /bin/zsh -ic "unset CLAUDECODE; VERCEL_PLUGIN_LOG_LEVEL=debug x '...' --settings .claude/settings.json; exec zsh"
wezterm cli spawn --cwd .../superhero-origin-$TS -- /bin/zsh -ic "unset CLAUDECODE; VERCEL_PLUGIN_LOG_LEVEL=debug x '...' --settings .claude/settings.json; exec zsh"TMPDIR=$(node -e "import {tmpdir} from 'os'; console.log(tmpdir())" --input-type=module)
CLAIMDIR="$TMPDIR/vercel-plugin-<session-id>-seen-skills.d"
# List all injected skills
ls "$CLAIMDIR"
# Count
ls "$CLAIMDIR" | wc -l
# Check specific skill
ls "$CLAIMDIR/workflow" && echo "YES" || echo "NO"LOG=~/.claude/debug/<session-id>.txt # SessionStart hooks grep -c 'SessionStart.*success' "$LOG" # PreToolUse calls and injections grep -c 'executePreToolHooks' "$LOG" # total calls grep -c 'provided additionalContext' "$LOG" # actual injections # PostToolUse validation catches grep 'VALIDATION' "$LOG" | head -10 # UserPromptSubmit grep -c 'UserPromptSubmit.*success' "$LOG"
TMPDIR=$(node -e "import {tmpdir} fromComprehensive Vercel ecosystem plugin — relational knowledge graph, skills for every major product, specialized agents, and Vercel conventions. Turns any AI agent into a Vercel expert.
Repo: vercel-labs/vercel-plugin
End-to-end benchmark suite for vercel-plugin. Runs realistic projects through skill…
Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels.…
Create and launch benchmark test projects to exercise vercel-plugin skill injection across…
Audit vercel-plugin performance on real-world projects. Extracts tool calls from Claude Code…
Release vercel-plugin — run gates, bump version, generate artifacts, commit, and push. Use…
Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection,…