benchmark-agents
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features —…
Audit vercel-plugin performance on real-world projects. Extracts tool calls from Claude Code conversation logs, tests hook matching against actual inputs, identifies pattern coverage gaps, and checks plugin cache staleness. Use when asked to audit, test, or investigate plugin
$ npx -y skills add vercel/vercel-plugin --skill plugin-audit --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/plugin-auditContext preview
The summary Claude sees to decide when to auto-load this skill.
Audit vercel-plugin performance on real-world projects. Extracts tool calls from Claude Code conversation logs, tests hook matching against actual inputs, identifies pattern coverage gaps, and checks plugin cache staleness. Use when asked to audit, test, or investigate plugin
name: plugin-audit description: Audit vercel-plugin performance on real-world projects. Extracts tool calls from Claude Code conversation logs, tests hook matching against actual inputs, identifies pattern coverage gaps, and checks plugin cache staleness. Use when asked to audit, test, or investigate plugin skill injection on a real project.
Audit how well vercel-plugin skill injection performs on real-world Claude Code sessions.
Find JSONL conversation logs for a target project:
ls -lt ~/.claude/projects/-Users-*-<project-name>/*.jsonl
The path uses the project's absolute path with slashes replaced by hyphens and a leading hyphen.
Parse the JSONL log to extract all tool_use entries. Each line is a JSON object with `message.content[]` containing `type: "tool_use"` blocks. Extract `name` and `input` fields. Group by tool type (Bash, Read, Write, Edit).
Use the exported pipeline functions directly — do NOT shell out to the hook script for each test. Import from the hooks directory:
import { loadSkills, matchSkills } from "./hooks/pretooluse-skill-inject.mjs";
import { createLogger } from "./hooks/logger.mjs";Call `loadSkills()` once, then `matchSkills(toolName, toolInput, compiledSkills)` for each tool call. This is fast and gives exact match results.
Compare matched skills against what SHOULD have matched based on the project's technology stack. Common gap categories:
Compare the installed plugin cache against the dev version:
# Cache location ~/.claude/plugins/cache/vercel-labs-vercel-plugin/vercel-plugin/<version>/ # Compare skill content diff <(grep 'pattern' skills/<skill>/SKILL.md) <(grep 'pattern' ~/.claude/plugins/cache/.../skills/<skill>/SKILL.md)
Check `~/.claude/plugins/installed_plugins.json` for version and git SHA.
Produce a structured report with:
1. **Session summary**: Project, date, tool call count, model 2. **Match matrix**: Table of tool calls × matched skills (with match type) 3. **Coverage gaps**: Unmatched tool calls that should have matched, with suggested pattern additions 4. **Dedup timeline**: Order of skill injections and what got deduped 5. **Cache status**: Whether installed version matches dev, with specific diffs
Comprehensive Vercel ecosystem plugin — relational knowledge graph, skills for every major product, specialized agents, and Vercel conventions. Turns any AI agent into a Vercel expert.
Repo: vercel-labs/vercel-plugin
Advanced AI agent benchmark scenarios that push Vercel's cutting-edge platform features —…
End-to-end benchmark suite for vercel-plugin. Runs realistic projects through skill…
Run vercel-plugin eval scenarios in Vercel Sandboxes instead of local WezTerm panels.…
Create and launch benchmark test projects to exercise vercel-plugin skill injection across…
Release vercel-plugin — run gates, bump version, generate artifacts, commit, and push. Use…
Run live eval sessions against the vercel-plugin to verify hook behavior, skill injection,…