prompt-evaluation-runn…
Use when evaluating prompts, LLM outputs, red-team suites, or model behavior with local eval configs and safe provider/cost controls.
Use when an agent has modified logic, API routes, or data-transformation code, a bug was just fixed and must not be reintroduced, or behavior must stay identical across execution paths such as sandbox versus production.
$ npx -y skills add yeaight7/agent-powerups --skill ai-regression-testing --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/ai-regression-testingContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when an agent has modified logic, API routes, or data-transformation code, a bug was just fixed and must not be reintroduced, or behavior must stay identical across execution paths such as sandbox versus production.
name: ai-regression-testing description: Use when an agent has modified logic, API routes, or data-transformation code, a bug was just fixed and must not be reintroduced, or behavior must stay identical across execution paths such as sandbox versus production.
When an agent writes code and then reviews it, it carries the same assumptions into both steps. Automated tests break this cycle.
Agent writes fix → Agent reviews fix → Agent says "looks correct" → Bug still present
The most common blind spot: an agent fixes the production path but leaves the sandbox/mock path unchanged, or vice versa.
Run in order. Do not skip to agent review if automated steps fail.
npm test # or: pytest, cargo test, go test ./... npm run build # TypeScript build / type check
With tests passing, do a focused review for patterns agents commonly miss:
1. **Execution path parity**: Do all code paths (sandbox, production, feature-flag on/off) return the same response shape? 2. **Query completeness**: Are all fields used in the response present in the query or selection? 3. **Error state cleanup**: On error, is stale state cleared before the error is surfaced? 4. **Optimistic update rollback**: If an API call fails, is the optimistic UI change reverted?
For every bug found and fixed, add a test immediately:
Bug: <description> File: <path> Regression test: <test name and what it asserts>
If you cannot write a test, document why:
Bug: <description> Regression test: DEFERRED — <reason> (e.g., requires E2E harness not yet in place)
Do not silently skip. Every real bug should either have a test or an explicit deferral note.
Test the contract, not the implementation:
// Test what the consumer receives, not how it's computed
const REQUIRED_RESPONSE_FIELDS = ["id", "email", "settings", "created_at"];
it("profile endpoint returns all required fields", async () => {
const res = await GET(createRequest("/api/user/profile"));
const json = await res.json();
for (const field of REQUIRED_RESPONSE_FIELDS) {
expect(json.data).toHaveProperty(field);
}
});Name tests after the bug category, not the fix:
it("sandbox path returns same field set as production path (BUG-CLASS: path-parity)")
it("notification_settings is not undefined after SELECT * removal (regression)")| Pattern | Check | Priority | |---------|-------|----------| | Execution path parity | Same response shape across all paths | High | | Query field omission | All response fields present in DB query | High | | Error state leakage | State cleared before error is returned | Medium | | Missing rollback | Previous state restored on API failure | Medium |
Do not aim for coverage percentage. Write tests only for bugs that were found. Bug clusters naturally: if three bugs appeared in `/api/user/profile`, that endpoint needs tests. An endpoint that has never had a bug does not need tests yet.
Tests added this way grow organically with the bug history and cannot be gamed by coverage metrics.
Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more
Repo: yeaight7/agent-powerups
Use when evaluating prompts, LLM outputs, red-team suites, or model behavior with local eval configs and safe provider/cost controls.
Use when creating or reviewing red-team eval plugins, attack templates, grader rubrics, safety fixtures, or model-risk test metadata.
Use when designing, running, debugging, or hardening deterministic eval suites for agent skills, prompts, tool workflows, or MCP-backed cases.
Use when designing tool definitions for a new agent or subagent, an agent shows high retry rates, ambiguous tool invocations, or silent failures, or an…
Use when routing a prompt to a local provider CLI for a second opinion, review, or plan -- you are about to call a provider directly, need the response saved…
Use when starting work in an unfamiliar area of a codebase, spawning a subagent that needs targeted file context, a first search pass missed the relevant file,…