prompt-evaluation-runn…
Use when evaluating prompts, LLM outputs, red-team suites, or model behavior with local eval configs and safe provider/cost controls.
Use when confronted with an unknown failure in CI or production, before committing to a deep debugging approach.
$ npx -y skills add yeaight7/agent-powerups --skill failure-triage --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/failure-triageContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when confronted with an unknown failure in CI or production, before committing to a deep debugging approach.
name: failure-triage description: Use when confronted with an unknown failure in CI or production, before committing to a deep debugging approach.
Before diving deep into a stack trace or spending hours reproducing a bug, triage it to determine the blast radius, subsystem, and debugging approach. Do not start writing fixes until you have explicitly stated your triage hypothesis and confirmed the category.
1. **Categorize the failure:**
2. **Locate the origin.** Scan the stack trace. Ignore framework/library internals. Find the highest frame that belongs to the *first-party application code*:
# surface first-party frames (adjust the path filter to the repo layout) grep -n "src/" stacktrace.txt | head -20
3. **Check recent changes.** Most bugs are in the newest code:
git log -n 5 --oneline git diff HEAD~5 --stat git log -n 10 --oneline -- <suspect-file-or-dir>
4. **Formulate a hypothesis.** State clearly: "I suspect this is an environment error caused by missing configuration, originating in `src/config.ts`."
Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more
Repo: yeaight7/agent-powerups
Use when evaluating prompts, LLM outputs, red-team suites, or model behavior with local eval configs and safe provider/cost controls.
Use when creating or reviewing red-team eval plugins, attack templates, grader rubrics, safety fixtures, or model-risk test metadata.
Use when designing, running, debugging, or hardening deterministic eval suites for agent skills, prompts, tool workflows, or MCP-backed cases.
Use when designing tool definitions for a new agent or subagent, an agent shows high retry rates, ambiguous tool invocations, or silent failures, or an…
Use when routing a prompt to a local provider CLI for a second opinion, review, or plan -- you are about to call a provider directly, need the response saved…
Use when starting work in an unfamiliar area of a codebase, spawning a subagent that needs targeted file context, a first search pass missed the relevant file,…