prompt-evaluation-runn…
Use when evaluating prompts, LLM outputs, red-team suites, or model behavior with local eval configs and safe provider/cost controls.
Use when implementing any feature or bugfix, before writing implementation code
$ npx -y skills add yeaight7/agent-powerups --skill test-driven-development --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/test-driven-developmentContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when implementing any feature or bugfix, before writing implementation code
name: test-driven-development description: Use when implementing any feature or bugfix, before writing implementation code
Write the test first. Watch it fail. Write minimal code to pass.
**Core principle:** If you didn't watch the test fail, you don't know if it tests the right thing.
**Violating the letter of the rules is violating the spirit of the rules.**
**Always:**
**Exceptions (ask your human partner):**
Thinking "skip TDD just this once"? Stop. That's rationalization.
NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST
Write code before the test? Delete it. Start over.
**No exceptions:**
Implement fresh from tests. Period.
graph LR
red["RED<br/>Write failing test"]
style red fill:#ffcccc
verify_red{"Verify fails<br/>correctly"}
green["GREEN<br/>Minimal code"]
style green fill:#ccffcc
verify_green{"Verify passes<br/>All green"}
refactor["REFACTOR<br/>Clean up"]
style refactor fill:#ccccff
next_step(["Next"])
red --> verify_red
verify_red -->|yes| green
verify_red -->|"wrong<br/>failure"| red
green --> verify_green
verify_green -->|yes| refactor
verify_green -->|no| green
refactor -->|"stay<br/>green"| verify_green
verify_green --> next_step
next_step --> redWrite one minimal test showing what should happen.
**Requirements:**
**MANDATORY. Never skip.**
Confirm:
**Test passes?** You're testing existing behavior. Fix test.
**Test errors?** Fix error, re-run until it fails correctly.
Write simplest code to pass the test. Don't add features, refactor other code, or "improve" beyond the test.
**MANDATORY.**
Confirm:
**Test fails?** Fix code, not test.
After green only:
Keep tests green. Don't add behavior.
Next failing test for next feature.
**"I'll write tests after to verify it works"**
Tests written after code pass immediately — proving nothing. You never saw the test catch the bug.
**"I already manually tested all the edge cases"**
Manual testing is ad-hoc. No record, can't re-run, easy to forget cases under pressure.
**"Deleting X hours of work is wasteful"**
Sunk cost fallacy. Working code without real tests is technical debt.
| Excuse | Reality | | --- | --- | | "Too simple to test" | Simple code breaks. Test takes 30 seconds. | | "I'll test after" | Tests passing immediately prove nothing. | | "Tests after achieve same goals" | Tests-after = "what does this do?" Tests-first = "what should this do?" | | "Already manually tested" | Ad-hoc ≠ systematic. No record, can't re-run. | | "Deleting X hours is wasteful" | Sunk cost fallacy. Keeping unverified code is technical debt. | | "Keep as reference, write tests first" | You'll adapt it. That's testing after. Delete means delete. | | "TDD will slow me down" | TDD faster than debugging. Pragmatic = test-first. |
**All of these mean: Delete code. Start over with TDD.**
Before marking work complete:
Can't check all boxes? You skipped TDD. Start over.
| Problem | Solution | | --- | --- | | Don't know how to test | Write wished-for API. Write assertion first. Ask your human partner. | | Test too complicated | Design too complicated. Simplify interface. | | Must mock everything | Code too coupled. Use dependency injection. | | Test setup huge | Extract helpers. Still complex? Simplify design. |
Bug found? Write failing test reproducing it. Follow TDD cycle. Test proves fix and prevents regression.
Never fix bugs without a test.
Production code → test exists and failed first Otherwise → not TDD
No exceptions without your human partner's permission.
Curated power-ups for coding agents: skills, slash commands, MCP configs, hooks, AGENTS.md templates, and workflows for serious software engineering. Claude Code, Codex, Antigravity CLI, Cursor and more
Repo: yeaight7/agent-powerups
Use when evaluating prompts, LLM outputs, red-team suites, or model behavior with local eval configs and safe provider/cost controls.
Use when creating or reviewing red-team eval plugins, attack templates, grader rubrics, safety fixtures, or model-risk test metadata.
Use when designing, running, debugging, or hardening deterministic eval suites for agent skills, prompts, tool workflows, or MCP-backed cases.
Use when designing tool definitions for a new agent or subagent, an agent shows high retry rates, ambiguous tool invocations, or silent failures, or an…
Use when routing a prompt to a local provider CLI for a second opinion, review, or plan -- you are about to call a provider directly, need the response saved…
Use when starting work in an unfamiliar area of a codebase, spawning a subagent that needs targeted file context, a first search pass missed the relevant file,…