test-types-and-pyramid
Not all tests are equal. Choosing the right *level* for a given behavior is the difference between a suite that gives fast, reliable signal and one that is slow, flaky, and ignored.
$ npx -y skills add vanara-agents/skills --agent claude-codeHow it fires
How this agent gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
Context preview
The summary Claude sees to decide when to auto-load this agent.
Not all tests are equal. Choosing the right *level* for a given behavior is the difference between a suite that gives fast, reliable signal and one that is slow, flaky, and ignored.
Agent definition
test-types-and-pyramid.mdTest Types and the Test Pyramid
Not all tests are equal. Choosing the right *level* for a given behavior is the difference between a suite that gives fast, reliable signal and one that is slow, flaky, and ignored.
The pyramid
/\ E2E / UI few, slow, high-confidence, brittle
/ \
/----\ Integration some, medium speed, real collaborators
/ \
/--------\ Unit many, fast, isolated, cheapPush the bulk of your assertions **down** to the unit level where they are fast and stable. Use the higher levels sparingly to prove the pieces wire together, not to re-test logic that unit tests already cover. An inverted pyramid (mostly E2E) is slow and flaky.
Unit tests
- **Scope:** one function, method, or class in isolation. Collaborators that cross a process
or I/O boundary are replaced with doubles.
- **Cost:** milliseconds. Run thousands on every save.
- **Owns:** business logic, pure functions, edge cases, error branches, algorithms.
- **Rule of thumb:** if a behavior can be tested at the unit level, test it there.
Integration tests
- **Scope:** several real units working together, or one unit against a real adjacent
dependency (a database, an in-memory queue, a local HTTP server).
- **Cost:** tens to hundreds of milliseconds. Often need setup/teardown.
- **Owns:** the seams — ORM queries against a real schema, repository wiring, serialization,
framework request/response handling, transaction behavior.
- **Avoid:** re-testing branch logic here that a unit test already covers; keep these about
the *wiring*, not the *rules*.
End-to-end (E2E) tests
- **Scope:** the whole system through its real entry point — a browser, a CLI, a public API.
- **Cost:** seconds. Slow, infrastructure-heavy, the most prone to flakiness.
- **Owns:** a handful of critical user journeys (sign up, checkout, the money path).
- **Hand off:** browser-driven E2E belongs to the **e2e-playwright** skill, not this agent.
What goes where — a decision guide
| Question | Level | |---|---| | Does `applyDiscount(49.99, 10)` round correctly? | Unit | | Does an invalid percentage throw `RangeError`? | Unit | | Does the repository persist and re-read an order from Postgres? | Integration | | Does `POST /orders` return `201` with a `Location` header? | Integration | | Can a user add to cart and complete checkout in the browser? | E2E |
Coverage as a signal, not a goal
The 80% target is a floor that flags untested regions, not a trophy. 100% line coverage with assertion-free tests proves nothing. Chase **branch and behavior** coverage — every error path and boundary — over raw line percentage. Use `scripts/check-coverage.mjs` to gate the floor in CI, and read the uncovered lines to find the cases you forgot.
Read more
Test Types and the Test Pyramid
Not all tests are equal. Choosing the right *level* for a given behavior is the difference between a suite that gives fast, reliable signal and one that is slow, flaky, and ignored.
The pyramid
/\ E2E / UI few, slow, high-confidence, brittle
/ \
/----\ Integration some, medium speed, real collaborators
/ \
/--------\ Unit many, fast, isolated, cheapPush the bulk of your assertions **down** to the unit level where they are fast and stable. Use the higher levels sparingly to prove the pieces wire together, not to re-test logic that unit tests already cover. An inverted pyramid (mostly E2E) is slow and flaky.
Unit tests
- **Scope:** one function, method, or class in isolation. Collaborators that cross a process
or I/O boundary are replaced with doubles.
- **Cost:** milliseconds. Run thousands on every save.
- **Owns:** business logic, pure functions, edge cases, error branches, algorithms.
- **Rule of thumb:** if a behavior can be tested at the unit level, test it there.
Integration tests
- **Scope:** several real units working together, or one unit against a real adjacent
dependency (a database, an in-memory queue, a local HTTP server).
- **Cost:** tens to hundreds of milliseconds. Often need setup/teardown.
- **Owns:** the seams — ORM queries against a real schema, repository wiring, serialization,
framework request/response handling, transaction behavior.
- **Avoid:** re-testing branch logic here that a unit test already covers; keep these about
the *wiring*, not the *rules*.
End-to-end (E2E) tests
- **Scope:** the whole system through its real entry point — a browser, a CLI, a public API.
- **Cost:** seconds. Slow, infrastructure-heavy, the most prone to flakiness.
- **Owns:** a handful of critical user journeys (sign up, checkout, the money path).
- **Hand off:** browser-driven E2E belongs to the **e2e-playwright** skill, not this agent.
What goes where — a decision guide
| Question | Level | |---|---| | Does `applyDiscount(49.99, 10)` round correctly? | Unit | | Does an invalid percentage throw `RangeError`? | Unit | | Does the repository persist and re-read an order from Postgres? | Integration | | Does `POST /orders` return `201` with a `Location` header? | Integration | | Can a user add to cart and complete checkout in the browser? | E2E |
Coverage as a signal, not a goal
The 80% target is a floor that flags untested regions, not a trophy. 100% line coverage with assertion-free tests proves nothing. Chase **branch and behavior** coverage — every error path and boundary — over raw line percentage. Use `scripts/check-coverage.mjs` to gate the floor in CI, and read the uncovered lines to find the cases you forgot.
🐒 Free agents, skills & packs for Claude Code One subscription. An army of Claude Code agents. 30 production-grade agents, skills, and packs for Claude Code — free, Apache-2.0, install with one command.
Repo: vanara-agents/skills
Other agents on vanara-agents-skills.
- AGENT
Use when designing a new HTTP/GraphQL API or changing an existing one — modeling resources, defining endpoint contracts, choosing status codes, pagination, filtering, error envelopes, versioning, and idempotency. Produces a reviewable API contract plus an OpenAPI snippet, not
Open agent - review-notes
This shows how the api-designer agent reviews a flawed draft. Findings are severity-ranked so the implementer fixes the contract-breakers first. Severity legend: **CRITICAL** (breaks clients / data risk), **HIGH** (real bug or inconsistency), **MEDIUM** (maintainability),
Open agent - contract-and-openapi
The contract is the deliverable. Express it as an **OpenAPI 3.1** document so it is human-readable *and* machine-checkable. This reference covers how to structure that document and what `scripts/lint-openapi.mjs` enforces.
Open agent - design-checklist
Run through this before declaring an API contract done. It is ordered the way you should *design*: resources first, cross-cutting rules last. Every box is a place real APIs go wrong in production.
Open agent - versioning-and-evolution
APIs are forever once published: a consumer you've never met may depend on any field you expose. Design so you can **add without breaking**, and version explicitly when you must break.
Open agent - pr-comment-template
Copy-paste templates for leaving review comments. Keep each comment to one finding: an anchor, the problem, and the fix.
Open agent

