/app-user-story-qa
End-to-end app feature inventory and user-story testing workflow with a canonical tracker. Use when the user asks to audit every feature, derive expected behavior from code, test user journeys, or explicitly fix and retest documented UX or logistical defects.
$ npx -y skills add majiayu000/spellbook --skill app-user-story-qa --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/app-user-story-qa
Context preview
The summary Claude sees to decide when to auto-load this skill.
End-to-end app feature inventory and user-story testing workflow with a canonical tracker. Use when the user asks to audit every feature, derive expected behavior from code, test user journeys, or explicitly fix and retest documented UX or logistical defects.
SKILL.md
app-user-story-qa.SKILL.mdname: app-user-story-qa
description: "End-to-end app feature inventory and user-story testing workflow with a canonical tracker. Use when the user asks to audit every feature, derive expected behavior from code, test user journeys, or explicitly fix and retest documented UX or logistical defects."
App User Story QA
Purpose
Use this skill to turn a broad "test every feature in this app" request into a controlled loop with one source of truth. Anchor every feature to code, track expected behavior in a canonical spreadsheet, run tests against each user story, and classify failures. Apply fixes only when the current request explicitly authorizes them.
Route
Use `plan_first` for most runs. Switch to `clarify_first` only when the app boundary, writable checkout, production risk, or acceptance criteria are unclear enough that a wrong assumption would waste substantial work.
Before editing:
1. Run a repo state snapshot: current path, branch, latest remote, dirty files, open PRs when relevant. 2. If the checkout is dirty or user work is present, first decide which state the user asked to test. Preserve current user work in the isolated test worktree when the request targets the working tree; use a clean base only when the user asked for the base branch. 3. Read applicable repo instructions such as `AGENTS.md`, `README`, architecture docs, and feature entrypoints. 4. Search for existing QA trackers before creating a new one. 5. Define the app boundary by user-facing surfaces, not internal helper functions.
Operating Contract
Select one mode before editing production code:
- `report_only` is the default for audit, inventory, test, diagnose, or tracker requests. It may create or update the requested canonical tracker and run tests, but it must not change product behavior.
- `apply_fixes` requires the current user request to explicitly ask for fixes. It covers only defects already reproduced and classified within the agreed app boundary. Earlier approval and a generic request to "test everything" do not authorize fixes.
If the mode is ambiguous, use `report_only` and record proposed fixes in the tracker.
Direct actions:
- Inventory local code and docs, create or update the single canonical tracker, run local tests, and document failures.
- In `apply_fixes` only, fix narrow in-scope logistical or UX defects and add focused coverage when user-observable behavior changes.
Escalate before:
- Testing production systems, using paid external services, changing credentials or deployment state, deleting user data, or broadening the app boundary beyond the user's request.
- Editing shared test infrastructure, weakening assertions, or changing product requirements instead of the implementation.
Evidence-backed pushback:
- Challenge requests to skip the tracker, mark untested rows as passed, or fix symptoms without reproducing the failure. Cite the tracker row, command output, code anchor, or missing environment precondition.
Feedback loop:
- Promote repeated setup failures, brittle manual paths, or recurring test gaps into tracker notes, scripts, docs, or follow-up issues instead of leaving them only in chat.
Gotchas
- Do not create scattered notes or duplicate trackers. One canonical tracker is the audit log.
- Do not invent expected behavior from product hopes. Expected behavior comes from current code, docs, and visible user surfaces.
- Do not mark a feature passed from stale output. Fresh evidence from the current session is required.
- Do not "fix" environment blockers unless the user asked for environment repair.
- Do not infer fix authorization from words such as "audit", "test", "check", or "create a tracker". Keep those runs in `report_only`.
Canonical Tracker
Create or update exactly one tracker. Prefer a real spreadsheet when the runtime supports it; otherwise use a CSV and treat it as the canonical spreadsheet. Do not scatter status across side reports.
Use these columns:
Feature ID,Surface,Feature / capability,User story,Expected behavior based on code,Code anchors,Initial test approach,Initial test command or route,Status,Test result,Errors,Fix status,Retest result,Notes
Rules:
- Assign stable IDs such as `F001`, `F002`.
- Require at least one code anchor per row. A row without an anchor is not complete.
- Write expected behavior from current code and docs, not from aspirational specs.
- Keep the tracker parseable: validate CSV/spreadsheet row widths after edits.
- Update the tracker during the loop, not only at the end.
Status Values
Use a small, consistent status set:
- `Not Tested`
- `Passed`
- `Failed - Product`
- `Failed - UX`
- `Failed - Logistical`
- `Failed - Test Infra`
- `Blocked - Env`
- `Fixed`
- `Passed after fix`
Use `Failed - Logistical` for repo-owned setup, script, port, packaging, or local workflow defects that block a valid user path. Use `Blocked - Env` for local machine issues such as missing credentials, occupied services outside the repo, stale PATH binaries, or unavailable optional runtimes. Do not "fix" the user's environment unless they explicitly ask.
Defect Taxonomy
Classify every failure before fixing:
- `Product`: implemented behavior violates the user story or loses data.
- `UX`: behavior works but the user path is confusing, brittle, or poorly surfaced.
- `Logistical`: scripts, ports, setup, packaging, or local workflow make valid behavior hard to exercise.
- `Test Infra`: the test itself is flaky, racy, or asserts the wrong contract.
- `Env`: external setup blocks execution and is not a repo defect.
Fix only defects that are in scope for the task. Record out-of-scope defects in the tracker with clear rationale.
Workflow
1. Inventory
Map all user-visible surfaces first:
- CLI commands and flags
- Web or desktop UI routes
- API endpoints
- background jobs or hooks that affect users
- plugin, package, install, or release surfaces
- import/export, backup, migration, and governance flows
Read more
name: app-user-story-qa description: "End-to-end app feature inventory and user-story testing workflow with a canonical tracker. Use when the user asks to audit every feature, derive expected behavior from code, test user journeys, or explicitly fix and retest documented UX or logistical defects."
App User Story QA
Purpose
Use this skill to turn a broad "test every feature in this app" request into a controlled loop with one source of truth. Anchor every feature to code, track expected behavior in a canonical spreadsheet, run tests against each user story, and classify failures. Apply fixes only when the current request explicitly authorizes them.
Route
Use `plan_first` for most runs. Switch to `clarify_first` only when the app boundary, writable checkout, production risk, or acceptance criteria are unclear enough that a wrong assumption would waste substantial work.
Before editing:
1. Run a repo state snapshot: current path, branch, latest remote, dirty files, open PRs when relevant. 2. If the checkout is dirty or user work is present, first decide which state the user asked to test. Preserve current user work in the isolated test worktree when the request targets the working tree; use a clean base only when the user asked for the base branch. 3. Read applicable repo instructions such as `AGENTS.md`, `README`, architecture docs, and feature entrypoints. 4. Search for existing QA trackers before creating a new one. 5. Define the app boundary by user-facing surfaces, not internal helper functions.
Operating Contract
Select one mode before editing production code:
- `report_only` is the default for audit, inventory, test, diagnose, or tracker requests. It may create or update the requested canonical tracker and run tests, but it must not change product behavior.
- `apply_fixes` requires the current user request to explicitly ask for fixes. It covers only defects already reproduced and classified within the agreed app boundary. Earlier approval and a generic request to "test everything" do not authorize fixes.
If the mode is ambiguous, use `report_only` and record proposed fixes in the tracker.
Direct actions:
- Inventory local code and docs, create or update the single canonical tracker, run local tests, and document failures.
- In `apply_fixes` only, fix narrow in-scope logistical or UX defects and add focused coverage when user-observable behavior changes.
Escalate before:
- Testing production systems, using paid external services, changing credentials or deployment state, deleting user data, or broadening the app boundary beyond the user's request.
- Editing shared test infrastructure, weakening assertions, or changing product requirements instead of the implementation.
Evidence-backed pushback:
- Challenge requests to skip the tracker, mark untested rows as passed, or fix symptoms without reproducing the failure. Cite the tracker row, command output, code anchor, or missing environment precondition.
Feedback loop:
- Promote repeated setup failures, brittle manual paths, or recurring test gaps into tracker notes, scripts, docs, or follow-up issues instead of leaving them only in chat.
Gotchas
- Do not create scattered notes or duplicate trackers. One canonical tracker is the audit log.
- Do not invent expected behavior from product hopes. Expected behavior comes from current code, docs, and visible user surfaces.
- Do not mark a feature passed from stale output. Fresh evidence from the current session is required.
- Do not "fix" environment blockers unless the user asked for environment repair.
- Do not infer fix authorization from words such as "audit", "test", "check", or "create a tracker". Keep those runs in `report_only`.
Canonical Tracker
Create or update exactly one tracker. Prefer a real spreadsheet when the runtime supports it; otherwise use a CSV and treat it as the canonical spreadsheet. Do not scatter status across side reports.
Use these columns:
Feature ID,Surface,Feature / capability,User story,Expected behavior based on code,Code anchors,Initial test approach,Initial test command or route,Status,Test result,Errors,Fix status,Retest result,Notes
Rules:
- Assign stable IDs such as `F001`, `F002`.
- Require at least one code anchor per row. A row without an anchor is not complete.
- Write expected behavior from current code and docs, not from aspirational specs.
- Keep the tracker parseable: validate CSV/spreadsheet row widths after edits.
- Update the tracker during the loop, not only at the end.
Status Values
Use a small, consistent status set:
- `Not Tested`
- `Passed`
- `Failed - Product`
- `Failed - UX`
- `Failed - Logistical`
- `Failed - Test Infra`
- `Blocked - Env`
- `Fixed`
- `Passed after fix`
Use `Failed - Logistical` for repo-owned setup, script, port, packaging, or local workflow defects that block a valid user path. Use `Blocked - Env` for local machine issues such as missing credentials, occupied services outside the repo, stale PATH binaries, or unavailable optional runtimes. Do not "fix" the user's environment unless they explicitly ask.
Defect Taxonomy
Classify every failure before fixing:
- `Product`: implemented behavior violates the user story or loses data.
- `UX`: behavior works but the user path is confusing, brittle, or poorly surfaced.
- `Logistical`: scripts, ports, setup, packaging, or local workflow make valid behavior hard to exercise.
- `Test Infra`: the test itself is flaky, racy, or asserts the wrong contract.
- `Env`: external setup blocks execution and is not a repo defect.
Fix only defects that are in scope for the task. Record out-of-scope defects in the tracker with clear rationale.
Workflow
1. Inventory
Map all user-visible surfaces first:
- CLI commands and flags
- Web or desktop UI routes
- API endpoints
- background jobs or hooks that affect users
- plugin, package, install, or release surfaces
- import/export, backup, migration, and governance flows
Cross-runtime skills for Claude Code, Codex, and multi-agent workflows.
Repo: majiayu000/spellbook
Other skills on spellbook.
- /agentsmd-optimize
Audit AND optimize a CLAUDE.md / AGENTS.md instruction file — score it against the five high-leverage patterns, flag anti-patterns, then apply approved fixes in place. Use when the user says 优化 CLAUDE.md / 优化 AGENTS.md / optimize my agent doc / 帮我改 claudemd, or after an audit
Open skill - /agentsmd-scaffold
Generate or update repository-specific AGENTS.md instruction files from real repo evidence. Use when asked to create, design, scaffold, split, or improve root or scoped AGENTS.md files for Codex/Claude/agent workflows, especially when a repo needs directory-specific rules,
Open skill - /api-design
REST/GraphQL/gRPC API design best practices. Use when designing APIs, defining contracts, handling versioning. Covers OpenAPI 3.2, GraphQL Federation, gRPC streaming.
Open skill - /app-ui-design
Mobile app UI design expert for iOS and Android. Use when designing app interfaces, creating design systems, ensuring accessibility, or following platform guidelines. Covers Material Design 3, Human Interface Guidelines, color theory, typography, and 2025 trends.
Open skill - /architecture-foundation
Design architecture foundations before implementation. Use when asked to design or refactor architecture, choose Rust/Go crate, package, module, runtime, workflow, or service boundaries, compare mature project architecture, prevent stacked one-off PRs, audit migration debt in
Open skill - /ask-opencli
Ask Grok or Gemini via opencli (browser-session based, no API cost). Use when the user says "问 grok", "问 gemini", "ask grok", "ask gemini", "用 grok 查", "grok 怎么看", "让 gemini 分析", or wants a second opinion from Grok/Gemini without paying API tokens. Routes to opencli which drives
Open skill

