a11y
Accessibility audit + auto-fix (WCAG 2.2 A/AA). Scans built/static HTML for screen-reader, keyboard, and structure failures, fixes the deterministic ones, and…
Production Incident Commander — diagnose and recover from production incidents. Use when something is broken in production, site is down, errors spiking, or user reports a critical bug.
$ npx -y skills add Houseofmvps/ultraship --skill rescue --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/rescueContext preview
The summary Claude sees to decide when to auto-load this skill.
Production Incident Commander — diagnose and recover from production incidents. Use when something is broken in production, site is down, errors spiking, or user reports a critical bug.
name: rescue description: "Production Incident Commander — diagnose and recover from production incidents. Use when something is broken in production, site is down, errors spiking, or user reports a critical bug." argument-hint: "<url-or-error-description>"
When production is down, every minute costs trust. This skill runs an incident like a principal SRE — fast triage, clear decision-making, structured recovery, and prevention so it never happens again.
Before doing anything, classify the incident:
| Severity | Definition | Response Time | Example | |---|---|---|---| | **SEV-1** | Service completely down, all users affected | Immediately | Site returns 500, database unreachable | | **SEV-2** | Major feature broken, many users affected | Within 15 min | Auth broken, payments failing, data loss | | **SEV-3** | Minor feature broken, some users affected | Within 1 hour | One API endpoint slow, email not sending | | **SEV-4** | Cosmetic or edge case | Next business day | UI glitch on one browser, non-critical error log |
Severity determines urgency. SEV-1/2: restore first, investigate later. SEV-3/4: investigate first, then fix.
Ask for: 1. **Production URL** (if not already known) 2. **What's happening?** (down, slow, errors, specific feature broken) 3. **When did it start?** (narrows the commit search window) 4. **What changed recently?** (deploy, config change, dependency update, traffic spike)
If the user is panicking, skip questions and use whatever info is available. Speed > completeness for SEV-1.
node ${CLAUDE_PLUGIN_ROOT}/tools/incident-commander.mjs <project-directory> --url=<production-url>Parse the JSON output.
**If a Sentry MCP server is connected** (check your available tools — search for `sentry` tools), pull live production errors before guessing at code: list the most recent / most frequent issues since the incident window, read the top stack traces, and map each frame back to a file and line in this repo. A real stack trace from production beats inferring the culprit from recent commits. Use the actual error signature to narrow the suspect commit. If no Sentry server is connected, continue with the static diagnostics above (and mention that connecting Sentry would sharpen this step).
Present findings in order of urgency:
**Site Status:**
**Likely Culprit:**
**Error Patterns Found:**
**Resource Issues:**
Present in order of speed — for SEV-1/2, always recommend Option 1 first:
**Option 1: Rollback (fastest — 2-5 min)**
git revert <culprit-hash> --no-edit && git push
This is almost always the right first move. Restore service, then investigate.
**When NOT to rollback:**
**Option 2: Hot Fix (5-15 min)** If the error pattern is clear and the fix is small:
**Option 3: Traffic Management (immediate)** If the issue is load-related:
**Option 4: Investigate Further** If the cause isn't clear:
After applying a fix:
node ${CLAUDE_PLUGIN_ROOT}/tools/health-check.mjs <production-url>Confirm the site is back to healthy status. Check:
For SEV-1/2, the user needs to communicate with their users:
**Status page update template:**
[Investigating] We're aware of [issue description] and are actively working on a fix. [Identified] We've identified the cause and are deploying a fix. [Resolved] The issue has been resolved. [Brief explanation]. We apologize for the disruption.
**If the user has a status page:** help them post the update. **If they don't:** suggest setting up a simple one (Instatus, Betteruptime, or a static page).
Generate a post-mortem document from the incident-commander output:
# Incident Post-Mortem — [Date] ## Summary - **What happened:** [One sentence] - **Severity:** SEV-[N] - **Duration:** [start time] to [end time] ([N] minutes) - **Impact:** [Who was affected, what they experienced] - **Root cause:** [One sentence] ## Timeline | Time | Event | |---|---| | HH:MM | Issue detected (how: monitoring/user report/deploy) | | HH:MM | Investigation started | | HH:MM | Root cause identified | | HH:MM | Fix deployed | | HH:MM | Servic
"ULTRASHIP" Claude Code plugin — 39 skills, 33 tools, 11 agents for ship-ready workflows: planning, review, pentesting, safety guardrails, canary monitoring, SEO/AI-readiness check, penetration testing, code review, competitive analysis, incident response. 1 dependency. 180 tests. MIT.
Repo: Houseofmvps/ultraship
Accessibility audit + auto-fix (WCAG 2.2 A/AA). Scans built/static HTML for screen-reader, keyboard, and structure failures, fixes the deterministic ones, and…
Living Architecture Map — auto-generate Mermaid diagrams of your codebase. Use when user wants to visualize architecture, understand code structure, generate…
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent,…
Post-deploy canary monitoring — checks site health, detects regressions, monitors for errors after deployment. Use after deploying to verify production is…
Learn From the Best — analyze patterns from any codebase and apply them to yours. Use when user wants to adopt best practices from another repo, compare code…
Code review with principal-engineer-level depth. Reviews for correctness, performance, security, maintainability, and architecture. Use when completing tasks,…