/incident-slo-runbook
Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication. Use when defining production reliability, preparing launch readiness, responding to an outage, writing a runbook, tuning
$ npx -y skills add majiayu000/spellbook --skill incident-slo-runbook --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/incident-slo-runbook
Context preview
The summary Claude sees to decide when to auto-load this skill.
Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication. Use when defining production reliability, preparing launch readiness, responding to an outage, writing a runbook, tuning
SKILL.md
incident-slo-runbook.SKILL.mdname: incident-slo-runbook
description: Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication. Use when defining production reliability, preparing launch readiness, responding to an outage, writing a runbook, tuning alerts, or closing the loop after an incident.
Incident SLO Runbook
Purpose
Use this skill to connect observability to action. Metrics and logs are not enough; each critical user journey needs an SLO, alert, owner, response path, and post-incident learning loop.
SLO Design
Define:
1. User journey or system capability. 2. SLI: request success, latency, freshness, durability, or job completion. 3. SLO target and measurement window. 4. Error budget and burn-rate alerts. 5. Exclusions with rationale. 6. Dashboard and data source. 7. Owner and escalation path.
Avoid vanity metrics. Prefer user-visible success and latency over internal counters unless internal counters are the only reliable proxy.
Runbook Requirements
Each runbook should include:
- Symptom and alert name.
- Impacted users or systems.
- First 5-minute checks.
- Triage decision tree.
- Mitigation steps with commands.
- Rollback or failover path.
- Escalation owner.
- Customer/support communication note.
- Postmortem trigger.
Commands must be safe to run or explicitly labeled destructive.
Incident Flow
1. Declare severity and incident commander. 2. Confirm impact from live evidence. 3. Stabilize with the lowest-risk mitigation. 4. Communicate status on a fixed cadence. 5. Preserve evidence before cleanup. 6. Write a blameless postmortem with action items and owners.
Output Shape
service_or_journey:
slo:
alerts:
dashboard_or_queries:
runbook:
escalation:
postmortem_template:
verification:
Read more
name: incident-slo-runbook description: Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication. Use when defining production reliability, preparing launch readiness, responding to an outage, writing a runbook, tuning alerts, or closing the loop after an incident.
Incident SLO Runbook
Purpose
Use this skill to connect observability to action. Metrics and logs are not enough; each critical user journey needs an SLO, alert, owner, response path, and post-incident learning loop.
SLO Design
Define:
1. User journey or system capability. 2. SLI: request success, latency, freshness, durability, or job completion. 3. SLO target and measurement window. 4. Error budget and burn-rate alerts. 5. Exclusions with rationale. 6. Dashboard and data source. 7. Owner and escalation path.
Avoid vanity metrics. Prefer user-visible success and latency over internal counters unless internal counters are the only reliable proxy.
Runbook Requirements
Each runbook should include:
- Symptom and alert name.
- Impacted users or systems.
- First 5-minute checks.
- Triage decision tree.
- Mitigation steps with commands.
- Rollback or failover path.
- Escalation owner.
- Customer/support communication note.
- Postmortem trigger.
Commands must be safe to run or explicitly labeled destructive.
Incident Flow
1. Declare severity and incident commander. 2. Confirm impact from live evidence. 3. Stabilize with the lowest-risk mitigation. 4. Communicate status on a fixed cadence. 5. Preserve evidence before cleanup. 6. Write a blameless postmortem with action items and owners.
Output Shape
service_or_journey: slo: alerts: dashboard_or_queries: runbook: escalation: postmortem_template: verification:
Cross-runtime skills for Claude Code, Codex, and multi-agent workflows.
Repo: majiayu000/spellbook
Other skills on spellbook.
- /agentsmd-optimize
Audit AND optimize a CLAUDE.md / AGENTS.md instruction file — score it against the five high-leverage patterns, flag anti-patterns, then apply approved fixes in place. Use when the user says 优化 CLAUDE.md / 优化 AGENTS.md / optimize my agent doc / 帮我改 claudemd, or after an audit
Open skill - /agentsmd-scaffold
Generate or update repository-specific AGENTS.md instruction files from real repo evidence. Use when asked to create, design, scaffold, split, or improve root or scoped AGENTS.md files for Codex/Claude/agent workflows, especially when a repo needs directory-specific rules,
Open skill - /api-design
REST/GraphQL/gRPC API design best practices. Use when designing APIs, defining contracts, handling versioning. Covers OpenAPI 3.2, GraphQL Federation, gRPC streaming.
Open skill - /app-ui-design
Mobile app UI design expert for iOS and Android. Use when designing app interfaces, creating design systems, ensuring accessibility, or following platform guidelines. Covers Material Design 3, Human Interface Guidelines, color theory, typography, and 2025 trends.
Open skill - /app-user-story-qa
End-to-end app feature inventory and user-story testing workflow with a canonical tracker. Use when the user asks to audit every feature, derive expected behavior from code, test user journeys, or explicitly fix and retest documented UX or logistical defects.
Open skill - /architecture-foundation
Design architecture foundations before implementation. Use when asked to design or refactor architecture, choose Rust/Go crate, package, module, runtime, workflow, or service boundaries, compare mature project architecture, prevent stacked one-off PRs, audit migration debt in
Open skill

