Skip to content
Agent Orchestration
Skill

/omh-live-incident-response

[omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary mitigation, verified recovery, and the customer notice. Use when the user says:

BOOST
From plugin
oh-my-hermes
3.2k145 skills
Install
$ npx -y skills add rlaope/oh-my-hermes --skill omh-live-incident-response --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/omh-live-incident-response

Context preview

The summary Claude sees to decide when to auto-load this skill.

[omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary mitigation, verified recovery, and the customer notice. Use when the user says:

SKILL.md

omh-live-incident-response.SKILL.md
name: "omh-live-incident-response"
description: "[omh] Production is down or an incident is open: command an incident that is still open -- severity as declared live state, commander and roles, an append-only timeline, a recorded temporary mitigation, verified recovery, and the customer notice. Use when the user says: live-incident-response, live incident response, incident response, active incident, ongoing incident, open incident, incident open, incident commander."
metadata:
  hermes:
    tags: [workflow, oh-my-hermes, reliability]
    category: reliability
    phase: live-incident-command
    role: operator
    quality_tier: incident-command-gated

Live Incident Response

This is a Hermes-native `live-incident-response` workflow skill.

Why This Exists

`live-incident-response` exists because an incident that is still open had no owner. `support-operations` sent an active incident to `reliability-review`, and `reliability-review` reviews incident notes after the fact, so the one skill that saw the request handed it to a postmortem while the outage was still running.

Do Not Use When

  • The incident is over and the request is the postmortem, the SLO or error-budget consequence, or remediation follow-up; use `reliability-review`.
  • The request is one customer's support case needing a reply, a severity opinion, and an escalation path, with no incident declared; use `support-operations`.
  • The request is a release being rolled out and watched -- deploy checklist, health signals, rollback criteria -- and nothing has been declared broken; use `deploy-and-monitor`.
  • The request is to send the page, publish the status-page update, or deliver the customer notice; use `connector-operator`, which records a send as observed only on a returned result.
  • The request asks whether a release is ready across rollout, rollback, and observability, before anything broke; use `production-audit`.

Examples

Good example:

  • Prompt: we have a production outage right now, declare severity and assign an incident commander
  • Expected behavior: Prepare live_incident_record/v1: ask for the user-visible symptom and blast radius, declare the severity with the observation that set it, name the commander and the remaining roles, open the append-only timeline, and state which signal decides recovery.
  • Why: The incident is open, so severity and command are live state rather than findings to review later.

Bad example:

  • Prompt: live-incident-response write up the postmortem for last week's outage and what it cost the error budget
  • Expected behavior: Route to `reliability-review`: a closed incident is reviewed, never commanded.
  • Why: Severity, roles, and a running timeline have no subject once the incident is over.

Completion Checklist

  • Severity is declared, carries the observation that set it, and every change appended rather than overwrote the previous level.
  • One commander is named; operations, communications, and scribe each name a person or read unfilled.
  • Every timeline entry is timestamped, attributed, and typed, and no earlier entry was edited.
  • Each mitigation reads temporary or permanent, and a temporary one names what removes it.
  • Recovery cites the named signal, its healthy value, the observed value, and the observer, never the mitigation alone.
  • Paging, status-page, and customer-send entries read prepared unless a connector result was observed.

Recovery Notes

  • If nobody is named commander, ask for one before anything else; an incident without a commander produces opinions instead of decisions.
  • If the recovery signal is not stated, ask which signal and which value counts as healthy before calling anything recovered.
  • If the incident turns out to be closed, hand the postmortem to `reliability-review` and leave this record as the timeline it reads.
  • If a connector call fails or returns nothing, keep the send prepared and name the channel that is unconfirmed instead of assuming delivery.

Workflow Lane

  • Current lane: **Automation and status** (`achievements`, `workspace-audit`, `production-audit`, `live-incident-response`, `automation-blueprint`, `github-event-ops`, `github-issue-intake`, `buzz`, `+39 more`) - schedules, status, health, and ops review.
  • If intent belongs to another lane, hand back to `oh-my-hermes` or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: `omh-routing/references/skill-common-rail.md`.

Use When

Use when an incident is open right now and the user needs it commanded: severity declared as live state, a commander and the other roles assigned, an append-only timeline kept, a temporary mitigation recorded as temporary, recovery verified against a named signal, and the customer notice drafted. The incident is still running; once it is closed the work is a review.

Strong routing signals: `live-incident-response`, `live incident response`, `incident response`, `active incident`, `ongoing incident`, `open incident`, `incident open`, `incident commander`, `incident command`, `incident bridge`, `incident channel`, `incident timeline`, `incident roles`, `declare severity`, `declare an incident`, `declare the incident`, `sev1`, `sev2`, `sev3`, `production outage`, `production is down`, `the site is down`, `service is down`, `we have an outage`, `outage right now`, `war room`, `stop the bleeding`, `temporary mitigation`, `page the on-call`, `page on-call`, `who is the incident commander`, `assign an incident commander`, `verify recovery`

Catalog Metadata

Category: `reliability` Phase: `live-incident-command` Hermes role: `operator` Quality tier: `incident-command-gated` Reasoning demand: `standard`

Quality bar:

  • Declare severity from the observed blast radius and record the observation that set it; an undeclared severity is not a severity, and a changed one appends rather than replaces.
  • Name one commander before anything else, then name or mark unfilled each of operations, communications,
Read more
Ships withoh-my-hermes

English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.

Get the whole plugin
Stats
3,250
Stars
244
Forks
Active
Maintenance
Python
Language
MIT
License
13h ago
Last commit
4mo ago
Created
3d ago
Added

Repo: rlaope/oh-my-hermes

Other skills on oh-my-hermes.