Skip to content
Agent Orchestration
Skill

/omh-app-debugging

[omh] Application code misbehaves -- a wrong value, a flaky test, a lost update: reproduce it first, form competing hypotheses, discriminate them with the cheapest observation, and only then fix the demonstrated root cause. Use when the user says: app-debugging, app debugging,

BOOST
From plugin
oh-my-hermes
3.2k145 skills
Install
$ npx -y skills add rlaope/oh-my-hermes --skill omh-app-debugging --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/omh-app-debugging

Context preview

The summary Claude sees to decide when to auto-load this skill.

[omh] Application code misbehaves -- a wrong value, a flaky test, a lost update: reproduce it first, form competing hypotheses, discriminate them with the cheapest observation, and only then fix the demonstrated root cause. Use when the user says: app-debugging, app debugging,

SKILL.md

omh-app-debugging.SKILL.md
name: "omh-app-debugging"
description: "[omh] Application code misbehaves -- a wrong value, a flaky test, a lost update: reproduce it first, form competing hypotheses, discriminate them with the cheapest observation, and only then fix the demonstrated root cause. Use when the user says: app-debugging, app debugging, application debugging, root cause, root-cause, find the root cause, root cause analysis, flaky test."
metadata:
  hermes:
    tags: [workflow, oh-my-hermes, verification]
    category: verification
    phase: app-root-cause
    role: reviewer
    quality_tier: root-cause-evidence-gated

App Debugging

This is a Hermes-native `app-debugging` workflow skill.

Why This Exists

`app-debugging` exists because a wrong result in application code had no owner: `native-debugging` covers native binaries, `build-failure-triage` covers red builds, and `agent-debug` covers agent misbehaviour, so the most common debugging request dispatched straight to an execution lane that skipped the diagnosis.

First Steps

  • Ask for, or plan, the one command that shows the wrong behaviour, and record its observed output before naming a cause.
  • Refuse to prepare a fix while the reproduction reads not_observed; plan the observation that would establish it instead.

Do Not Use When

  • The fault is a crash, memory corruption, or core dump in a compiled native binary that needs a debugger session; use `native-debugging`.
  • The build, compile, or CI job fails the same way on every run; use `build-failure-triage`.
  • The subject is an agent or workflow run that misbehaved rather than the application code; use `agent-debug`.
  • The cause is already demonstrated and the request is to judge whether the fix is proven; use `verification-gate`.

Examples

Good example:

  • Prompt: a test fails one run in five in CI, how do I find out why
  • Expected behavior: Prepare reproduction_record/v1 with the loop that measures the failure rate, competing_hypotheses/v1 across ordering, shared state, timing, and environment, and the cheapest observation that splits them; no fix yet.
  • Why: The failure is intermittent, so the first deliverable is a measured reproduction rather than a patch.

Bad example:

  • Prompt: add a sleep before the assertion so the flaky test passes
  • Expected behavior: Record the sleep as a symptom mask, keep root cause open, and plan the observation that names the race.
  • Why: A timing change that hides the failure leaves the fault in place and removes the reproduction.

Completion Checklist

  • The reproduction names its command, observed output, expected output, and hit rate, and reads observed before any fix is prepared.
  • At least three hypotheses on distinct axes were written before the first observation was chosen.
  • Each observation records which hypotheses its result eliminated.
  • The root cause cites the demonstrating observation and the observation that ruled out each rival.
  • The fix handoff carries the reproduction as a regression test that fails before and passes after.

Recovery Notes

  • If the fault does not reproduce, make reproduction the first hypothesis and plan the loop, seed, or ordering that would establish it.
  • If every hypothesis is eliminated, record that, widen the axes, and keep root cause unclaimed rather than promoting the last survivor.

Workflow Lane

  • Current lane: **Coding handoff** (`idea-to-deploy`, `llm-app-dev`, `cto-loop`, `deploy-and-monitor`, `code-review`, `build-failure-triage`, `verification-gate`, `security-safety-review`, `+28 more`) - coding owners, handoffs, review, CI, and merge evidence.
  • If intent belongs to another lane, hand back to `oh-my-hermes` or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: `omh-routing/references/skill-common-rail.md`.

Use When

Use when application code -- a Python, TypeScript, Go, or JVM service, library, or test -- behaves wrongly and the cause is unknown: a wrong value, an intermittent or order-dependent test, a race or lost update, or a bug that moves when observed. The work is a demonstrated root cause: an observed reproduction, competing hypotheses, the cheapest observation that separates them, and only then a fix.

Strong routing signals: `app-debugging`, `app debugging`, `application debugging`, `root cause`, `root-cause`, `find the root cause`, `root cause analysis`, `flaky test`, `flaky tests`, `test is flaky`, `intermittent test failure`, `fails intermittently`, `fails one run in`, `passes locally but fails in ci`, `heisenbug`, `bug disappears`, `disappears when i add a print`, `race condition`, `lost update`, `update is lost`, `wrong return value`, `returns the wrong value`, `reproduce the bug`, `minimal reproduction`

Catalog Metadata

Category: `verification` Phase: `app-root-cause` Hermes role: `reviewer` Quality tier: `root-cause-evidence-gated` Reasoning demand: `standard`

Quality bar:

  • State the symptom and the expected behaviour separately from any suspected cause.
  • Record the reproduction with its hit rate; an intermittent fault is reproduced when its rate over N runs is measured, not when it happened once.
  • Load `references/hypothesis-and-race-method.md` for the hypothesis table, flaky-test tactics, and race patterns instead of improvising them.
  • Pick the next observation by cost and by how many hypotheses its result eliminates, and record the eliminations.
  • Keep reproduction, root cause, fix, and verification as separate observed states.

Handoff policy:

Keep the symptom statement, the hypothesis set, the discriminating observations, and the root-cause verdict in Hermes. Record every reproduction run, probe output, and fix verification only from executor or wrapper observed evidence.

Required inputs:

  • the wrong behaviour as observed, and the behaviour that was expected
  • the command, request, or test that shows it, and how often it shows it
  • language, framework, and what changed recently when known
  • logs, stac
Read more
Ships withoh-my-hermes

English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.

Get the whole plugin
Stats
3,206
Stars
244
Forks
Active
Maintenance
Python
Language
MIT
License
4h ago
Last commit
4mo ago
Created
9h ago
Added

Repo: rlaope/oh-my-hermes

Other skills on oh-my-hermes.