Skip to content
Agent Orchestration
Skill

/ulw-qa

[omh] Hostile scenario testing: adversarial QA and fix loops. Use when the user says: ultraqa, adversarial qa, hostile scenarios, e2e qa, real-world qa, qa scenario, release qa, 敵対的QA.

BOOST
From plugin
oh-my-hermes
3.2k145 skills
Install
$ npx -y skills add rlaope/oh-my-hermes --skill ulw-qa --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/ulw-qa

Context preview

The summary Claude sees to decide when to auto-load this skill.

[omh] Hostile scenario testing: adversarial QA and fix loops. Use when the user says: ultraqa, adversarial qa, hostile scenarios, e2e qa, real-world qa, qa scenario, release qa, 敵対的QA.

SKILL.md

ulw-qa.SKILL.md
name: "ulw-qa"
description: "[omh] Hostile scenario testing: adversarial QA and fix loops. Use when the user says: ultraqa, adversarial qa, hostile scenarios, e2e qa, real-world qa, qa scenario, release qa, 敵対的QA."
metadata:
  hermes:
    tags: [workflow, oh-my-hermes, verification]
    category: verification
    phase: qa
    role: reviewer
    quality_tier: scenario-gated

Ultraqa

This is a Hermes-native `ultraqa` workflow skill.

Why This Exists

`ultraqa` exists to keep `verification` work explicit, evidence-backed, and inside the Hermes/executor boundary instead of relying on ad hoc chat narration.

Do Not Use When

  • The request is casual chat, a status-only acknowledgement, or another workflow has stronger routing evidence.
  • The user needs implementation, review, CI, merge, or external publishing evidence that has not been delegated or observed.

Examples

Good example:

  • Prompt: $ultraqa test the setup wizard with hostile install paths, stale config, and missing PATH cases.
  • Expected behavior: Generate adversarial QA scenarios, expected signals, observed results, and fix-or-retry routing.
  • Why: The request asks for verification pressure and hostile scenarios.

Bad example:

  • Prompt: ultraqa: treat casual chat or unaccepted work as if this workflow already produced verified results.
  • Expected behavior: Ask a clarification question or route to a narrower workflow instead of forcing `ultraqa`.
  • Why: The request lacks the required inputs or would overclaim work that Hermes did not observe.

Completion Checklist

  • The scenario, expected behavior, observed result, and pass/fail basis are named.
  • Proposed fixes are separated from observed QA evidence.
  • Missing or failed verification routes back to plan, fix, or a narrower test.

Recovery Notes

  • If the expected behavior is unclear, route back to plan before running adversarial checks.
  • If verification fails, return to fix or research with the failed signal instead of advancing.

Workflow Lane

  • Current lane: **Coding handoff** (`idea-to-deploy`, `llm-app-dev`, `cto-loop`, `deploy-and-monitor`, `code-review`, `build-failure-triage`, `verification-gate`, `security-safety-review`, `+28 more`) - coding owners, handoffs, review, CI, and merge evidence.
  • If intent belongs to another lane, hand back to `oh-my-hermes` or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: `omh-routing/references/skill-common-rail.md`.

Use When

Use when the task needs adversarial test scenarios, verification, and fix loops.

Strong routing signals: `ultraqa`, `$ultraqa`, `adversarial qa`, `hostile scenarios`, `e2e qa`, `real-world qa`, `qa scenario`, `release qa`, `敵対的QA`, `リリース前QA`, `障害シナリオ`, `장애 상황`, `쿠버네티스 장애`, `적절히 진단`, `검증 체크리스트`, `릴리즈 전 gate`, `对抗式测试`, `发布前测试`, `故障场景`

Catalog Metadata

Category: `verification` Phase: `qa` Hermes role: `reviewer` Quality tier: `scenario-gated` Reasoning demand: `standard`

Quality bar:

  • Do not start this engine as an automatic continuation of another skill's output: an accepted plan, a clarified brief, or a routing recommendation is planning evidence, not permission. Unless the user explicitly invoked this engine themselves, restate in one line what will start (engine, scope, selected executor) and wait for the user's explicit go-ahead first.
  • A mid-run user message is an interjection, not a stop: answer it briefly and, in the same reply, continue the run — re-read the phase todo when one is active and dispatch or advance the next pending step, or name the armed wait it is waiting on -- handle, bound completion signal, deadline -- instead of re-reading status. Only the user's explicit stop or cancel, or the engine's own completion gate, ends the run; when the interjection changes scope, say so and update the declared plan or todo instead of silently abandoning it. A mid-run message is the latest steering for the active task, not automatically a replacement objective: it replaces the objective when the user says so and steers the current one otherwise.
  • A follow-up that needs new authority, materially expands the scope, or changes external state not already authorized is described first and started only on the user's approval: the turn ends by naming that next action and asking whether to take it, as one question carrying the choices the user has, never by declaring what will not be done; persistence never broadens the authorized scope. A refused escalation is answered the same way, with a safer alternative inside the boundary or the authorization the boundary asks for — never a workaround or an indirect execution.
  • The closing brief scales to the change: one or two sentences plus the observed validation for a simple change, more only when the complexity earns it. Lead with the result or decision, in the user's words; omit abandoned approaches unless they explain a tradeoff the reader needs; narrate no internal bookkeeping (todo transitions, waits). When the work stops at a boundary or at a decision the user owns, end with the next action offered as a question, and state what was left undone as the option it leaves open, never as a refusal. Required closing lines stay outside this scaling: the observed run summary, and any prepared-not-observed or unmerged work, are stated whatever the brief's length.
  • Generate hostile scenarios from changed behavior and known risk areas.
  • Report pass/fail evidence separately from proposed fixes.
  • For native `omh_todo` checkpoints, load the todo-checklist closing recipe; `record` then `recall` this qa declaration. Stored declarations are not proof.
  • Delegate code mutations discovered by QA to the selected coding executor.
  • A check no automation can hold gets a manual test guide, never a done claim: load `references/manual-test-guide.md` for the step shape (setup, do, expect, broken), the four reasons a step may stay manual, and why it stays `prepared_not_observed` until a run records a fre
Read more
Ships withoh-my-hermes

English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.

Get the whole plugin
Stats
3,206
Stars
244
Forks
Active
Maintenance
Python
Language
MIT
License
4h ago
Last commit
4mo ago
Created
9h ago
Added

Repo: rlaope/oh-my-hermes

Other skills on oh-my-hermes.