Skip to content

/qa

General-purpose QA verdict for any artifact type

From plugin
ouroboros
5.8k22 skills21 agents3 hooks1 MCP
Install
$ npx -y skills add Q00/ouroboros --skill qa --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/qa

Context preview

The summary Claude sees to decide when to auto-load this skill.

General-purpose QA verdict for any artifact type

SKILL.md

qa.SKILL.md
name: qa
description: "General-purpose QA verdict for any artifact type"

/ouroboros:qa

Standalone quality assessment for any artifact — code, documents, API responses, test output, or custom content. Unlike `ooo evaluate` (3-stage formal verification pipeline), `ooo qa` is a fast single-pass verdict with actionable suggestions.

Usage

ooo qa [file_path | artifact_text]
ooo qa                                     # evaluate recent execution output
/ouroboros:qa [file_path | artifact_text]   # plugin mode

**Trigger keywords:** "ooo qa", "qa check", "quality check"

How It Works

The QA Judge evaluates an artifact against a quality bar and returns a structured verdict:

1. **Parse the Quality Bar** — What EXACTLY must be true to pass? 2. **Assess Dimensions** — Correctness, Completeness, Quality, Intent Alignment, Domain-Specific 3. **Render Verdict** — Score (0.0-1.0) with PASS / REVISE / FAIL 4. **Determine Loop Action** — `done` (pass), `continue` (revise), `escalate` (fail)

Verdict Thresholds

| Score Range | Verdict | Loop Action | |--------------|---------|-------------| | >= 0.80 | PASS | done | | 0.40 - 0.79 | REVISE | continue | | < 0.40 | FAIL | escalate |

Instructions

When the user invokes this skill:

Step 0: Determine execution mode

This skill works in two modes. Determine which one **before** attempting any tool calls:

  • **MCP mode** — If the QA MCP tool is available (already exposed, or loadable via discovery), use it:
  tool discovery query: "+ouroboros qa"

If found (typically named `mcp__plugin_ouroboros_ouroboros__ouroboros_qa`), proceed with **QA Steps** below.

  • **Fallback mode** — Only if the QA MCP tool is genuinely absent (no Ouroboros MCP server) skip to the **Fallback** section; an empty discovery result for an already-exposed tool is expected — call it directly rather than falling back. This skill is designed to work without MCP setup.

QA Steps (MCP mode)

1. **Determine the artifact to evaluate:**

  • If user provides a file path: Read the file with Read tool
  • If user provides inline text: Use that directly
  • If no artifact specified: Look for the most recent execution output in conversation context
  • Ask user if unclear what to evaluate

2. **Determine the quality bar:**

  • If a seed YAML is available in context: Extract acceptance criteria from it
  • If user specifies a quality bar: Use that
  • If neither: Ask the user "What does 'good' mean for this artifact?"

3. **Determine artifact type:**

  • `code` — source code files
  • `test_output` — test results, CI output
  • `document` — specs, docs, READMEs
  • `api_response` — API responses, JSON payloads
  • `screenshot` — visual artifacts
  • `custom` — anything else

3.5. **Acting verification fan-out — probe in parallel, then judge (do not skip for behaviour-bearing artifacts):** A text judge can be fooled by a hopeful log line. When the artifact actually *does* something (code, an app, an API, a UI), fan out empirical probes using the host's native parallel sub-agent primitive — one probe sub-agent per acting modality the runtime actually exposes, all spawned **in the same message so they run concurrently**:

  • **process probe** (`Bash`/shell): run the command / start the app / run

the declared smoke commands with bounded timeouts; capture exit codes and real output.

  • **browser probe** (browser-use tools, when the artifact serves HTTP or is

a web UI): load it, click the primary flows, capture what actually renders and any console/network errors.

  • **computer-use probe** (desktop computer-use tools, when the artifact is

a GUI/TUI): drive it like a user, screenshot the observed states.

  • **artifact probe** (file reads): verify declared files/paths exist with

real content, not placeholders. Each probe returns structured evidence only — commands run, observed effects, screenshots/paths, pass/fail per probed behaviour. Every probe also hits the applicable adversarial classes (the QA tool lists them): `misleading_output` (claimed success vs. real effect), `hung_command` (bounded timeout?), `malformed_input`, `stale_state`, `dirty_worktree`. Skip a modality only when its tools are absent or the artifact type makes it meaningless — and say which modalities were skipped and why.

Await all probes, then pass the merged evidence into the judge as `reference` (prefer observed behaviour over source text as the `artifact` when they disagree). **Empirical evidence outranks the judge**: if the judge scores PASS but any probe observed the behaviour failing, present the verdict as REVISE/FAIL on that evidence and say so explicitly — a score contradicted by observation is not a pass. If no acting tools are available at all, judge on the text alone but flag that behaviour was not observed.

4. **Call the `ouroboros_qa` MCP tool:**

   Tool: ouroboros_qa
   Arguments:
     artifact: <the content to evaluate>
     quality_bar: <what 'pass' means>
     artifact_type: "code"  (or other type)
     reference: <observed-behaviour evidence from step 3.5, plus any reference>
     pass_threshold: 0.80  (adjustable)
     seed_content: <seed YAML if available>

5. **Present results clearly:**

  • Show the score and verdict prominently
  • List dimension scores
  • Highlight specific differences found
  • Show actionable suggestions
  • End with next step guidance based on verdict:
  • **PASS (done)**: `Next: Your artifact meets the quality bar. Proceed with confidence.`
  • **REVISE (continue)**: `Next: Address the suggestions above, then run ooo qa again to re-check.`
  • **FAIL (escalate)**: `Next: Fundamental issues detected. Consider ooo interview to re-examine requirements, or ooo unstuck to challenge assumptions.`

Iterative QA Loop

For iterative usage, track the `qa_session_id` and `it

Read more
Ships withouroboros

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot, Kiro and more.

Get the whole plugin, auto-invoked

Other skills on ouroboros.