Skip to content
Security
Skill

/multi-pass-self-critique

Meta-skill for /audit-strict and high-stakes audits. Run two independent passes with different starting contexts, then keep only consensus findings. Aggressively cuts false positives.

From plugin
rugproof
952 skills23 agents45 commands4 hooks
Install
$ npx -y skills add omermaksutii/RugProof --skill multi-pass-self-critique --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/multi-pass-self-critique

Context preview

The summary Claude sees to decide when to auto-load this skill.

Meta-skill for /audit-strict and high-stakes audits. Run two independent passes with different starting contexts, then keep only consensus findings. Aggressively cuts false positives.

SKILL.md

multi-pass-self-critique.SKILL.md
name: multi-pass-self-critique
description: Meta-skill for /audit-strict and high-stakes audits. Run two independent passes with different starting contexts, then keep only consensus findings. Aggressively cuts false positives.

Multi-pass self-critique (meta-skill)

This skill governs the protocol for high-precision audits where false positives are unacceptable.

When to use

  • `/audit-strict` invocations
  • Pre-launch audits where the user has explicitly opted into slower, higher-precision review
  • Re-audits where prior tools have been noisy

The protocol

Pass A — Skill-driven, bottom-up

Read the code line-by-line. Apply the vuln-skill library. Emit candidate findings with reasoning traces.

Pass B — Exploit-driven, top-down

Fresh context. Pretend you have no prior findings. Approach the contract as an attacker: "What would I steal here? What's the cheapest exploit?" Emit findings.

Compare

For each Pass-A finding, check Pass-B:

  • Did Pass B independently identify this issue (under any name)?
  • Does Pass B's exploit narrative match this issue's mechanism?

For each Pass-B finding, check Pass-A:

  • Did the skill library flag this?

Categorize

| Pass A | Pass B | Result | |---|---|---| | ✓ | ✓ | **Consensus** — Confidence HIGH, keep | | ✓ | ✗ | Single-source A — Confidence MEDIUM, keep with note | | ✗ | ✓ | Single-source B — Confidence MEDIUM, keep with note | | ✗ | ✗ | Not reported |

Synthesize

Output the consensus findings as primary, single-source findings as secondary. Be explicit about which is which.

Why this works

  • Each pass has different blind spots. Skill-based misses novel patterns; exploit-based misses subtle CWE patterns.
  • Their intersection is the *high-precision* set.
  • Their union is the *high-recall* set. Sometimes you want union; for `/audit-strict`, you want intersection.

Anti-patterns

  • **Sharing findings between passes.** Pass B must not see Pass A's findings; that defeats independence.
  • **Counting "Pass A finds X, Pass A also finds X in a different file" as consensus.** Same pass = same blind spots.
  • **Using the same model temperature for both passes.** Vary the approach, not just the seed.

Output

Multi-pass audit:

  Pass A (skill-driven) findings:     12
  Pass B (exploit-driven) findings:    9
  Consensus (both):                    7   ← HIGH confidence
  Pass-A-only (no exploit found):      5   ← MEDIUM, flagged for review
  Pass-B-only (skills missed):         2   ← MEDIUM, possibly novel patterns

Final report:
  → 7 HIGH-confidence findings (action recommended)
  → 7 MEDIUM-confidence findings (requires user judgment)

Related

  • [[confidence-scoring]] — output of this skill drives the confidence label
  • [[known-good-comparison]] — third axis: also check against reference impls
  • /audit-strict (command)
Read more
Ships withrugproof

Rugproof your code before someone else does. 🌐 Live site: omermaksutii.github.io/RugProof 📦 Latest: v1.0.0 — 45 commands · 23 agents · 45 skills · 13 MCP servers · tested, offline-first, with rule packs, a benchmark, non-EVM coverage, and post-deploy

Get the whole plugin

Other skills on rugproof.