Skip to content
Development
Skill

/reality-check

Check whether a claimed shipped feature, repo state or goal status holds up in evidence. Use when: comparing a claim with what exists; a gap report is not a verdict.

From plugin
agentops
44234 skills7 agents1 hook
Install
$ npx -y skills add boshu2/agentops --skill reality-check --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/reality-check

Context preview

The summary Claude sees to decide when to auto-load this skill.

Check whether a claimed shipped feature, repo state or goal status holds up in evidence. Use when: comparing a claim with what exists; a gap report is not a verdict.

SKILL.md

reality-check.SKILL.md
name: reality-check
description: 'Check whether a claimed shipped feature, repo state or goal status holds up in evidence. Use when: comparing a claim with what exists; a gap report is not a verdict.'
practices: [design-by-contract, evidence-based-engineering]
hexagonal_role: domain
consumes: [caller-question, native-source-evidence]
produces: [reality-check-report.v1, goal-measurement-report, native-status-snapshot]
context_rel:
- kind: supplier-to
  with: plan
skill_api_version: 1
user-invocable: true
metadata:
  tier: judgment
  dependencies: []
  capabilities: [compare_claim_to_evidence, measure_declared_goals, report_native_status]
  effects: [write_advisory_gap_report, write_goal_snapshot, write_requested_rendered_spec]
  canonical_status: canonical
  disposition: keep_strategy
output_contract: cited claim comparison; validated reality-check-report.v1 for durable gap reports; measured goal results or observable native status

Reality Check

Compare an expected state with observable evidence, measure declared goals, or report native status. Select the requested question; a snapshot needs no invented completion claim. Return facts and gaps without selecting work.

Claim comparison

1. Read the exact claim and its source. For a completion claim, enumerate every stated goal, including work that was never started. Give each a disposition: confirmed with evidence, concrete gap or unverifiable. 2. Inspect relevant files, command outcomes and artifacts. Separate confirmed behavior, concrete gaps, incomplete evidence and changed assumptions. Name the missing evidence instead of resolving an untestable claim by assertion. 3. Compare proposed scope with the original goal when asked about a plan. Report additions that lack authority as scope escalation; the report cannot approve them. Repeated measurements use the same question and criteria; a changed question starts a different comparison. 4. Return the cited findings with checked and not-checked scope. Keep native tracker, Git, runtime, deterministic checks and semantic judgments distinct.

A quick answer can be inline. A selected durable gap report retains `reality-check-report.v1`: write `reality-check-report.json` under the caller's chosen destination, default `.agents/scratch/reality-check/<run-id>/`, and run `skills/reality-check/scripts/validate-output.sh <report.json>`. Include the checked claim, evidence-backed finding kinds and goal-by-goal dispositions for completion/status claims. This format permits no `verdict`, `readiness` or `PASS` field; observations are not independent semantic judgment.

Goal measurement

Inspect the declared goals source; prefer `GOALS.md` when it and legacy YAML both exist. Preserve directive and gate identities and report each executable check with its actual outcome. Run the requested `ao goals` command once: `measure --json`, `validate --json`, `drift`, `history`, `export`, `meta --json`, `scenarios` or `render`.

These commands do not edit the goals source, but `measure`, `drift` and `export` may write best-effort derived snapshots under `.agents/ao/goals/baselines/`. `render --out <file>` writes a caller-selected spec; never target the goals source or another non-derived file. Use stdout when no output file is requested. Return command, exit code, goal-level results, aggregate measurement, missing evidence and checked/not-checked scope. Do not add, remove, prioritize, migrate or repair goals, or turn a measurement gap into assigned work.

Native status

Use `ao status` for the local evidence-store view. It validates content-addressed intent and verdict artifacts before counting them, reports corruption or unavailable sources, and shows evidence recency. Its durable stores are `.agents/ao/intents/sha256` and `.agents/ao/verdicts/sha256`; a count is not a per-artifact digest inventory. Inspect a specific digest or timestamp only when that artifact is part of the requested question.

Report caller-supplied subject manifests from their named location. Otherwise mark manifests, runtime phase, elapsed execution, tool-call activity and remaining work as not checked. An artifact's recent timestamp proves evidence recency, not an active worker. Read other tracker, Git or factory facts only from their own authorized source; do not blend factory completion, green checks and a fresh verdict into one health judgment. Report unavailable evidence explicitly.

Boundary

Return the selected report or snapshot. This skill neither changes native state nor issues semantic PASS, repairs records, schedules or retries work. The documented goal snapshots and requested report/spec writes are its only output side effects. A native caller pursuing an authorized outcome uses these facts and continues its work; the reporting mode does not decide completion for it.

Read more
Ships withagentops

Agent work you can verify and build on. AgentOps means agent operations: applying years of DevOps experience to how coding agents plan, implement, validate, and hand off work.

Get the whole plugin

Other skills on agentops.

cass
Skill

cass

Search agent session logs and cited episodes with CASS. Use when: past prompts, decisions or failures may answer a question; repeated text is not a proven…

@boshu2@boshu2View Skill
cc-hooks
Skill

cc-hooks

Configure Claude Code hooks and narrow enforcement guards. Use when: the caller requests hook installation, repair or policy changes; a hook is not required to…

@boshu2@boshu2View Skill