context-surfing
Monitors context window health during large, long-running, multi-session, or explicitly…
Runs project compile, test, and lint commands between implementation and quality review. Gates simplify-and-harden behind machine verification. If checks fail, routes back to implementation with diagnostics for a fix loop. If checks pass, signals ready for the quality pass. Use
$ npx -y skills add pskoett/pskoett-ai-skills --skill verify-gate --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/verify-gateContext preview
The summary Claude sees to decide when to auto-load this skill.
Runs project compile, test, and lint commands between implementation and quality review. Gates simplify-and-harden behind machine verification. If checks fail, routes back to implementation with diagnostics for a fix loop. If checks pass, signals ready for the quality pass. Use
name: verify-gate description: "Runs project compile, test, and lint commands between implementation and quality review. Gates simplify-and-harden behind machine verification. If checks fail, routes back to implementation with diagnostics for a fix loop. If checks pass, signals ready for the quality pass. Use after any implementation work completes and before simplify-and-harden. Essential for the inner loop's verify step."
Machine verification gate between implementation and quality review. Runs the project's compile, test, and lint commands. If any fail, enters a fix loop. If all pass, unblocks simplify-and-harden.
This is the inner loop's **verify** step. Without it, the agent hands off code with zero machine signal about whether it actually works.
[implementation] → verify-gate → simplify-and-harden → self-improvement
↻ fix loop — on failure, hands diagnostics to self-healing
↳ self-healing (diagnose → patch → verify → file HEAL); verify-gate re-checksRead the project's configuration to find verification commands. Check these sources in order:
1. **Project instruction files** (CLAUDE.md, AGENTS.md, .github/copilot-instructions.md) — look for a `## Verification` or `## Test Commands` section 2. **package.json** — `scripts.test`, `scripts.lint`, `scripts.typecheck`, `scripts.build`. Also check for a `bun.lock` / `bun.lockb` alongside it → prefer `bun run <script>` over `npm run <script>` when present. Check for `pnpm-lock.yaml` → prefer `pnpm run`. Check for `yarn.lock` → prefer `yarn`. 3. **Makefile** / **Justfile** — `test`, `lint`, `check`, `build` targets 4. **Cargo.toml** — `cargo build`, `cargo test`, `cargo clippy` 5. **pyproject.toml** / **setup.cfg** — `pytest`, `mypy`, `ruff` 6. **go.mod** — `go build ./...`, `go test ./...`, `go vet ./...` 7. **deno.json** / **deno.jsonc** — `deno task <name>` for any defined tasks
If no commands are discoverable, ask the user once and suggest they add a `## Verification` section to their project instruction files (CLAUDE.md, AGENTS.md, or equivalent) for future sessions:
## Verification - Build: `npm run build` - Test: `npm test` - Lint: `npm run lint` - Type check: `npx tsc --noEmit`
Before running commands, translate each user acceptance criterion into observable evidence. For host-integrated or user-facing work, record the full identity chain:
Do not substitute intermediate proof for the requested outcome: an iframe `src` is not a rendered app, a successful build is not proof that the host loaded it, an expanded-only view does not prove the ordinary view, an instant synthetic click does not prove a human-duration press, and a queued or delivered acknowledgement is not a visible turn in the target conversation.
If the environment cannot observe a required outcome, record the exact acceptance gap. The gate cannot report `PASSED` for that criterion.
Run discovered commands in this order. Stop at the first failure category.
Run the build or type-check command. These catch structural errors before wasting time on tests.
Exit 0 → proceed to Phase 2 Exit non-zero → enter fix loop with compiler output
Run the test command. Scope to changed files if the test runner supports it.
Exit 0 → proceed to Phase 3 Exit non-zero → enter fix loop with test output
Run the lint command. Lint failures are lower severity but still worth catching.
Exit 0 → proceed to Phase 4 Exit non-zero → enter fix loop with lint output
Run configured custom verification tools after the standard phases. Each custom command must prove its stated invariant rather than merely exit successfully for unrelated input.
Exercise every acceptance check defined in Step 1 through the requested host and ordinary user path. Capture enough evidence to identify the host, workspace, loaded artifact, interaction, and observed result. A proxy or intermediate state fails this phase even when compile, tests, and lint are green.
When a phase fails:
1. **Read the output.** Parse the error output for actionable diagnostics — file paths, line numbers, error messages. 2. **Scope the fix.** Only fix what the verification caught. Do not refactor, improve, or touch unrelated code. 3. **Apply the fix.** Make the minimal change to resolve the failure. 4. **Re-run the failed phase.** Not all phases — just the one that failed. 5. **If it passes**, continue to the next phase. 6. **If it fails again**, increment the attempt counter.
A collection of skills for AI agents. Follows the Agent Skills specification and ships an Agent Plugins 1.0 portable package. This repository is my personal skill testing ground.
Monitors context window health during large, long-running, multi-session, or explicitly…
Control-plane workflow for coordinating multi-agent, multi-session project work from a single…
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Runs all eval…
[Beta] Creates permanent eval cases from promoted learnings and runs regression checks…
Frames coding-agent work sessions with explicit intent capture and drift monitoring. Use when…
[Beta] CI-only learning aggregation workflow using gh-aw (GitHub Agentic Workflows). Scans…