/strict-enforcement
Use when running claudikins-kernel:verify, checking implementation quality, deciding pass/fail verdicts, or enforcing cross-command gates — requires actual evidence of code working, not just passing tests
$ npx -y skills add elb-pr/claudikins-kernel --skill strict-enforcement --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.
- You can call itInvoke it directly when you want it.
- Slash command
/strict-enforcement
Context preview
The summary Claude sees to decide when to auto-load this skill.
Use when running claudikins-kernel:verify, checking implementation quality, deciding pass/fail verdicts, or enforcing cross-command gates — requires actual evidence of code working, not just passing tests
SKILL.md
strict-enforcement.SKILL.mdname: strict-enforcement
description: Use when running claudikins-kernel:verify, checking implementation quality, deciding pass/fail verdicts, or enforcing cross-command gates — requires actual evidence of code working, not just passing tests
allowed-tools:
- Read
- Grep
- Glob
- Bash
- WebFetch
- Skill
- mcp__plugin_claudikins-tool-executor_tool-executor__search_tools
- mcp__plugin_claudikins-tool-executor_tool-executor__get_tool_schema
- mcp__plugin_claudikins-tool-executor_tool-executor__execute_code
Strict Enforcement Verification Methodology
When to use this skill
Use this skill when you need to:
- Run the `claudikins-kernel:verify` command
- Validate implementation before shipping
- Decide pass/fail verdicts
- Check code integrity after changes
- Enforce cross-command gates
Core Philosophy
> "Evidence before assertions. Always." - Verification philosophy
Never claim code works without seeing it work. Tests passing is not enough. Claude must SEE the output.
The Three Laws
1. **See it working** - Screenshots, curl responses, CLI output. Actual evidence. 2. **Human checkpoint** - No auto-shipping. Human reviews evidence and decides. 3. **Exit code 2 gates** - Verification failures block claudikins-kernel:ship. No exceptions.
Verification Phases
Phase 1: Automated Quality Checks
Run the automated checks first. Fast feedback.
| Check | Command Pattern | What It Catches | | ----- | ------------------------------------ | ----------------------------------- | | Tests | `npm test` / `pytest` / `cargo test` | Logic errors, regressions | | Lint | `npm run lint` / `ruff` / `clippy` | Style issues, common bugs | | Types | `tsc` / `mypy` / `cargo check` | Type mismatches, interface drift | | Build | `npm run build` / `cargo build` | Compilation errors, bundling issues |
**Flaky Test Detection (C-12):**
Test fails?
├── Re-run failed tests
├── Pass 2nd time?
│ └── Yes → STOP: [Accept flakiness] [Fix tests] [Abort]
└── Fail 2nd time?
├── Run isolated
└── Still fail? → STOP: [Fix] [Skip] [Abort]Phase 2: Output Verification (catastrophiser)
This is the feedback loop that makes Claude's code actually work.
| Project Type | Verification Method | Evidence | | ------------ | ------------------------------------ | ------------------------------ | | Web app | Start server, screenshot, test flows | Screenshots, console logs | | API | Curl endpoints, check responses | Status codes, response bodies | | CLI | Run commands, capture output | stdout, stderr, exit codes | | Library | Run examples, check results | Output values, test coverage | | Service | Check logs, verify health endpoint | Log patterns, health responses |
**Fallback Hierarchy (A-3):**
If primary method unavailable, fall back:
1. Start server + screenshot (preferred for web) 2. Curl endpoints (preferred for API) 3. Run CLI commands (preferred for CLI) 4. Run tests only (fallback) 5. Code review only (last resort)
**Timeout:** 30 seconds per verification method (CMD-30).
Phase 3: Code Simplification (Optional)
After verification passes, optionally run cynic for polish.
**Prerequisites:**
- Phase 2 (catastrophiser) must PASS
- Human approves: "Run cynic for polish pass?"
**cynic Rules:**
- Preserve exact behaviour (tests MUST still pass)
- Remove unnecessary abstraction
- Improve naming clarity
- Delete dead code
- Flatten nested conditionals
**If tests fail after simplification:**
- Log failure reasons
- Show human
- Proceed anyway (A-5) with caveat
See [cynic-rollback.md](references/cynic-rollback.md) for recovery patterns.
Phase 4: Klaus Escalation
If stuck during verification:
Is mcp__claudikins-klaus available? (E-16)
├── No →
│ Offer: [Manual review] [Ask Claude differently] (E-17)
│ Fallback: [Accept with uncertainty] [Max retries, abort] (E-18)
└── Yes →
Spawn klaus via SubagentStop hookPhase 5: Human Checkpoint
The final gate. Present comprehensive evidence.
Verification Report
-------------------
Tests: ✓ 47/47 passed
Lint: ✓ 0 issues
Types: ✓ 0 errors
Build: ✓ success
Evidence:
- Screenshot: .claude/evidence/login-flow.png
- API test: POST /api/auth → 200 OK
- CLI test: mycli --help → exit 0
[Ready to Ship] [Needs Work] [Accept with Caveats]
Human decides. If approved, set `unlock_ship = true`.
Rationalizations to Resist
Agents under pressure find excuses. These are all violations:
| Excuse | Reality | | ------------------------------------------ | --------------------------------------------------------------------- | | "Tests pass, that's good enough" | Tests aren't enough. SEE it working. Screenshots, curl, output. | | "I'll verify after shipping" | Verify BEFORE ship. That's the whole point. | | "The type checker caught everything" | Types don't catch runtime issues. Get evidence. | | "Screenshot failed but it probably works" | "Probably" isn't evidence. Fix the screenshot or use fallback. | | "Human checkpoint is just a formality" | Human checkpoint is the gate. No auto-shipping. | | "Code review is enough for this change" | Code review is last resort fallback. Try harder. | | "Tests are flaky, I'll ignore the failure" | Flaky tests hide real failures. Fix or explicitly accept with caveat. | | "Exit code 2 is too strict" | Exit code 2 exists to block bad ships. Pass properly. |
**All of these mean: Get evidence. Human decides. No shortcuts.**
Red Flags — STOP and Reassess
If you're thinking any of these, you're about to viola
Read more
name: strict-enforcement description: Use when running claudikins-kernel:verify, checking implementation quality, deciding pass/fail verdicts, or enforcing cross-command gates — requires actual evidence of code working, not just passing tests allowed-tools: - Read - Grep - Glob - Bash - WebFetch - Skill - mcp__plugin_claudikins-tool-executor_tool-executor__search_tools - mcp__plugin_claudikins-tool-executor_tool-executor__get_tool_schema - mcp__plugin_claudikins-tool-executor_tool-executor__execute_code
Strict Enforcement Verification Methodology
When to use this skill
Use this skill when you need to:
- Run the `claudikins-kernel:verify` command
- Validate implementation before shipping
- Decide pass/fail verdicts
- Check code integrity after changes
- Enforce cross-command gates
Core Philosophy
> "Evidence before assertions. Always." - Verification philosophy
Never claim code works without seeing it work. Tests passing is not enough. Claude must SEE the output.
The Three Laws
1. **See it working** - Screenshots, curl responses, CLI output. Actual evidence. 2. **Human checkpoint** - No auto-shipping. Human reviews evidence and decides. 3. **Exit code 2 gates** - Verification failures block claudikins-kernel:ship. No exceptions.
Verification Phases
Phase 1: Automated Quality Checks
Run the automated checks first. Fast feedback.
| Check | Command Pattern | What It Catches | | ----- | ------------------------------------ | ----------------------------------- | | Tests | `npm test` / `pytest` / `cargo test` | Logic errors, regressions | | Lint | `npm run lint` / `ruff` / `clippy` | Style issues, common bugs | | Types | `tsc` / `mypy` / `cargo check` | Type mismatches, interface drift | | Build | `npm run build` / `cargo build` | Compilation errors, bundling issues |
**Flaky Test Detection (C-12):**
Test fails?
├── Re-run failed tests
├── Pass 2nd time?
│ └── Yes → STOP: [Accept flakiness] [Fix tests] [Abort]
└── Fail 2nd time?
├── Run isolated
└── Still fail? → STOP: [Fix] [Skip] [Abort]Phase 2: Output Verification (catastrophiser)
This is the feedback loop that makes Claude's code actually work.
| Project Type | Verification Method | Evidence | | ------------ | ------------------------------------ | ------------------------------ | | Web app | Start server, screenshot, test flows | Screenshots, console logs | | API | Curl endpoints, check responses | Status codes, response bodies | | CLI | Run commands, capture output | stdout, stderr, exit codes | | Library | Run examples, check results | Output values, test coverage | | Service | Check logs, verify health endpoint | Log patterns, health responses |
**Fallback Hierarchy (A-3):**
If primary method unavailable, fall back:
1. Start server + screenshot (preferred for web) 2. Curl endpoints (preferred for API) 3. Run CLI commands (preferred for CLI) 4. Run tests only (fallback) 5. Code review only (last resort)
**Timeout:** 30 seconds per verification method (CMD-30).
Phase 3: Code Simplification (Optional)
After verification passes, optionally run cynic for polish.
**Prerequisites:**
- Phase 2 (catastrophiser) must PASS
- Human approves: "Run cynic for polish pass?"
**cynic Rules:**
- Preserve exact behaviour (tests MUST still pass)
- Remove unnecessary abstraction
- Improve naming clarity
- Delete dead code
- Flatten nested conditionals
**If tests fail after simplification:**
- Log failure reasons
- Show human
- Proceed anyway (A-5) with caveat
See [cynic-rollback.md](references/cynic-rollback.md) for recovery patterns.
Phase 4: Klaus Escalation
If stuck during verification:
Is mcp__claudikins-klaus available? (E-16)
├── No →
│ Offer: [Manual review] [Ask Claude differently] (E-17)
│ Fallback: [Accept with uncertainty] [Max retries, abort] (E-18)
└── Yes →
Spawn klaus via SubagentStop hookPhase 5: Human Checkpoint
The final gate. Present comprehensive evidence.
Verification Report ------------------- Tests: ✓ 47/47 passed Lint: ✓ 0 issues Types: ✓ 0 errors Build: ✓ success Evidence: - Screenshot: .claude/evidence/login-flow.png - API test: POST /api/auth → 200 OK - CLI test: mycli --help → exit 0 [Ready to Ship] [Needs Work] [Accept with Caveats]
Human decides. If approved, set `unlock_ship = true`.
Rationalizations to Resist
Agents under pressure find excuses. These are all violations:
| Excuse | Reality | | ------------------------------------------ | --------------------------------------------------------------------- | | "Tests pass, that's good enough" | Tests aren't enough. SEE it working. Screenshots, curl, output. | | "I'll verify after shipping" | Verify BEFORE ship. That's the whole point. | | "The type checker caught everything" | Types don't catch runtime issues. Get evidence. | | "Screenshot failed but it probably works" | "Probably" isn't evidence. Fix the screenshot or use fallback. | | "Human checkpoint is just a formality" | Human checkpoint is the gate. No auto-shipping. | | "Code review is enough for this change" | Code review is last resort fallback. Try harder. | | "Tests are flaky, I'll ignore the failure" | Flaky tests hide real failures. Fix or explicitly accept with caveat. | | "Exit code 2 is too strict" | Exit code 2 exists to block bad ships. Pass properly. |
**All of these mean: Get evidence. Human decides. No shortcuts.**
Red Flags — STOP and Reassess
If you're thinking any of these, you're about to viola
Showing the first part of this file.
SRE thinking applied to Claude Code, based on Boris Cherny's Q&A. It enforces a strict 4-stage pipeline with gates between each step. You literally cannot skip verification. You cannot ship without approval.
Other skills on claudikins-kernel.
- /brain-jam-plan
Use when running claudikins-kernel:outline, brainstorming implementation approaches, gathering requirements iteratively, structuring complex technical plans, or facing analysis paralysis with too many options — provides iterative human-in-the-loop planning with explicit
Open skill - /git-workflow
Use when running claudikins-kernel:execute, decomposing plans into tasks, setting up two-stage review, deciding batch sizes, or handling stuck agents — enforces isolation, verification, and human checkpoints; prevents runaway parallelization and context death
Open skill - /shipping-methodology
Use when running claudikins-kernel:ship, preparing PRs, writing changelogs, deciding merge strategy, or handling CI failures — enforces GRFP-style iterative approval, code integrity validation, and human-gated merges
Open skill

