context-surfing
Monitors context window health during large, long-running, multi-session, or explicitly…
[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them. Turns failures into test cases that prevent silent regression. This is the outer loop's regress-test step. Use when a learning is promoted and has a clear pass/fail condition, or
$ npx -y skills add pskoett/pskoett-ai-skills --skill eval-creator --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/eval-creatorContext preview
The summary Claude sees to decide when to auto-load this skill.
[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them. Turns failures into test cases that prevent silent regression. This is the outer loop's regress-test step. Use when a learning is promoted and has a clear pass/fail condition, or
name: eval-creator description: "[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them. Turns failures into test cases that prevent silent regression. This is the outer loop's regress-test step. Use when a learning is promoted and has a clear pass/fail condition, or on cadence to verify promoted rules still hold."
Turns promoted learnings into permanent eval cases. Runs regression checks to verify promoted rules hold. This is the outer loop's **regress-test** step.
The blog says: "If a failure taught you something important, it should become a permanent test case. Otherwise the knowledge is still fragile."
.evals/
EVAL_INDEX.md # Index of all eval cases with status
cases/
eval-YYYYMMDD-001.md # Individual eval case
eval-YYYYMMDD-002.md
...From harness-updater or manually:
--- id: eval-YYYYMMDD-NNN pattern-key: [from learning] source: [LRN-YYYYMMDD-001, ERR-YYYYMMDD-003] promoted-rule: "[the rule text in project instruction files]" promoted-to: CLAUDE.md # or AGENTS.md, .github/copilot-instructions.md, or equivalent created: YYYY-MM-DD last-run: YYYY-MM-DD last-result: pass | fail | skip evidence-level: presence | structural | behavioral --- ## What This Tests [One sentence: what failure this eval prevents from recurring] ## Precondition [What must be true for this eval to be runnable] - File X exists - Project uses framework Y - etc. ## Verification Method [One of: grep-check, command-check, file-check, rule-check, behavior-check] ### grep-check Search for a pattern that should (or should not) exist:
target: src/**/*.ts pattern: "hardcoded-secret-pattern" expect: not_found
### command-check Run a command and check the exit code or output:
command: npm run typecheck expect_exit: 0
### file-check Verify a file or section exists:
target: CLAUDE.md # or AGENTS.md, .github/copilot-instructions.md section: "## Verification" expect: exists
### rule-check Verify a rule exists in an instruction file:
target: CLAUDE.md # or AGENTS.md, .github/copilot-instructions.md contains: "[the promoted rule text or key phrase]" expect: found
`rule-check` proves only that guidance is installed. It must not be reported as proof that an agent follows the guidance. ### behavior-check Exercise the relevant behavior with synthetic positive and negative controls:
command: bash bench/run-contract-case.sh bounded-retry expect_exit: 0 positive_control: "transient failure is retried within budget" negative_control: "deterministic failure changes input or stops"
A behavioral pass requires observable evidence for both controls. Prefer the repository's existing test, bench, or fixture mechanism; do not create a new test framework for one eval. ## Expected Result **Pass:** [What "good" looks like] **Fail:** [What regression looks like] ## Recovery Action If this eval fails: 1. [Specific step to diagnose] 2. [Specific step to fix] 3. Re-run this eval to verify
Read `.evals/EVAL_INDEX.md`, iterate through all cases, execute each verification method.
Filter to evals matching a specific pattern.
Filter to evals whose source files match an area (frontend, backend, etc.).
For each eval case:
1. **Check precondition** — if not met, mark as `skip` 2. **Execute verification method:**
require the positive control to pass and the negative control to be rejected
3. **Compare result** to expected 4. **Check evidence strength** — a presence/structural method cannot satisfy a behavioral assertion. Mark the case `fail` with `insufficient_evidence` rather than upgrading the claim. 5. **Update `last-run` and `last-result`** in the eval case file 6. **Update `EVAL_INDEX.md`** with the result
## Eval Run: YYYY-MM-DD **Total:** N evals **Passed:** N **Failed:** N **Skipped:** N ### Failures #### eval-YYYYMMDD-001 — [pattern-key] - **What regressed:** [description] - **Expected:** [X] - **Got:** [Y] - **Recovery action:** [from eval case] ### Summary [All green / N regressions need attention]
`.evals/EVAL_INDEX.md`:
# Eval Index | ID | Pattern-Key | Rule Summary | Last Run | Result | Created | |----|-------------|-------------|----------|--------|---------| | eval-YYYYMMDD-001 | auth-middleware-lock | Run migrations on test DB first | YYYY-MM-DD | pass | YYYY-MM-DD | | eval-YYYYMMDD-002 | pnpm-not-npm | Use pnpm in this repo | YYYY-MM-DD | fail | YYYY-MM-DD |
A collection of skills for AI agents. Follows the Agent Skills specification and ships an Agent Plugins 1.0 portable package. This repository is my personal skill testing ground.
Monitors context window health during large, long-running, multi-session, or explicitly…
Control-plane workflow for coordinating multi-agent, multi-session project work from a single…
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Runs all eval…
Frames coding-agent work sessions with explicit intent capture and drift monitoring. Use when…
[Beta] CI-only learning aggregation workflow using gh-aw (GitHub Agentic Workflows). Scans…
[Beta] Cross-session analysis of accumulated .learnings/ files. Reads all entries, groups by…