/eval-creator-ci
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Runs all eval cases in .evals/ on a schedule or per-PR, reports pass/fail results, and can block merges on regressions. Also creates new eval cases from promoted patterns flagged by
$ npx -y skills add pskoett/pskoett-ai-skills --skill eval-creator-ci --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/eval-creator-ci
Context preview
The summary Claude sees to decide when to auto-load this skill.
[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Runs all eval cases in .evals/ on a schedule or per-PR, reports pass/fail results, and can block merges on regressions. Also creates new eval cases from promoted patterns flagged by
SKILL.md
eval-creator-ci.SKILL.mdname: eval-creator-ci
description: "[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Runs all eval cases in .evals/ on a schedule or per-PR, reports pass/fail results, and can block merges on regressions. Also creates new eval cases from promoted patterns flagged by learning-aggregator-ci. Use when: you want automated regression testing of promoted rules in CI/headless pipelines. For interactive eval creation and runs, use eval-creator."
Eval Creator CI
Install
gh skill install pskoett/pskoett-skills eval-creator-ci
For interactive sessions, use:
gh skill install pskoett/pskoett-skills eval-creator
Fallback using the Agent Skills CLI:
npx skills add pskoett/pskoett-skills/skills/eval-creator-ci
npx skills add pskoett/pskoett-skills/skills/eval-creator
Purpose
Runs the outer loop's **regress-test** step in CI. Executes all eval cases in `.evals/`, reports pass/fail results, and optionally blocks merges on regressions. Can also create new eval cases from promotion candidates flagged by `learning-aggregator-ci`.
The interactive `eval-creator` skill is designed for in-session use where the user creates evals and runs them with immediate feedback. This CI variant runs on schedule or per-PR and posts results as check annotations.
Context Limitation (Important)
CI agents do not have implementation context. They execute mechanical verification methods (grep checks, command checks, file checks, rule checks) defined in eval case files. They do not interpret results beyond pass/fail — nuanced judgment is left to human review of the posted report.
Prerequisites
- GitHub Actions enabled on the repository
- `gh` CLI authenticated with repo access
- `gh-aw` extension installed (`gh extension install github/gh-aw`, v0.40.1+)
- `.evals/` directory with eval cases (created by `eval-creator` or `eval-creator-ci`)
- `.evals/EVAL_INDEX.md` with eval case index
CI Contract
Hard rules for headless execution:
1. **Eval execution is read-only for code** — eval cases read files and run check commands but do not modify source code 2. **Eval case creation writes to `.evals/` only** — when creating new evals from promotion candidates 3. **Headless** — no interactive prompts, no approval gates 4. **Structured output** — emit results as YAML under `eval_creator_ci` key 5. **Gate policy** — can fail the check run on eval regressions (configurable) 6. **Single comment** — post one consolidated results comment per run
Authoring Workflow (gh-aw)
1. Copy `references/workflow-example.md` into `.github/workflows/eval-creator-ci.md` 2. Customize trigger and gate policy 3. Validate: `gh aw compile` (add `--actionlint --zizmor` for security scan) 4. Push to enable
Persistence and Chaining
- **`cache-memory:`** stores eval run history (last-run dates, result trends) across runs. Avoids re-running evals that haven't changed.
- **`workflow_call:`** in create mode, triggered by `learning-aggregator-ci` via `call-workflow`. Receives promotion candidates as workflow inputs.
- **`upload-artifact:`** persists eval results YAML for downstream consumption.
Workflow Rules
The CI agent follows these rules in order:
Mode: Run Evals (default)
1. Read `.evals/EVAL_INDEX.md` to get the list of all eval cases 2. For each eval case file in `.evals/cases/`: a. Read the eval case metadata and verification method b. Check preconditions — if not met, mark as `skip` c. Execute the verification method:
- `grep-check`: Search target files for pattern, compare to expected (found/not_found)
- `command-check`: Run the command, check exit code and/or output
- `file-check`: Verify file or section exists
- `rule-check`: Read target file, search for expected content
d. Compare result to expected outcome e. Record pass/fail/skip 3. Update `.evals/EVAL_INDEX.md` with `last-run` date and `last-result` for each case 4. Emit structured YAML under key `eval_creator_ci` 5. Post results summary as a PR comment or check annotation 6. If gate policy is enabled and any eval fails: fail the check run
Mode: Create Evals (from promotion candidates)
1. Read the `learning_aggregator_ci` artifact or gap report from the most recent learning-aggregator-ci run 2. For each promotion-ready pattern with `eval_candidate: true`: a. Determine the appropriate verification method based on the pattern type b. Create the eval case file in `.evals/cases/` with proper frontmatter c. Add the entry to `.evals/EVAL_INDEX.md` 3. Commit the new eval cases (if running with write permissions) 4. Report created evals in the output
Output Schema
eval_creator_ci:
version: "0.1.0"
source:
run_id: "<workflow run ID>"
trigger: "pull_request | schedule | workflow_dispatch"
run_date: "YYYY-MM-DD"
mode: "run | create | both"
run_results:
total: 12
passed: 10
failed: 1
skipped: 1
failures:
- id: "eval-20260301-001"
pattern_key: "harden.input_validation"
rule_summary: "Always validate external inputs"
expected: "not_found"
actual: "found"
target: "src/api/handler.ts"
recovery_action: "Add input validation to new handler endpoint"
skips:
- id: "eval-20260315-003"
reason: "Precondition not met: project does not use TypeScript"
create_results:
created: 2
cases:
- id: "eval-20260411-001"
pattern_key: "simplify.dead_code"
verification_method: "grep-check"
source_learning: "LRN-20260301-001"
- id: "eval-20260411-002"
pattern_key: "harden.authorization"
verification_method: "rule-check"
source_learning: "LRN-20260315-003"
summary:
regressions: 1
new_evals_created: 2
gate_result: "fail"
followup_required: trueRecommended Outputs
| Output | Destination | Content | |--------|------------|---------| | Eval results | PR c
Read more
name: eval-creator-ci description: "[Beta] CI-only eval regression runner using gh-aw (GitHub Agentic Workflows). Runs all eval cases in .evals/ on a schedule or per-PR, reports pass/fail results, and can block merges on regressions. Also creates new eval cases from promoted patterns flagged by learning-aggregator-ci. Use when: you want automated regression testing of promoted rules in CI/headless pipelines. For interactive eval creation and runs, use eval-creator."
Eval Creator CI
Install
gh skill install pskoett/pskoett-skills eval-creator-ci
For interactive sessions, use:
gh skill install pskoett/pskoett-skills eval-creator
Fallback using the Agent Skills CLI:
npx skills add pskoett/pskoett-skills/skills/eval-creator-ci npx skills add pskoett/pskoett-skills/skills/eval-creator
Purpose
Runs the outer loop's **regress-test** step in CI. Executes all eval cases in `.evals/`, reports pass/fail results, and optionally blocks merges on regressions. Can also create new eval cases from promotion candidates flagged by `learning-aggregator-ci`.
The interactive `eval-creator` skill is designed for in-session use where the user creates evals and runs them with immediate feedback. This CI variant runs on schedule or per-PR and posts results as check annotations.
Context Limitation (Important)
CI agents do not have implementation context. They execute mechanical verification methods (grep checks, command checks, file checks, rule checks) defined in eval case files. They do not interpret results beyond pass/fail — nuanced judgment is left to human review of the posted report.
Prerequisites
- GitHub Actions enabled on the repository
- `gh` CLI authenticated with repo access
- `gh-aw` extension installed (`gh extension install github/gh-aw`, v0.40.1+)
- `.evals/` directory with eval cases (created by `eval-creator` or `eval-creator-ci`)
- `.evals/EVAL_INDEX.md` with eval case index
CI Contract
Hard rules for headless execution:
1. **Eval execution is read-only for code** — eval cases read files and run check commands but do not modify source code 2. **Eval case creation writes to `.evals/` only** — when creating new evals from promotion candidates 3. **Headless** — no interactive prompts, no approval gates 4. **Structured output** — emit results as YAML under `eval_creator_ci` key 5. **Gate policy** — can fail the check run on eval regressions (configurable) 6. **Single comment** — post one consolidated results comment per run
Authoring Workflow (gh-aw)
1. Copy `references/workflow-example.md` into `.github/workflows/eval-creator-ci.md` 2. Customize trigger and gate policy 3. Validate: `gh aw compile` (add `--actionlint --zizmor` for security scan) 4. Push to enable
Persistence and Chaining
- **`cache-memory:`** stores eval run history (last-run dates, result trends) across runs. Avoids re-running evals that haven't changed.
- **`workflow_call:`** in create mode, triggered by `learning-aggregator-ci` via `call-workflow`. Receives promotion candidates as workflow inputs.
- **`upload-artifact:`** persists eval results YAML for downstream consumption.
Workflow Rules
The CI agent follows these rules in order:
Mode: Run Evals (default)
1. Read `.evals/EVAL_INDEX.md` to get the list of all eval cases 2. For each eval case file in `.evals/cases/`: a. Read the eval case metadata and verification method b. Check preconditions — if not met, mark as `skip` c. Execute the verification method:
- `grep-check`: Search target files for pattern, compare to expected (found/not_found)
- `command-check`: Run the command, check exit code and/or output
- `file-check`: Verify file or section exists
- `rule-check`: Read target file, search for expected content
d. Compare result to expected outcome e. Record pass/fail/skip 3. Update `.evals/EVAL_INDEX.md` with `last-run` date and `last-result` for each case 4. Emit structured YAML under key `eval_creator_ci` 5. Post results summary as a PR comment or check annotation 6. If gate policy is enabled and any eval fails: fail the check run
Mode: Create Evals (from promotion candidates)
1. Read the `learning_aggregator_ci` artifact or gap report from the most recent learning-aggregator-ci run 2. For each promotion-ready pattern with `eval_candidate: true`: a. Determine the appropriate verification method based on the pattern type b. Create the eval case file in `.evals/cases/` with proper frontmatter c. Add the entry to `.evals/EVAL_INDEX.md` 3. Commit the new eval cases (if running with write permissions) 4. Report created evals in the output
Output Schema
eval_creator_ci:
version: "0.1.0"
source:
run_id: "<workflow run ID>"
trigger: "pull_request | schedule | workflow_dispatch"
run_date: "YYYY-MM-DD"
mode: "run | create | both"
run_results:
total: 12
passed: 10
failed: 1
skipped: 1
failures:
- id: "eval-20260301-001"
pattern_key: "harden.input_validation"
rule_summary: "Always validate external inputs"
expected: "not_found"
actual: "found"
target: "src/api/handler.ts"
recovery_action: "Add input validation to new handler endpoint"
skips:
- id: "eval-20260315-003"
reason: "Precondition not met: project does not use TypeScript"
create_results:
created: 2
cases:
- id: "eval-20260411-001"
pattern_key: "simplify.dead_code"
verification_method: "grep-check"
source_learning: "LRN-20260301-001"
- id: "eval-20260411-002"
pattern_key: "harden.authorization"
verification_method: "rule-check"
source_learning: "LRN-20260315-003"
summary:
regressions: 1
new_evals_created: 2
gate_result: "fail"
followup_required: trueRecommended Outputs
| Output | Destination | Content | |--------|------------|---------| | Eval results | PR c
A collection of skills for AI agents. Follows the Agent Skills specification. This repository is my personal skill testing ground.
Other skills on pskoett-ai-skills.
- /agent-teams-simplify-and-harden
Implementation + audit loop using parallel agent teams with structured simplify, harden, and document passes. Spawns implementation agents to do the work, then audit agents to find complexity, security gaps, and spec deviations, then loops until code compiles cleanly, all tests
Open skill - /context-surfing
Monitors context window health throughout a session and rides peak context quality for maximum output fidelity. Activates automatically after plan-interview and intent-framed-agent. Stays active through execution and hands off cleanly to simplify-and-harden and self-improvement
Open skill - /control-session-orchestrator
Control-plane workflow for coordinating multi-agent, multi-session project work from a single Codex, GitHub Copilot, or agent-app control session. Use this skill whenever the user asks to orchestrate agents, create or steer worker sessions, run a workflow-like effort, fan out
Open skill - /eval-creator
[Beta] Creates permanent eval cases from promoted learnings and runs regression checks against them. Turns failures into test cases that prevent silent regression. This is the outer loop''s regress-test step. Use when a learning is promoted and has a clear pass/fail condition,
Open skill - /intent-framed-agent
Frames coding-agent work sessions with explicit intent capture and drift monitoring. Use when a session transitions from planning/Q&A to implementation for coding tasks, refactors, feature builds, bug fixes, or other multi-step execution where scope drift is a risk.
Open skill - /learning-aggregator-ci
[Beta] CI-only learning aggregation workflow using gh-aw (GitHub Agentic Workflows). Scans .learnings/ files on a schedule, groups entries by pattern_key, identifies promotion-ready patterns, and posts a gap report as a PR or issue comment. Use when: you want automated
Open skill

