code-assist
Guides implementation of code tasks using test-driven development in an Explore, Plan, Code, Commit workflow. Acts as a Technical Implementation Partner and…
Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
$ npx -y skills add mikeyobrien/ralph-orchestrator --skill evaluate-presets --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/evaluate-presetsContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues.
name: evaluate-presets description: Use when testing Ralph's hat collection presets, validating preset configurations, or auditing the preset library for bugs and UX issues. metadata: internal: true
Systematically test all hat collection presets using shell scripts. Direct CLI invocation—no meta-orchestration complexity.
**Evaluate a single preset:**
./tools/evaluate-preset.sh tdd-red-green claude
**Evaluate all presets:**
./tools/evaluate-all-presets.sh claude
**Arguments:**
**IMPORTANT:** When invoking these scripts via the Bash tool, use these settings:
Since preset evaluations can run for hours (especially the full suite), **always run in background mode** and use the `TaskOutput` tool to check progress periodically.
**Example invocation pattern:**
Bash tool with: command: "./tools/evaluate-preset.sh tdd-red-green claude" timeout: 600000 run_in_background: true
After launching, use `TaskOutput` with `block: false` to check status without waiting for completion.
1. Loads test task from `tools/preset-test-tasks.yml` (if `yq` available) 2. Creates merged config with evaluation settings 3. Runs Ralph with `--record-session` for metrics capture 4. Captures output logs, exit codes, and timing 5. Extracts metrics: iterations, hats activated, events published
**Output structure:**
.eval/ ├── logs/<preset>/<timestamp>/ │ ├── output.log # Full stdout/stderr │ ├── session.jsonl # Recorded session │ ├── metrics.json # Extracted metrics │ ├── environment.json # Runtime environment │ └── merged-config.yml # Config used └── logs/<preset>/latest -> <timestamp>
Runs all 12 presets sequentially and generates a summary:
.eval/results/<suite-id>/ ├── SUMMARY.md # Markdown report ├── <preset>.json # Per-preset metrics └── latest -> <suite-id>
| Preset | Test Task | |--------|-----------| | `tdd-red-green` | Add `is_palindrome()` function | | `adversarial-review` | Review user input handler for security | | `socratic-learning` | Understand `HatRegistry` | | `spec-driven` | Specify and implement `StringUtils::truncate()` | | `mob-programming` | Implement a `Stack` data structure | | `scientific-method` | Debug failing mock test assertion | | `code-archaeology` | Understand history of `config.rs` | | `performance-optimization` | Profile hat matching | | `api-design` | Design a `Cache` trait | | `documentation-first` | Document `RateLimiter` | | `incident-response` | Respond to "tests failing in CI" | | `migration-safety` | Plan v1 to v2 config migration |
**Exit codes from `evaluate-preset.sh`:**
**Metrics in `metrics.json`:**
**Critical:** Validate that hats get fresh context per Tenet #1 ("Fresh Context Is Reliability").
Each hat should execute in its **own iteration**:
Iter 1: Ralph → publishes starting event → STOPS Iter 2: Hat A → does work → publishes next event → STOPS Iter 3: Hat B → does work → publishes next event → STOPS Iter 4: Hat C → does work → LOOP_COMPLETE
**BAD:** Multiple hat personas in one iteration:
Iter 2: Ralph does Blue Team + Red Team + Fixer work
^^^ All in one bloated context!**1. Count iterations vs events in `session.jsonl`:**
# Count iterations grep -c "_meta.loop_start\|ITERATION" .eval/logs/<preset>/latest/output.log # Count events published grep -c "bus.publish" .eval/logs/<preset>/latest/session.jsonl
**Expected:** iterations ≈ events published (one event per iteration) **Bad sign:** 2-3 iterations but 5+ events (all work in single iteration)
**2. Check for same-iteration hat switching in `output.log`:**
grep -E "ITERATION|Now I need to perform|Let me put on|I'll switch to" \
.eval/logs/<preset>/latest/output.log**Red flag:** Hat-switching phrases WITHOUT an ITERATION separator between them.
**3. Check event timestamps in `session.jsonl`:**
cat .eval/logs/<preset>/latest/session.jsonl | jq -r '.ts'
**Red flag:** Multiple events with identical timestamps (published in same iteration).
| Pattern | Diagnosis | Action | |---------|-----------|--------| | iterations ≈ events | ✅ Good | Hat routing working | | iterations << events | ⚠️ Same-iteration switching | Check prompt has STOP instruction | | iterations >> events | ⚠️ Recovery loops | Agent not publishing required events | | 0 events | ❌ Broken | Events not being read from JSONL |
If hat routing is broken:
1. **Check workflow prompt** in `hatless_ralph.rs`:
2. **Check hat instructions** propagation:
3. **Check
A hat-based orchestration framework that keeps AI agents in a loop until the task is done. "Me fail English? That's unpossible!" - Ralph Wiggum
Repo: mikeyobrien/ralph-orchestrator
Guides implementation of code tasks using test-driven development in an Explore, Plan, Code, Commit workflow. Acts as a Technical Implementation Partner and…
Generates structured .code-task.md files from descriptions or PDD implementation plans. Auto-detects input type, creates properly formatted tasks with…
Lists all code tasks in the repository with their status, dates, and metadata. Useful for getting an overview of pending work or finding specific tasks.
Transforms a rough idea into a detailed design document with implementation plan. Follows Prompt-Driven Development — iterative requirements clarification,…
Browser automation via Playwriter (remorses) using persistent Chrome sessions and the full Playwright Page API.
Use when creating animated demos (GIFs) for pull requests or documentation. Covers terminal recording with asciinema and conversion to GIF/SVG for GitHub…