/harness-bench
Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants
$ npx -y skills add ruvnet/ruflo --skill harness-bench --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/harness-bench
Context preview
The summary Claude sees to decide when to auto-load this skill.
Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants
SKILL.md
harness-bench.SKILL.mdname: harness-bench
description: Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.
argument-hint: "--op create --repo <path> [--out <path>] | --op verify --suite <path>"
allowed-tools: Bash
Surfaces `metaharness-darwin bench <create|verify>` — the supporting verb for `harness-evolve --bench`. Use when you want evolution scored against a fixed corpus (independent of `npm test`) so champion fitness is comparable across commits or across forks of the same harness.
When to use
- Setting up a new evolution pipeline for a repo whose `npm test` is
flaky, slow, or undersized — scaffold a deterministic bench suite once, then evolve against it repeatedly.
- CI: `bench verify` the checked-in suite on every PR that touches it
(cheap; ~5s).
- Forking a harness to a new domain: copy and edit the suite to retarget
the evaluation without losing comparability to the parent.
Algorithm
Implementation: [`scripts/bench.mjs`](../../scripts/bench.mjs).
`--op create`
1. Resolve `--repo` path; reject if missing. 2. Shell to `metaharness-darwin bench create <repo> [--out <suite.json>]`. 3. Default output path: `<repo>/.metaharness/bench/suite.json` (chosen by upstream). 4. Suite shape (per upstream): array of `{ input, expectedOutput, weight }` tasks derived from existing test cases.
`--op verify`
1. Resolve `--suite` path; reject if missing. 2. Shell to `metaharness-darwin bench verify <suite.json>`. 3. Exit 1 if any task malformed (upstream's signal).
Output shape
{
"success": true,
"data": {
"op": "verify",
"taskCount": 42,
"wellFormed": true,
"durationMs": 870
}
}Exit codes
| Code | Meaning | |---|---| | 0 | OK (or degraded — Darwin absent) | | 1 | `--op verify` and suite malformed | | 2 | Config error or upstream invocation failure |
Graceful degradation
When `@metaharness/darwin` is absent, emits the standard `{degraded: true, reason: 'metaharness-darwin-not-available'}` payload and exits 0.
Read more
name: harness-bench description: Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent. argument-hint: "--op create --repo <path> [--out <path>] | --op verify --suite <path>" allowed-tools: Bash
Surfaces `metaharness-darwin bench <create|verify>` — the supporting verb for `harness-evolve --bench`. Use when you want evolution scored against a fixed corpus (independent of `npm test`) so champion fitness is comparable across commits or across forks of the same harness.
When to use
- Setting up a new evolution pipeline for a repo whose `npm test` is
flaky, slow, or undersized — scaffold a deterministic bench suite once, then evolve against it repeatedly.
- CI: `bench verify` the checked-in suite on every PR that touches it
(cheap; ~5s).
- Forking a harness to a new domain: copy and edit the suite to retarget
the evaluation without losing comparability to the parent.
Algorithm
Implementation: [`scripts/bench.mjs`](../../scripts/bench.mjs).
`--op create`
1. Resolve `--repo` path; reject if missing. 2. Shell to `metaharness-darwin bench create <repo> [--out <suite.json>]`. 3. Default output path: `<repo>/.metaharness/bench/suite.json` (chosen by upstream). 4. Suite shape (per upstream): array of `{ input, expectedOutput, weight }` tasks derived from existing test cases.
`--op verify`
1. Resolve `--suite` path; reject if missing. 2. Shell to `metaharness-darwin bench verify <suite.json>`. 3. Exit 1 if any task malformed (upstream's signal).
Output shape
{
"success": true,
"data": {
"op": "verify",
"taskCount": 42,
"wellFormed": true,
"durationMs": 870
}
}Exit codes
| Code | Meaning | |---|---| | 0 | OK (or degraded — Darwin absent) | | 1 | `--op verify` and suite malformed | | 2 | Config error or upstream invocation failure |
Graceful degradation
When `@metaharness/darwin` is absent, emits the standard `{degraded: true, reason: 'metaharness-darwin-not-available'}` payload and exits 0.
An agent meta-harness for Claude Code and Codex. Agent = Model + Harness. The model writes; the harness gives it tools, memory, loops, sandboxes, and controls so it can actually work.
Repo: ruvnet/ruflo
Other skills on claude-flow.
- /agentdb-advanced
Master advanced AgentDB features including QUIC synchronization, multi-database management, custom distance metrics, hybrid search, and distributed systems integration. Use when building distributed AI systems, multi-agent coordination, or advanced vector search applications.
Open skill - /agentdb-learning
Create and train AI learning plugins with AgentDB's 9 reinforcement learning algorithms. Includes Decision Transformer, Q-Learning, SARSA, Actor-Critic, and more. Use when building self-learning agents, implementing RL, or optimizing agent behavior through experience.
Open skill - /agentdb-memory-patterns
Implement persistent memory patterns for AI agents using AgentDB. Includes session memory, long-term storage, pattern learning, and context management. Use when building stateful agents, chat systems, or intelligent assistants.
Open skill - /agentdb-optimization
Optimize AgentDB performance with quantization (4-32x memory reduction), HNSW indexing (150x faster search), caching, and batch operations. Use when optimizing memory usage, improving search speed, or scaling to millions of vectors.
Open skill - /agentdb-vector-search
Implement semantic vector search with AgentDB for intelligent document retrieval, similarity matching, and context-aware querying. Use when building RAG systems, semantic search engines, or intelligent knowledge bases.
Open skill - /agentic-jujutsu
Quantum-resistant, self-learning version control for AI agents with ReasoningBank intelligence and multi-agent coordination
Open skill

