Skip to content
Testing
Skill

/rewardkit

Write Harbor task verifiers using Reward Kit. Use when creating or editing a

From plugin
lhtb
5285 skills
Install
$ npx -y skills add zli12321/LHTB --skill rewardkit --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/rewardkit

Context preview

The summary Claude sees to decide when to auto-load this skill.

Write Harbor task verifiers using Reward Kit. Use when creating or editing a

SKILL.md

rewardkit.SKILL.md
name: rewardkit
description: Write Harbor task verifiers using Reward Kit. Use when creating or editing a 
  task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing 
  verifiers that produce a reward score.

Help the user write task verifiers with Reward Kit. Reward Kit is a lightweight Python package that turns a directory of criteria files into a reward score. Each criterion is a Python function call or a TOML judge file; folders become separate rewards.

Setup in a Harbor task

Put criteria alongside `test.sh` in the task's `tests/` directory:

tests/
├── test.sh
├── checks.py         # programmatic criteria
└── judge.toml        # optional LLM/agent judge

`tests/test.sh`:

#!/bin/bash
uvx --from 'harbor-rewardkit==0.1.*' rewardkit /tests

This runs all criteria in `/tests/` against the workspace at `/app` and writes `/logs/verifier/reward.json`. Defaults match Harbor's conventions — no extra config needed.

If judge criteria need API keys, pass them through `task.toml`:

[verifier.env]
ANTHROPIC_API_KEY = "${ANTHROPIC_API_KEY}"

Ask whether Reward Kit should run in the agent's shared environment or in a separate verifier environment. Prefer a separate verifier environment when judge prompts, grading dependencies, API keys, or clean-room checks should not be available to the agent:

[verifier]
environment_mode = "separate"

[verifier.environment]
docker_image = "python:3.12-slim"
allow_internet = true

In separate mode, `tests/` is the verifier image build context and must provide `/tests/test.sh` at runtime; Harbor does not upload `tests/` into the running verifier container.

Programmatic criteria

Call built-ins from any `.py` file in `tests/`:

import rewardkit as rk

rk.file_exists("output.txt")
rk.file_contains("output.txt", "hello")
rk.command_succeeds("python main.py", weight=2.0)
rk.json_key_equals("result.json", "status", "ok")

All criteria accept `weight` (default `1.0`) and `isolated` (default `False`, runs in overlayfs so side effects don't leak).

Available built-ins

  • **Files**: `file_exists`, `file_not_exists`, `file_contains`, `file_contains_regex`,

`file_matches`, `files_equal`, `diff_ratio`

  • **Commands**: `command_succeeds`, `command_output_contains`, `command_output_matches`,

`command_output_matches_regex` (30s default timeout, optional `cwd`)

  • **Data**: `json_key_equals`, `json_path_equals`, `csv_cell_equals`, `xlsx_cell_equals`

(needs `[office]` extra), `sqlite_query_equals`

  • **HTTP**: `http_status_equals`, `http_response_contains`
  • **Images**: `image_similarity`, `image_size_equals` (needs `[image]` extra)
  • **Trajectory**: `trajectory_tool_used`, `trajectory_tool_not_used`, `trajectory_turn_count`

For extras, install with `uv tool install harbor-rewardkit[all]`.

Custom criteria

Use the `@criterion` decorator. First parameter is always `workspace: Path`. Returns `bool` or `float`:

from pathlib import Path
from rewardkit import criterion

@criterion
def has_valid_output(workspace: Path) -> bool:
    return (workspace / "output.txt").read_text().strip() != ""

Zero-parameter criteria auto-register. Criteria with extra args must be called via `rk`:

@criterion(description="output has at least {n} lines")
def has_n_lines(workspace: Path, n: int) -> bool:
    return len((workspace / "output.txt").read_text().splitlines()) >= n

rk.has_n_lines(10, weight=2.0)
rk.has_n_lines(50, weight=1.0)

For criteria shared across reward subdirs, define with `shared=True` in a root-level file and call from subdirs.

Judge criteria (LLM or agent-as-a-judge)

For subjective checks (quality, readability, edge cases), create a TOML file:

[judge]
judge = "anthropic/claude-sonnet-4-6"   # LiteLLM model string
files = ["/app/main.py"]

[[criterion]]
description = "Is the code correct?"
type = "binary"

[[criterion]]
description = "How readable is the code?"
type = "likert"
points = 5
weight = 2.0

Criterion types:

  • `binary` — yes/no → 1.0 or 0.0
  • `likert` — 1..points, normalized to [0, 1]
  • `numeric` — min..max, normalized to [0, 1]

Agent judges

Agent judges shell out to a CLI and can explore the filesystem:

[judge]
judge = "claude-code"
model = "anthropic/claude-sonnet-4-6"
isolated = true

[[criterion]]
description = "Does the solution handle edge cases?"
type = "binary"

Slower and more expensive than LLM judges, but they can run commands and inspect files.

Useful `[judge]` options

`timeout` (default 300), `reasoning_effort` (`low`|`medium`|`high`), `reference` (path to reference solution), `atif-trajectory` (evaluate the agent's trajectory), `weight`, `prompt_template` (custom prompt with `{criteria}` placeholder).

Scoring aggregation

[scoring]
aggregation = "all_pass"   # weighted_mean | all_pass | any_pass | threshold
threshold = 0.7             # only for threshold

Only affects aggregation *within* this TOML file.

Multi-reward tasks

Put criteria in subdirectories — each becomes a separate reward:

tests/
├── test.sh
├── correctness/
│   └── check.py
├── structure/
│   └── files_exist.py
└── quality/
    └── quality.toml

Produces:

{ "correctness": 0.75, "structure": 1.0, "quality": 0.6 }

Output files

  • `/logs/verifier/reward.json` — per-reward scores
  • `/logs/verifier/reward-details.json` — per-criterion results, judge reasoning, errors

Multi-step tasks

In a multi-step task, each step has its own `tests/` under `steps/{name}/tests/`, and the verifier runs once per step. Reward Kit behaves the same as in a single-step task: for each step it reads `/tests`, runs the criteria against `/app`, and writes `/logs/verifier/reward.json` for that step. Harbor then aggregates per-step results into a trial-level reward via `multi_step_reward_strategy` in `task.toml` — aggregation happens *outside* Reward Kit, so don't try to encode

Read more
Ships withlhtb

Long-Horizon Terminal-Bench is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps.

Get the whole plugin
Stats
529
Stars
34
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
4d ago
Last commit
1mo ago
Created

Repo: zli12321/LHTB