configs
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI shorthands, run directory, resume, logs and metrics. Use when asked to evaluate a model or checkpoint on an environment, smoke-test an environment, or check on an eval run.
$ npx -y skills add primeintellect-ai/prime-rl --skill eval --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/evalContext preview
The summary Claude sees to decide when to auto-load this skill.
Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI shorthands, run directory, resume, logs and metrics. Use when asked to evaluate a model or checkpoint on an environment, smoke-test an environment, or check on an eval run.
name: eval description: Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI shorthands, run directory, resume, logs and metrics. Use when asked to evaluate a model or checkpoint on an environment, smoke-test an environment, or check on an eval run.
`uv run eval` evaluates a model in one or more environments and exits after one epoch per source. It reuses the orchestrator's eval pipeline: one env server per source, a concurrency band, every episode through the monitors, and a trace stream that `--resume` continues from. Online evals of training runs are the `training` skill.
The user launches runs; hand over the command unless told otherwise. Always cap smokes (`-n`, `-r`).
uv run eval gsm8k -n 32 -r 4 # Prime Inference uv run eval gsm8k -n 32 -r 4 -c 8 --env.agent.harness.id bash # pin concurrency; set a field of the env block uv run eval gsm8k -n 32 -r 4 -m Qwen/Qwen3-4B --client.base_url http://localhost:8000/v1 # a local `uv run inference` server uv run eval @ eval.toml --run.name my-eval # multi-source TOML uv run eval @ eval.toml --run.name my-eval --resume # continue an interrupted run uv run eval @ eval.toml --dry-run # resolve and write the config, exit uv run eval @ eval.toml --monitors.prime # stream each source's epoch to the platform
Minimal multi-source TOML (the eval block is flattened to the top level; per-source `select`, `group_size`, `sampling` override the top level):
model = "Qwen/Qwen3-4B" group_size = 4 [select] limit = 32 [client] base_url = "http://localhost:8000/v1" [[source]] env.taskset.id = "gsm8k" env.agent.harness.id = "bash" [[source]] env.taskset.id = "aime25" env.agent.harness.id = "null" env.agent.runtime.type = "subprocess"
`select` picks which tasks of a source's taskset run. The steps apply in a fixed order: `include` and `exclude` (by `idx` position, `ids`, `keys` or `names`), then `shuffle` (under `seed`, default 0), `skip`, `limit`. Set it once for every source of a group (`[select]` in an eval TOML, `[orchestrator.eval.select]`, `[orchestrator.train.select]`), or per source (`select.limit = 50` in a `[[source]]`); a source's own fields win.
uv run eval gsm8k -n 50 # the first 50 tasks uv run eval gsm8k -n 50 -s # a random subset of 50 (same tasks every run) uv run eval gsm8k -n 50 -s --select.seed 1 # another random subset uv run eval gsm8k --select.include.idx 0:10,42 # tasks by position (ints and Python slices) uv run eval terminal-bench-2 --select.include.names '["fix-git"]' # tasks by name; `ids` and `keys` work the same uv run eval terminal-bench-2 --select.exclude.keys '["<task.key>"]' # drop known-broken tasks (keys are on every trace)
Disjoint train/test splits from one taskset:
# contiguous: eval on the first 100 tasks, train on the rest [[orchestrator.eval.source]] env.taskset.id = "gsm8k" select.include.idx = [":100"] [[orchestrator.train.source]] env.taskset.id = "gsm8k" select.exclude.idx = [":100"] # random: the same shuffle on both sides; eval takes 100, train skips them # eval: select.shuffle = true, select.limit = 100 # train: select.shuffle = true, select.skip = 100
The shuffle comes before `skip` and `limit`, so a larger `limit` extends the same selection: resuming a `-n 50 -s` run with `-n 100` keeps the 50 landed tasks and adds 50 more.
Run dir: `output_dir / run.name` (auto `<envs>--<model>--<short-id>`; `ls -t outputs | head -1` finds the latest). The console stays quiet while the eval runs; everything is in the files and the dashboard (`dashboard` skill).
{run_dir}/
├── configs/latest/ # command.txt, the launch TOML, resolved/eval.json
├── logs/latest/
│ ├── eval.log # the eval process
│ └── envs/eval/{name}.log # one log per env server
└── monitors/file/ # metrics.jsonl, the trace stream, traces/live/ (one file per live trace), plan.jsontail -F {run_dir}/logs/latest/eval.log
grep -E "WARNING|ERROR" {run_dir}/logs/latest/eval.log {run_dir}/logs/latest/envs/eval/*.log
grep SUCCESS {run_dir}/logs/latest/eval.log # one "Evaluated <env> ... Reward 0.xxxx" line per source
uv run python -m prime_rl.monitors.file.traces {run_dir} [<trace_id>] # live rollouts by phase, or one assembled live traceThe progress line in `eval.log` counts live rollouts by phase (`- boot 1 · running 3`). A rollout stuck in `boot` for minutes is waiting on its sandbox; one in `running` with a frozen turn count is waiting on a model call or tool.
Metrics live under `eval/<env>/all/<agent>/…` (`reward/mean`, `is_truncated/mean` — raise `sampling.max_completion_tokens` when
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
Find, start, use, and stop the local run dashboard for metrics, configs, traces, logs, and…
How to install prime-rl and its optional dependencies. Use when setting up the project,…
How prime-rl vendors, builds, and ships CUDA kernels (the `deps/prime-kernels` submodule and…
How to prepare and publish GitHub releases for prime-rl. Use when drafting release notes,…
Launch and monitor prime-rl training runs. Use when starting, supervising, or debugging an…