Skip to content
Machine Learning
Skill

/eval

Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI shorthands, run directory, resume, logs and metrics. Use when asked to evaluate a model or checkpoint on an environment, smoke-test an environment, or check on an eval run.

BOOST
From plugin
prime-rl
2.1k9 skills
Install
$ npx -y skills add primeintellect-ai/prime-rl --skill eval --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/eval

Context preview

The summary Claude sees to decide when to auto-load this skill.

Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI shorthands, run directory, resume, logs and metrics. Use when asked to evaluate a model or checkpoint on an environment, smoke-test an environment, or check on an eval run.

SKILL.md

eval.SKILL.md
name: eval
description: Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI shorthands, run directory, resume, logs and metrics. Use when asked to evaluate a model or checkpoint on an environment, smoke-test an environment, or check on an eval run.

Eval

`uv run eval` evaluates a model in one or more environments and exits after one epoch per source. It reuses the orchestrator's eval pipeline: one env server per source, a concurrency band, every episode through the monitors, and a trace stream that `--resume` continues from. Online evals of training runs are the `training` skill.

Start an eval

The user launches runs; hand over the command unless told otherwise. Always cap smokes (`-n`, `-r`).

uv run eval gsm8k -n 32 -r 4                                       # Prime Inference
uv run eval gsm8k -n 32 -r 4 -c 8 --env.agent.harness.id bash      # pin concurrency; set a field of the env block
uv run eval gsm8k -n 32 -r 4 -m Qwen/Qwen3-4B --client.base_url http://localhost:8000/v1   # a local `uv run inference` server
uv run eval @ eval.toml --run.name my-eval                          # multi-source TOML
uv run eval @ eval.toml --run.name my-eval --resume                 # continue an interrupted run
uv run eval @ eval.toml --dry-run                                   # resolve and write the config, exit
uv run eval @ eval.toml --monitors.prime                            # stream each source's epoch to the platform
  • Config: `EvalConfig` (`packages/prime-rl-configs/src/prime_rl/configs/eval.py`); `uv run eval -h` lists the fields.
  • Entrypoint: `src/prime_rl/entrypoints/eval.py`; implementation `src/prime_rl/eval/eval.py`, shared engine `src/prime_rl/eval/runner.py`.
  • Shorthands (single-source runs): `<taskset-id>`, `--env.<field> <value>`, `-n` `select.limit`, `-s` `select.shuffle`, `-r` `group_size`, `-m` `model`, `-c N` pins `concurrency.min_inflight = max_inflight = N`. They cannot be combined with a TOML that defines `[[source]]` blocks.
  • Concurrency: an API exposes no vLLM `/metrics`, so run pinned there; against `uv run inference` set `min_inflight < max_inflight` in `[concurrency]` to adapt to KV usage.
  • Env servers: spawned per source on OS-assigned ports (published to `configs/attempt_N/resolved/envs/eval/<name>.address`) unless the source sets `serve.address`.
  • Ready-made configs: `configs/debug/eval/*.toml`, one per shape (single turn, multi turn, resume, multi env, aime2026, tb2).

Minimal multi-source TOML (the eval block is flattened to the top level; per-source `select`, `group_size`, `sampling` override the top level):

model = "Qwen/Qwen3-4B"
group_size = 4

[select]
limit = 32

[client]
base_url = "http://localhost:8000/v1"

[[source]]
env.taskset.id = "gsm8k"
env.agent.harness.id = "bash"

[[source]]
env.taskset.id = "aime25"
env.agent.harness.id = "null"
env.agent.runtime.type = "subprocess"

Select tasks

`select` picks which tasks of a source's taskset run. The steps apply in a fixed order: `include` and `exclude` (by `idx` position, `ids`, `keys` or `names`), then `shuffle` (under `seed`, default 0), `skip`, `limit`. Set it once for every source of a group (`[select]` in an eval TOML, `[orchestrator.eval.select]`, `[orchestrator.train.select]`), or per source (`select.limit = 50` in a `[[source]]`); a source's own fields win.

uv run eval gsm8k -n 50                                   # the first 50 tasks
uv run eval gsm8k -n 50 -s                                # a random subset of 50 (same tasks every run)
uv run eval gsm8k -n 50 -s --select.seed 1                # another random subset
uv run eval gsm8k --select.include.idx 0:10,42            # tasks by position (ints and Python slices)
uv run eval terminal-bench-2 --select.include.names '["fix-git"]' # tasks by name; `ids` and `keys` work the same
uv run eval terminal-bench-2 --select.exclude.keys '["<task.key>"]' # drop known-broken tasks (keys are on every trace)

Disjoint train/test splits from one taskset:

# contiguous: eval on the first 100 tasks, train on the rest
[[orchestrator.eval.source]]
env.taskset.id = "gsm8k"
select.include.idx = [":100"]

[[orchestrator.train.source]]
env.taskset.id = "gsm8k"
select.exclude.idx = [":100"]

# random: the same shuffle on both sides; eval takes 100, train skips them
#   eval:  select.shuffle = true, select.limit = 100
#   train: select.shuffle = true, select.skip = 100

The shuffle comes before `skip` and `limit`, so a larger `limit` extends the same selection: resuming a `-n 50 -s` run with `-n 100` keeps the 50 landed tasks and adds 50 more.

Monitor an eval

Run dir: `output_dir / run.name` (auto `<envs>--<model>--<short-id>`; `ls -t outputs | head -1` finds the latest). The console stays quiet while the eval runs; everything is in the files and the dashboard (`dashboard` skill).

{run_dir}/
├── configs/latest/            # command.txt, the launch TOML, resolved/eval.json
├── logs/latest/
│   ├── eval.log               # the eval process
│   └── envs/eval/{name}.log   # one log per env server
└── monitors/file/             # metrics.jsonl, the trace stream, traces/live/ (one file per live trace), plan.json
tail -F {run_dir}/logs/latest/eval.log
grep -E "WARNING|ERROR" {run_dir}/logs/latest/eval.log {run_dir}/logs/latest/envs/eval/*.log
grep SUCCESS {run_dir}/logs/latest/eval.log                            # one "Evaluated <env> ... Reward 0.xxxx" line per source
uv run python -m prime_rl.monitors.file.traces {run_dir} [<trace_id>]   # live rollouts by phase, or one assembled live trace

The progress line in `eval.log` counts live rollouts by phase (`- boot 1 · running 3`). A rollout stuck in `boot` for minutes is waiting on its sandbox; one in `running` with a frozen turn count is waiting on a model call or tool.

Metrics live under `eval/<env>/all/<agent>/…` (`reward/mean`, `is_truncated/mean` — raise `sampling.max_completion_tokens` when

Read more
Ships withprime-rl

Agentic RL Training at Scale

Get the whole plugin
Stats
2,141
Stars
454
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
8m ago
Last commit
1y ago
Created
1d ago
Added

Repo: primeintellect-ai/prime-rl

Other skills on prime-rl.