configs
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
Monitor an ongoing prime-rl training run — find the output directory, tail logs, check key metrics, inspect SLURM jobs, and restart safely. Use when asked to check on a run, debug training, or investigate performance.
$ npx -y skills add primeintellect-ai/prime-rl --skill monitor-run --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/monitor-runContext preview
The summary Claude sees to decide when to auto-load this skill.
Monitor an ongoing prime-rl training run — find the output directory, tail logs, check key metrics, inspect SLURM jobs, and restart safely. Use when asked to check on a run, debug training, or investigate performance.
name: monitor-run description: Monitor an ongoing prime-rl training run — find the output directory, tail logs, check key metrics, inspect SLURM jobs, and restart safely. Use when asked to check on a run, debug training, or investigate performance.
1. Find the run dir and read the resolved configs at `{run_dir}/configs/latest/resolved/` (start with `rl.json`, or `orchestrator.json` on local runs). Read the launch command from `{run_dir}/configs/latest/command.txt`. The launch TOML is copied verbatim to `{run_dir}/configs/latest/rl.toml`. The run dir is `{output_dir}/{run_name}` — `run.name` auto-generates as `<envs>--<model>--<short-id>`, so if you only know the output dir, pick the most recently modified subdirectory (`ls -t {output_dir} | head -1`) or read `run.name` from the launch command. 2. Confirm all processes are alive and the run is making progress. 3. Write the initial summary into `{run_dir}/STATUS.md`.
Default cadence: **1 hour** (researcher can override). At each check-in:
1. Confirm processes are alive. 2. Grep logs for errors/warnings; note current step and key metrics. 3. **Append** an entry to `{run_dir}/STATUS.md` (never overwrite):
## YYYY-MM-DD HH:MM UTC
**Step**: {current_step} / {max_steps}
**Health**: {Healthy | Degraded | Down}
**Progress**: reward/mean, seq_len, truncation, eval scores, env-specific metrics.
**Stability**: entropy, mismatch_kl, grad_norm — flag spikes.
**Performance**: trainer vs orchestrator step time, env lag, inference pressure.
**Notes**: anything unusual (errors, restarts, hangs). Omit if nothing notable.In W&B, each project auto-gets an **"overview" saved view** (train / eval / stability / performance sections) on its first run — use it for a quick check instead of the auto-generated default workspace.
**Never restart unless the researcher explicitly asked.** Confirm the exact restart command and the conditions that warrant one.
**Never** run kill or launch commands yourself. Hand the researcher the exact command and let them run it; after a restart, verify all processes are back up and progress resumed before the next check-in.
---
`uv run dashboard [output_dir ...]` (default `outputs/`, or `$PRL_OUTPUT_DIR` if set; several dirs can be tracked at once) serves a local web dashboard at `http://localhost:7788` with four views per run: metrics (the W&B overview sections, read from `metrics.jsonl`), per-attempt config files, a rollout trace viewer with per-token overlays (advantage, trainer logprob, entropy, KL mismatch, stable/loss/content masks), and merged component logs. It only reads the run dirs — safe to run against a live run. `--port`/`--host` pick the bind address; a taken port automatically bumps to the next free one, so several dashboards run side by side without coordination. GPU deps live behind the `gpu` extra, so `uv sync --extra dashboard && uv run dashboard` works without the training stack (e.g. on a head node).
**Daemon (auto-start)**: launchers auto-start one dashboard per host per user and a live one absorbs each new run's output dir automatically — see the `dashboard` skill for discovery, kill/restart commands, and `--isolated`. The short version: the live port can differ from 7788 (a taken port bumps), so read the discovery file:
cat ~/.cache/prime-rl/dashboard/daemon.json # {"pid": ..., "url": "http://localhost:<actual port>"}
ps aux | grep PRL::Dashboard # the daemon's process titleVerify liveness with `curl -sf <url>/api/runs` and hand the researcher the `url`.
{run_dir}/logs/latest/
├── trainer.log # rank 0 stdout
├── orchestrator.log # orchestrator stdout
├── eval.log # SFT online-eval process (also the `uv run eval` process — see the `eval` skill)
├── inference.log # vLLM stdout
├── trainer/
│ ├── node_*.log # per-node (multi-node only)
│ └── torchrun/ # per-rank stdout/stderr
├── inference/
│ ├── node_*.log # per-node (multi-node only)
│ └── router.log # the single global router (multi-node only; single-node logs it in inference.log)
└── envs/{train,eval}/{env_name}.log # one log file per envSLURM batch logs are under `{run_dir}/launcher/logs/*job_*.log`.
Usually tailing `trainer.log`, `orchestrator.log`, and `inference.log` is enough. Drop into per-node or per-rank logs only when debugging. All logs are loguru with `HH:mm:ss LEVEL message`; levels: `DEBUG`, `INFO`, `SUCCESS`, `WARNING`, `ERROR`.
Scan for problems:
grep -E "WARNING|ERROR" {run_dir}/logs/latest/{trainer,orchestrator,eval,inference}.log
grep -E "WARNING|ERROR" {run_dir}/logs/latest/envs/{train,eval}/*.log`{run_dir}/monitors/file/traces/live/<trace_id>.jsonl` holds the env server's streamed deltas of one live rollout (train or eval) and disappears when its episode lands in the stream, so the directory is the live set. `uv run python -m prime_rl.monitors.file.traces {run_dir}` prints one row per live trace with `stage` (pending/boot/setup/running/finalize/scoring/done/error), turns, tokens, elapsed and the latest message; with a trace id it prints the assembled trace. The dashboard's traces tab shows them in the episode table (tinted, phase badge, growing counts; filter status → in flight) and opens them in the viewer.
All metrics print to the
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
Find, start, use, and stop the local run dashboard for metrics, configs, traces, logs, and…
Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI…
How to install prime-rl and its optional dependencies. Use when setting up the project,…
How prime-rl vendors, builds, and ships CUDA kernels (the `deps/prime-kernels` submodule and…
How to prepare and publish GitHub releases for prime-rl. Use when drafting release notes,…