Skip to content
Machine Learning
Skill

/monitor-run

Monitor an ongoing prime-rl training run — find the output directory, tail logs, check key metrics, inspect SLURM jobs, and restart safely. Use when asked to check on a run, debug training, or investigate performance.

BOOST
From plugin
prime-rl
2.1k9 skills
Install
$ npx -y skills add primeintellect-ai/prime-rl --skill monitor-run --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/monitor-run

Context preview

The summary Claude sees to decide when to auto-load this skill.

Monitor an ongoing prime-rl training run — find the output directory, tail logs, check key metrics, inspect SLURM jobs, and restart safely. Use when asked to check on a run, debug training, or investigate performance.

SKILL.md

monitor-run.SKILL.md
name: monitor-run
description: Monitor an ongoing prime-rl training run — find the output directory, tail logs, check key metrics, inspect SLURM jobs, and restart safely. Use when asked to check on a run, debug training, or investigate performance.

Monitor a run

Runbook

On launch

1. Find the run dir and read the resolved configs at `{run_dir}/configs/latest/resolved/` (start with `rl.json`, or `orchestrator.json` on local runs). Read the launch command from `{run_dir}/configs/latest/command.txt`. The launch TOML is copied verbatim to `{run_dir}/configs/latest/rl.toml`. The run dir is `{output_dir}/{run_name}` — `run.name` auto-generates as `<envs>--<model>--<short-id>`, so if you only know the output dir, pick the most recently modified subdirectory (`ls -t {output_dir} | head -1`) or read `run.name` from the launch command. 2. Confirm all processes are alive and the run is making progress. 3. Write the initial summary into `{run_dir}/STATUS.md`.

Recurring check-ins

Default cadence: **1 hour** (researcher can override). At each check-in:

1. Confirm processes are alive. 2. Grep logs for errors/warnings; note current step and key metrics. 3. **Append** an entry to `{run_dir}/STATUS.md` (never overwrite):

## YYYY-MM-DD HH:MM UTC

**Step**: {current_step} / {max_steps}
**Health**: {Healthy | Degraded | Down}

**Progress**: reward/mean, seq_len, truncation, eval scores, env-specific metrics.
**Stability**: entropy, mismatch_kl, grad_norm — flag spikes.
**Performance**: trainer vs orchestrator step time, env lag, inference pressure.

**Notes**: anything unusual (errors, restarts, hangs). Omit if nothing notable.

In W&B, each project auto-gets an **"overview" saved view** (train / eval / stability / performance sections) on its first run — use it for a quick check instead of the auto-generated default workspace.

Restarting a run

**Never restart unless the researcher explicitly asked.** Confirm the exact restart command and the conditions that warrant one.

**Never** run kill or launch commands yourself. Hand the researcher the exact command and let them run it; after a restart, verify all processes are back up and progress resumed before the next check-in.

---

Reference

Where to find things

  • `{run_dir}/configs/latest/` — the current attempt's command, launch TOML, and `resolved/` JSON files. Each launch stays under `configs/attempt_<n>/`.
  • `{run_dir}/logs/latest/` — the current attempt's logs (each launch gets `logs/attempt_<n>/`; resumes never overwrite earlier attempts). See below.
  • `{run_dir}/monitors/file/` — the metrics, and the traces with the annotations about them (see Episodes below).

Dashboard

`uv run dashboard [output_dir ...]` (default `outputs/`, or `$PRL_OUTPUT_DIR` if set; several dirs can be tracked at once) serves a local web dashboard at `http://localhost:7788` with four views per run: metrics (the W&B overview sections, read from `metrics.jsonl`), per-attempt config files, a rollout trace viewer with per-token overlays (advantage, trainer logprob, entropy, KL mismatch, stable/loss/content masks), and merged component logs. It only reads the run dirs — safe to run against a live run. `--port`/`--host` pick the bind address; a taken port automatically bumps to the next free one, so several dashboards run side by side without coordination. GPU deps live behind the `gpu` extra, so `uv sync --extra dashboard && uv run dashboard` works without the training stack (e.g. on a head node).

**Daemon (auto-start)**: launchers auto-start one dashboard per host per user and a live one absorbs each new run's output dir automatically — see the `dashboard` skill for discovery, kill/restart commands, and `--isolated`. The short version: the live port can differ from 7788 (a taken port bumps), so read the discovery file:

cat ~/.cache/prime-rl/dashboard/daemon.json   # {"pid": ..., "url": "http://localhost:<actual port>"}
ps aux | grep PRL::Dashboard                  # the daemon's process title

Verify liveness with `curl -sf <url>/api/runs` and hand the researcher the `url`.

Logs

{run_dir}/logs/latest/
├── trainer.log                # rank 0 stdout
├── orchestrator.log           # orchestrator stdout
├── eval.log                   # SFT online-eval process (also the `uv run eval` process — see the `eval` skill)
├── inference.log              # vLLM stdout
├── trainer/
│   ├── node_*.log             # per-node (multi-node only)
│   └── torchrun/              # per-rank stdout/stderr
├── inference/
│   ├── node_*.log             # per-node (multi-node only)
│   └── router.log             # the single global router (multi-node only; single-node logs it in inference.log)
└── envs/{train,eval}/{env_name}.log    # one log file per env

SLURM batch logs are under `{run_dir}/launcher/logs/*job_*.log`.

Usually tailing `trainer.log`, `orchestrator.log`, and `inference.log` is enough. Drop into per-node or per-rank logs only when debugging. All logs are loguru with `HH:mm:ss LEVEL message`; levels: `DEBUG`, `INFO`, `SUCCESS`, `WARNING`, `ERROR`.

Scan for problems:

grep -E "WARNING|ERROR" {run_dir}/logs/latest/{trainer,orchestrator,eval,inference}.log
grep -E "WARNING|ERROR" {run_dir}/logs/latest/envs/{train,eval}/*.log

Live rollouts

`{run_dir}/monitors/file/traces/live/<trace_id>.jsonl` holds the env server's streamed deltas of one live rollout (train or eval) and disappears when its episode lands in the stream, so the directory is the live set. `uv run python -m prime_rl.monitors.file.traces {run_dir}` prints one row per live trace with `stage` (pending/boot/setup/running/finalize/scoring/done/error), turns, tokens, elapsed and the latest message; with a trace id it prints the assembled trace. The dashboard's traces tab shows them in the episode table (tinted, phase badge, growing counts; filter status → in flight) and opens them in the viewer.

Metrics

All metrics print to the

Read more
Ships withprime-rl

Agentic RL Training at Scale

Get the whole plugin
Stats
2,141
Stars
454
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
8m ago
Last commit
1y ago
Created
1d ago
Added

Repo: primeintellect-ai/prime-rl

Other skills on prime-rl.