Skip to content
Machine Learning
Skill

/start-run

How to launch prime-rl training runs — the `rl`, `sft`, and `inference` entrypoints, their config classes, and single-node/SLURM/dry-run modes. Use when starting a run or picking the right entrypoint. Standalone evals are the `eval` skill.

BOOST
From plugin
prime-rl
2.1k9 skills
Install
$ npx -y skills add primeintellect-ai/prime-rl --skill start-run --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/start-run

Context preview

The summary Claude sees to decide when to auto-load this skill.

How to launch prime-rl training runs — the `rl`, `sft`, and `inference` entrypoints, their config classes, and single-node/SLURM/dry-run modes. Use when starting a run or picking the right entrypoint. Standalone evals are the `eval` skill.

SKILL.md

start-run.SKILL.md
name: start-run
description: How to launch prime-rl training runs — the `rl`, `sft`, and `inference` entrypoints, their config classes, and single-node/SLURM/dry-run modes. Use when starting a run or picking the right entrypoint. Standalone evals are the `eval` skill.

Start a run

All entrypoints run via `uv run <command>` and accept TOML configs via `@ path/to.toml` plus CLI overrides.

SLURM launches write generated scripts and coordination files under `<run_dir>/launcher/`, with batch logs under `launcher/logs/`. Local launches do not create this directory. Every launch writes configs and `command.txt` under `configs/attempt_<n>/`. `configs/latest` points to the current attempt. The command uses shell-safe quoting.

Run directories

`output_dir` (default `outputs`) groups related runs; each run writes all its artifacts (logs, configs, checkpoints, broadcasts, rollouts) to its own run directory `<output_dir>/<run_name>`. `run.name` auto-generates as `<envs>--<model>--<short-id>` (SFT: `<dataset>--<model>--<short-id>`), so every launch gets a fresh, readable run directory; `run.dir` overrides the directory leaf when it should differ from the name. Pass `--run.name <name>` to make the run directory predictable — required to resume the run later (`--resume`, or `--resume.step N`, reuses the named run directory; without `[ckpt]` it loads but saves no new checkpoints). Launching into a run directory that already contains artifacts fails unless resuming or `--clean` is set (which wipes only that run directory).

Config system at a glance

[`pydantic-config`](https://github.com/PrimeIntellect-ai/pydantic-config) — Pydantic-based TOML + CLI loader. Highlights (see the `configs` skill for full mechanics):

  • Config files via `@ path` (TOML / YAML / JSON); CLI args layer on top, deep-merged with class defaults.
  • Nested groups via dotted CLI paths — kebab-case on the CLI, snake_case in TOML.
  • Bool toggles: bare `--flag` enables, `--no-flag` disables (nested too).
  • Lists: space-separated or JSON literal. Dicts: JSON literal, deep-merged with file values.
  • Optional sub-configs (`WandbMonitorConfig | None`): bare `--monitors.wandb` enables defaults; `--monitors.wandb @ wandb.toml` enables from a file; `--no-monitors.wandb` disables.
  • Discriminated unions are switched by the `type` tag (e.g. `--optimizer.type muon`).
  • Validation aliases let renamed fields keep working; legacy keys can be remapped in a `model_validator(mode="before")`.
  • Auto-generated `--help` panels from `Field(description=...)` or PEP 224 docstrings.
  • Friendly errors: required-field boxes, validator errors point at the offending flag, unknown flags get a "did you mean" hint.
  • State-only optimizer offload remains enabled by default with `model.optim_cpu_offload = true`.
  • For gradients, FP32 masters, optimizer state, and optimizer-in-backward CPU execution, set

`model.optim_cpu_offload = false` and `model.full_offload = true`. This mode uses the native CPU optimizer kernel, only supports AdamW and SignSGD (SignSGD is stateless and halves the host RAM footprint), and disables gradient clipping. Use a `[model.full_offload]` table only to disable NUMA binding.

`rl` — RL training

Launches inference server, orchestrator, and trainer as subprocesses.

uv run rl @ examples/basic/reverse-text/rl.toml
uv run rl @ examples/basic/reverse-text/rl.toml --dry-run                                # write scripts, don't run
  • Config: `RLConfig` (`packages/prime-rl-configs/src/prime_rl/configs/rl.py`)
  • Entrypoint: `src/prime_rl/entrypoints/rl.py`
  • SLURM: single- and multi-node
  • Multi-node SLURM stops after `.trainer.done` for trainer-only fake-data runs. Runs with inference stop after both `.trainer.done` and `.orchestrator.done`.
  • NIXL on SLURM: install NIXL and ModelExpress with the provided scripts. The job starts ModelExpress and Redis unless `slurm.launch_modelexpress = false`.
  • Environment packages: before launching a config with a non-core verifier env id,

verify the package imports under `uv run` (for example `uv run python -c "import importlib.util; print(importlib.util.find_spec('r2e_gym'))"`). If a local env exists under `deps/prime-envs/environments/` or `deps/verifiers/environments/` but does not import, install the env workspace members with `uv sync --all-extras --all-packages` (all) or `uv sync --all-extras --package prime-rl --package <env>` (one) — they're auto-discovered, no `pyproject.toml` edit needed. Keep `--all-extras` for training so a targeted package sync does not prune accelerator dependencies from the environment.

`sft` — SFT training

Launches torchrun internally — never call torchrun directly.

uv run sft @ examples/basic/reverse-text/sft.toml
uv run sft @ examples/basic/reverse-text/sft.toml --slurm
uv run sft @ examples/basic/reverse-text/sft.toml --dry-run
  • Config: `SFTConfig` (`packages/prime-rl-configs/src/prime_rl/configs/sft.py`)
  • Entrypoint: `src/prime_rl/entrypoints/sft.py`
  • SLURM: single- and multi-node
  • Multi-node online evals use one SLURM job with `num_train_nodes + num_infer_nodes` nodes. The generated `launcher/sft.sbatch` assigns inference nodes first, then trainer nodes.

`inference` — vLLM server

Router session cleanup runs on successful, failed, and cancelled episodes, including discarded retry traces. Cleanup failures are logged per session and do not disable subsequent releases.

OpenAI-compatible API plus prime-rl custom endpoints (`/update_weights`, `/load_lora_adapter`, `/init_broadcaster`). Always use this entrypoint — never `vllm serve` directly. It starts a `vllm-router` on `server.port` (default 8000, the client-facing URL) fronting the engine on `backend_port` (default 8100); admin endpoints must target the engine port directly. The default `sticky_least_loaded` policy keeps each rollout on one replica while assigning new sessions to the least-loaded replica; RL and SFT online eval automatically r

Read more
Ships withprime-rl

Agentic RL Training at Scale

Get the whole plugin
Stats
2,141
Stars
454
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
8m ago
Last commit
1y ago
Created
1d ago
Added

Repo: primeintellect-ai/prime-rl

Other skills on prime-rl.