configs
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
How to launch prime-rl training runs — the `rl`, `sft`, and `inference` entrypoints, their config classes, and single-node/SLURM/dry-run modes. Use when starting a run or picking the right entrypoint. Standalone evals are the `eval` skill.
$ npx -y skills add primeintellect-ai/prime-rl --skill start-run --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/start-runContext preview
The summary Claude sees to decide when to auto-load this skill.
How to launch prime-rl training runs — the `rl`, `sft`, and `inference` entrypoints, their config classes, and single-node/SLURM/dry-run modes. Use when starting a run or picking the right entrypoint. Standalone evals are the `eval` skill.
name: start-run description: How to launch prime-rl training runs — the `rl`, `sft`, and `inference` entrypoints, their config classes, and single-node/SLURM/dry-run modes. Use when starting a run or picking the right entrypoint. Standalone evals are the `eval` skill.
All entrypoints run via `uv run <command>` and accept TOML configs via `@ path/to.toml` plus CLI overrides.
SLURM launches write generated scripts and coordination files under `<run_dir>/launcher/`, with batch logs under `launcher/logs/`. Local launches do not create this directory. Every launch writes configs and `command.txt` under `configs/attempt_<n>/`. `configs/latest` points to the current attempt. The command uses shell-safe quoting.
`output_dir` (default `outputs`) groups related runs; each run writes all its artifacts (logs, configs, checkpoints, broadcasts, rollouts) to its own run directory `<output_dir>/<run_name>`. `run.name` auto-generates as `<envs>--<model>--<short-id>` (SFT: `<dataset>--<model>--<short-id>`), so every launch gets a fresh, readable run directory; `run.dir` overrides the directory leaf when it should differ from the name. Pass `--run.name <name>` to make the run directory predictable — required to resume the run later (`--resume`, or `--resume.step N`, reuses the named run directory; without `[ckpt]` it loads but saves no new checkpoints). Launching into a run directory that already contains artifacts fails unless resuming or `--clean` is set (which wipes only that run directory).
[`pydantic-config`](https://github.com/PrimeIntellect-ai/pydantic-config) — Pydantic-based TOML + CLI loader. Highlights (see the `configs` skill for full mechanics):
`model.optim_cpu_offload = false` and `model.full_offload = true`. This mode uses the native CPU optimizer kernel, only supports AdamW and SignSGD (SignSGD is stateless and halves the host RAM footprint), and disables gradient clipping. Use a `[model.full_offload]` table only to disable NUMA binding.
Launches inference server, orchestrator, and trainer as subprocesses.
uv run rl @ examples/basic/reverse-text/rl.toml uv run rl @ examples/basic/reverse-text/rl.toml --dry-run # write scripts, don't run
verify the package imports under `uv run` (for example `uv run python -c "import importlib.util; print(importlib.util.find_spec('r2e_gym'))"`). If a local env exists under `deps/prime-envs/environments/` or `deps/verifiers/environments/` but does not import, install the env workspace members with `uv sync --all-extras --all-packages` (all) or `uv sync --all-extras --package prime-rl --package <env>` (one) — they're auto-discovered, no `pyproject.toml` edit needed. Keep `--all-extras` for training so a targeted package sync does not prune accelerator dependencies from the environment.
Launches torchrun internally — never call torchrun directly.
uv run sft @ examples/basic/reverse-text/sft.toml uv run sft @ examples/basic/reverse-text/sft.toml --slurm uv run sft @ examples/basic/reverse-text/sft.toml --dry-run
Router session cleanup runs on successful, failed, and cancelled episodes, including discarded retry traces. Cleanup failures are logged per session and do not disable subsequent releases.
OpenAI-compatible API plus prime-rl custom endpoints (`/update_weights`, `/load_lora_adapter`, `/init_broadcaster`). Always use this entrypoint — never `vllm serve` directly. It starts a `vllm-router` on `server.port` (default 8000, the client-facing URL) fronting the engine on `backend_port` (default 8100); admin endpoints must target the engine port directly. The default `sticky_least_loaded` policy keeps each rollout on one replica while assigning new sessions to the least-loaded replica; RL and SFT online eval automatically r
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
Find, start, use, and stop the local run dashboard for metrics, configs, traces, logs, and…
Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI…
How to install prime-rl and its optional dependencies. Use when setting up the project,…
How prime-rl vendors, builds, and ships CUDA kernels (the `deps/prime-kernels` submodule and…
How to prepare and publish GitHub releases for prime-rl. Use when drafting release notes,…