configs
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
How prime-rl vendors, builds, and ships CUDA kernels (the `deps/prime-kernels` submodule and the `prime-kernels` wheel). Use when adding a kernel, building it locally, calling one from training code, or publishing prebuilt wheels.
$ npx -y skills add primeintellect-ai/prime-rl --skill kernels --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/kernelsContext preview
The summary Claude sees to decide when to auto-load this skill.
How prime-rl vendors, builds, and ships CUDA kernels (the `deps/prime-kernels` submodule and the `prime-kernels` wheel). Use when adding a kernel, building it locally, calling one from training code, or publishing prebuilt wheels.
name: kernels description: How prime-rl vendors, builds, and ships CUDA kernels (the `deps/prime-kernels` submodule and the `prime-kernels` wheel). Use when adding a kernel, building it locally, calling one from training code, or publishing prebuilt wheels.
CUDA kernels live in their own monorepo, [prime-kernels](https://github.com/PrimeIntellect-ai/prime-kernels), checked out here as the git submodule `deps/prime-kernels`, alongside prime-rl's other submodules. That repo is the wheel root (`setup.py`, `pyproject.toml`) and `prime_kernels/` inside it is the importable package: one folder per kernel, holding the kernel's Python surface and, for compiled kernels, its C++/CUDA sources under `csrc/`, all declared in the single manifest `prime_kernels/kernels.toml`. See `deps/prime-kernels/README.md` once the submodule is initialized.
Nothing about a kernel lives in prime-rl. prime-rl pins a prime-kernels commit for local source builds and a prime-kernels release for installs. prime-kernels builds and publishes its own wheels. prime-rl stays a pure-Python wheel; never add compiled extensions to it.
Living under `deps/` means `tool.ruff.extend-exclude = ["deps"]` in `pyproject.toml` already covers it — prime-rl lints none of it.
Kernels are compiled for exact compute capabilities and may not be built at all, so always gate. Never import `prime_kernels.<name>` directly in training code:
import prime_kernels
if prime_kernels.is_available("flash_moe"):
flash_moe = prime_kernels.load("flash_moe")`prime_kernels.status()` maps every kernel to `"available"` or the reason it is not — log it once at startup rather than failing a run halfway through. `unavailable_reason(name)` is the same answer for one kernel (`None` when it is usable), which is what a test's skip guard wants; `is_available` is just that call compared to `None`.
`flash_moe` is a compiled fused MoE forward kernel (bf16 + mxfp8) on Blackwell tcgen05. Its trainer integration is dormant. `mxfp8_moe` is a Python-only registered kernel package for MXFP8 grouped GEMM and torch EP transport on SM100; it owns the MoE-specific torchao-derived orchestration and exports explicit BF16 boundaries instead of tensor-subclass interception.
What a kernel requires of its inputs — block sizes, alignments, shape constraints — belongs to prime-kernels, which exports it: `flash_moe.BLOCK_M`, `flash_moe.MXFP8_SCALE_BLOCK`, and `flash_moe.unsupported_shape_reason(dim, hidden_dim, mxfp8=...)`. When the trainer integration is restored, call this once during setup so an unsupported model fails before training rather than mid-step. Never hardcode a `128` on this side: then every requirement change is a change in both repos.
`uv sync --extra kernels` installs the prebuilt wheel (see "Pinning installs at the prebuilt wheels"); building from source is for changing kernels. It is manual by design — no `uv sync` may compile CUDA, so the extra resolves to release wheels, never to this source tree:
git submodule update --init deps/prime-kernels uv pip install --no-build-isolation -e deps/prime-kernels
Requirements: `nvcc` on `CUDA_HOME` with the **same CUDA major as torch** (torch refuses to build extensions otherwise).
Kernels whose toolkit is unsuitable are skipped with a message and reported unavailable at runtime — the build still succeeds. `PRIME_KERNELS=a,b` builds a subset; `PRIME_KERNELS_REQUIRE=1` turns any skip into an error.
The work happens in the prime-kernels repo, not here. Inside `deps/prime-kernels/`:
1. Commit the sources under `prime_kernels/<name>/csrc/`. 2. Add a `[<name>]` table to `prime_kernels/kernels.toml` — `sources`, `include-dirs`, `arch`, `cxx-std`; paths are relative to the kernel folder. A vendored Python kernel uses `python-only = true`, may declare import checks in `requires`, and omits compiled extension fields. 3. Write `prime_kernels/<name>/__init__.py` — `from . import _C`, then per op a wrapper calling `torch.ops.<ns>.<op>` and a `torch.library.register_fake`. No `torch.library.custom_op` decorator: that defines a *Python* op, and `TORCH_LIBRARY` has already defined these C++ side — only the fake (meta) kernel is missing. A kernel used in training also needs `torch.library.register_autograd`, since a schema carries no backward. `flash_moe` is forward only and currently has no trainer integration. 4. Nothing else: `setup.py` and the runtime registry both read the manifest.
Rules the build assumes:
`PYBIND11_MODULE(_C, m)` and registers ops with `TORCH_LIBRARY*`.
shipped to JIT from.
also installed as a standalone package (e.g. `prime_moe` from prime-flash-moe, where `flash_moe` originally came from), uninstall it.
Then land it in prime-rl as a submodule bump.
Kernel sources are pinned by the submodule commit, so picking up any kernel change — yours or someone else's — is a bump:
git -C deps/prime-kernels fetch origin git -C deps/prime-kernels log --oneline HEAD..origin/main git -C deps/prime-kernels checkout origin/main git add deps/prime-kernels
Then, in order:
what the caller must pass (weight layout, scale packing, argument order) is silently wrong numbers, not a build error, and prime-rl's call sites have to absorb it.
path that calls the changed kernel; the ABI is not chec
How the prime-rl config system works — TOML files, CLI overrides, composition, and special…
Find, start, use, and stop the local run dashboard for metrics, configs, traces, logs, and…
Launch and monitor prime-rl evals — the `uv run eval` entrypoint, its config and CLI…
How to install prime-rl and its optional dependencies. Use when setting up the project,…
How to prepare and publish GitHub releases for prime-rl. Use when drafting release notes,…
Launch and monitor prime-rl training runs. Use when starting, supervising, or debugging an…