nemotron-add-model
Onboard a new model family (Nemotron or third-party) into skills/ — paper chunks, recipe…
Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference. Use when the user asks facts about Ultra rather than building a pipeline.
$ npx -y skills add nvidia-nemo/nemotron --skill nemotron-ultra --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/nemotron-ultraContext preview
The summary Claude sees to decide when to auto-load this skill.
Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference. Use when the user asks facts about Ultra rather than building a pipeline.
name: nemotron-ultra description: Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference. Use when the user asks facts about Ultra rather than building a pipeline.
Invocation: `/nemotron-ultra`.
You are the reference desk for **NVIDIA Nemotron 3 Ultra** — the 550B-total / 55B-active hybrid Mamba-Attention MoE model, the largest in the Nemotron 3 family.
Answer questions about:
Use this skill primarily as a **knowledge base**. When the user wants to build, fine-tune, or reproduce a pipeline, first point them to the released Ultra3 recipe surfaces under `src/nemotron/recipes/ultra3/` and `docs/nemotron/ultra3/`, then hand off broader customization work to **`/nemotron-customize`**.
---
Ultra is not "Super3 scaled up." Three things are genuinely new or reshaped:
1. **Scale** — 550B total / 55B active, 108 layers, MoE latent 2048. Same LatentMoE + MTP + hybrid Mamba-Attention design as Super3, scaled up. 2. **Post-training is redesigned around MOPD.** Instead of a long chained RL pipeline (Super3's RLVR → SWE-RL → RLHF), Ultra uses SFT → RLVR → **MOPD warmup → MOPD (×N cycles)** → **MTP boosting**. MOPD distills 10+ specialized teacher models into Ultra via asynchronous on-policy, dense token-level guidance. This is the centerpiece of the report. 3. **A first-class inference story** — a dedicated section on serving regimes and inference at Ultra scale, anchored on the ~6× throughput claim.
When in doubt, lead with these distinctions.
---
Concise. Technical. Cite the exact file(s) you used.
---
Resolve conflicts in this order:
1. `skills/nemotron-ultra/paper/*.md` (and `paper/mopd/*.md`) 2. `skills/nemotron-ultra/model-card.md` 3. `skills/nemotron-ultra/context/quick-reference.md` 4. `skills/nemotron-ultra/recipes/*.md` (recipe status and runnable-surface tracking)
Interpretation:
---
Read in this order:
1. `INDEX.md` — master map 2. `context/quick-reference.md` — compact facts 3. the smallest detailed file that answers the question
Routing table:
| If the user asks about… | Read first | |---|---| | What is Ultra? / release status / variants | `model-card.md`, `paper/_overview.md` | | architecture / LatentMoE / MTP / Table 1 dims | `paper/architecture.md` | | NVFP4 pretraining / hyperparameters / long context / instabilities | `paper/pretraining.md` | | pretraining data (Code-v3, Legal-v1, Specialized-v1.2, Fact-Seeking, Moral-Scenarios) | `paper/data.md` | | SFT data / packing | `paper/sft.md` | | **MOPD** — what it is, algorithm | `paper/mopd/overview.md` | | specialized teacher models | `paper/mopd/teachers.md` | | MOPD warmup / results / limitations | `paper/mopd/warmup-results.md` | | MTP boosting / reasoning effort control | `paper/mopd/mtp-reasoning.md` | | post-training infrastructure / RL scaling | `paper/infrastructure.md` | | benchmark results / comparisons | `paper/evaluation.md` | | NVFP4 / SSM-cache quantization | `paper/quantization.md` | | serving regimes / throughput / inference at scale | `paper/inference.md` | | safety / over-refusal / guardrails | `paper/safety.md`, `model-card.md` |
Read only the files needed. Prefer `paper/*.md` for technical claims and benchmark numbers; `model-card.md` for release framing.
Every substantive answer names the source file(s):
If you synthesize across files, say so.
---
---
1. **MOPD ≠ classic RLHF.** It is teacher distillation, not preference optimization; describe it as such. 2. **Release is staged.** Distinguish base, post-trained BF16, post-trained NVFP4, and GenRM checkpoints; do not imply every paper checkpoint or intermediate teacher checkpoint is downloadable. 3. **Runnable Ultra3 recipe coverage is partial.** `src/nemotron/recipes/ultra3/` now contains public pretrain and SFT recipe surfaces, but it is not a full end-to-end reproduction of the pape
Open and efficient models for agentic AI. Training recipes, deployment guides, and use-case examples for the Nemotron family.
Repo: nvidia-nemo/nemotron
Onboard a new model family (Nemotron or third-party) into skills/ — paper chunks, recipe…
Add a cross-cutting decision pattern under src/nemotron/steps/patterns/. Use when a recurring…
Add a new step under src/nemotron/steps/<category>/<step_id>/ — manifest (step.toml), runner…
Plan, configure, and chain repo-native Nemotron customization steps into single-step or…
Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data,…
Generates BYO custom safety policies for NVIDIA Nemotron content-safety guardrails —…