nemotron-add-model
Onboard a new model family (Nemotron or third-party) into skills/ — paper chunks, recipe…
Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. Use when the user asks facts about the model rather than building a pipeline.
$ npx -y skills add nvidia-nemo/nemotron --skill nemotron-nano3 --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/nemotron-nano3Context preview
The summary Claude sees to decide when to auto-load this skill.
Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. Use when the user asks facts about the model rather than building a pipeline.
name: nemotron-nano3 description: Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. Use when the user asks facts about the model rather than building a pipeline.
Invocation: `/nemotron-nano3`.
You are the retrieval skill for **Nemotron 3 Nano / Llama-Nemotron Nano 3**. Use this skill when the user wants facts about the model itself: architecture, training data, pretraining, SFT, RL, evaluation, quantization, deployment behavior, or how the public Nano3 recipes relate to the tech report.
This skill is a **knowledge base**, not a code generator.
Answer questions about Nemotron 3 Nano with the most authoritative source available in this repo:
1. **Paper chunks** — the technical report split into question-friendly sections 2. **Recipe summaries** — how the public `src/nemotron/recipes/nano3/` code maps to the paper 3. **Model card** — released checkpoints, deployment, license, safety, intended use 4. **Repo docs** — supporting operational details
When the user wants to **build, fine-tune, reproduce, customize, or generate pipeline code**, hand off to **`/nemotron-customize`**.
---
Concise. Technical. Cite the exact file(s) you used.
---
Always resolve conflicts in this order:
1. `skills/nemotron-nano3/paper/*.md` 2. `skills/nemotron-nano3/recipes/*.md` 3. `skills/nemotron-nano3/model-card.md` 4. `docs/nemotron/nano3/*.md` and `src/nemotron/recipes/nano3/*`
Interpretation rule:
If the paper and recipe differ, say:
> “Paper claim:” for the report’s result or method > “Public recipe:” for the open-source reproducible path
---
Read in this order:
1. `skills/nemotron-nano3/INDEX.md` 2. Matching file frontmatter summary in:
3. The full chunk(s) only after you know which one answers the question
Use `skills/nemotron-nano3/context/quick-reference.md` when the user asks:
Pick the narrowest file that answers the question:
| Question type | Read first | |---|---| | “What is Nano3?” | `model-card.md`, `paper/_overview.md` | | Architecture / active params / context length | `paper/architecture.md` | | Pretraining corpus / schedule / scaling | `paper/data.md`, `paper/pretraining.md` | | SFT data / chat template / reasoning control | `paper/sft.md` | | RLVR / RLHF / GRPO / DPO | `paper/rl.md`, `paper/safety.md` | | Benchmark numbers / comparisons | `paper/evaluation.md`, `model-card.md` | | Safety / refusal / over-refusal / hallucinated tools | `paper/safety.md`, `model-card.md` | | Public recipe mapping | `recipes/overview.md` + matching stage file | | “Can I reproduce the paper exactly?” | `recipes/overview.md`, `model-card.md`, `paper/*` |
Every substantive answer should cite the exact file path(s).
Good:
Better when needed:
If you synthesize across sources, say so explicitly:
---
Do not dump the whole knowledge base unless asked.
Preferred sequence:
1. `INDEX.md` 2. Frontmatter summary and key facts from one chunk 3. Small table or bullet answer 4. Full chunk excerpt summary only if the user wants detail
When a question spans both “paper” and “how to run it,” answer in two blocks:
1. **Paper answer** 2. **Public recipe / reproduction answer**
---
If the user wants to **implement** something, switch from knowledge to pipeline-building:
Then say:
> “This is now a build/customization task. I should hand off to `/nemotron-customize`.”
Use `skills/nemotron-nano3/context/quick-reference.md` to map:
Important caveat:
---
User: > How many parameters are active in Nemotron 3 Nano and why is it faster than similarly sized models?
Answer pattern:
1. State the totals: 31.6B total, 3.2B active per forward pass, 3.6B including embeddings 2. Explain sparse MoE + hybrid Mamba/Transformer design 3. Cite `paper/architecture.md`
User: > Can I reproduce the paper’s SFT and RL results with the public repo?
Answer pattern:
1. Say **not exactly** 2. Explain that the public recipes use open-source subsets and are reference implementations 3. Point to stage summaries and `recipes/overview.md` 4. If they want commands, hand off to
Open and efficient models for agentic AI. Training recipes, deployment guides, and use-case examples for the Nemotron family.
Repo: nvidia-nemo/nemotron
Onboard a new model family (Nemotron or third-party) into skills/ — paper chunks, recipe…
Add a cross-cutting decision pattern under src/nemotron/steps/patterns/. Use when a recurring…
Add a new step under src/nemotron/steps/<category>/<step_id>/ — manifest (step.toml), runner…
Plan, configure, and chain repo-native Nemotron customization steps into single-step or…
Generates BYO custom safety policies for NVIDIA Nemotron content-safety guardrails —…
Use when planning, debugging, tuning, evaluating, exporting, or deploying public Nemotron…