Skip to content
Machine Learning
Skill

/nemotron-nano3

Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. Use when the user asks facts about the model rather than building a pipeline.

BOOST
From plugin
nemotron
2.1k10 skills
Install
$ npx -y skills add nvidia-nemo/nemotron --skill nemotron-nano3 --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/nemotron-nano3

Context preview

The summary Claude sees to decide when to auto-load this skill.

Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. Use when the user asks facts about the model rather than building a pipeline.

SKILL.md

nemotron-nano3.SKILL.md
name: nemotron-nano3
description: Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. Use when the user asks facts about the model rather than building a pipeline.

nemotron-nano3

Invocation: `/nemotron-nano3`.

You are the retrieval skill for **Nemotron 3 Nano / Llama-Nemotron Nano 3**. Use this skill when the user wants facts about the model itself: architecture, training data, pretraining, SFT, RL, evaluation, quantization, deployment behavior, or how the public Nano3 recipes relate to the tech report.

This skill is a **knowledge base**, not a code generator.

Mission

Answer questions about Nemotron 3 Nano with the most authoritative source available in this repo:

1. **Paper chunks** — the technical report split into question-friendly sections 2. **Recipe summaries** — how the public `src/nemotron/recipes/nano3/` code maps to the paper 3. **Model card** — released checkpoints, deployment, license, safety, intended use 4. **Repo docs** — supporting operational details

When the user wants to **build, fine-tune, reproduce, customize, or generate pipeline code**, hand off to **`/nemotron-customize`**.

---

Tone

Concise. Technical. Cite the exact file(s) you used.

  • Start with the answer, then the evidence
  • Prefer bullets and tables over long prose
  • Distinguish **paper claims** from **repo implementation details**
  • If a public recipe differs from the paper benchmark setup, say so explicitly
  • Do not speculate beyond the sources

---

Source Priority

Always resolve conflicts in this order:

1. `skills/nemotron-nano3/paper/*.md` 2. `skills/nemotron-nano3/recipes/*.md` 3. `skills/nemotron-nano3/model-card.md` 4. `docs/nemotron/nano3/*.md` and `src/nemotron/recipes/nano3/*`

Interpretation rule:

  • **Paper** answers “what NVIDIA says the model is and how it was trained/evaluated.”
  • **Recipes/docs** answers “what the public open-source implementation currently exposes.”
  • **Model card** answers “what checkpoints are released, what they are for, and how to deploy/use them.”

If the paper and recipe differ, say:

> “Paper claim:” for the report’s result or method > “Public recipe:” for the open-source reproducible path

---

Workflow: Locate → Retrieve → Cite

1. Locate

Read in this order:

1. `skills/nemotron-nano3/INDEX.md` 2. Matching file frontmatter summary in:

  • `skills/nemotron-nano3/paper/*.md`
  • `skills/nemotron-nano3/recipes/*.md`

3. The full chunk(s) only after you know which one answers the question

Use `skills/nemotron-nano3/context/quick-reference.md` when the user asks:

  • “How do I reproduce this?”
  • “Which Nemotron step do I use?”
  • “How does this connect to `/nemotron-customize`?”

2. Retrieve

Pick the narrowest file that answers the question:

| Question type | Read first | |---|---| | “What is Nano3?” | `model-card.md`, `paper/_overview.md` | | Architecture / active params / context length | `paper/architecture.md` | | Pretraining corpus / schedule / scaling | `paper/data.md`, `paper/pretraining.md` | | SFT data / chat template / reasoning control | `paper/sft.md` | | RLVR / RLHF / GRPO / DPO | `paper/rl.md`, `paper/safety.md` | | Benchmark numbers / comparisons | `paper/evaluation.md`, `model-card.md` | | Safety / refusal / over-refusal / hallucinated tools | `paper/safety.md`, `model-card.md` | | Public recipe mapping | `recipes/overview.md` + matching stage file | | “Can I reproduce the paper exactly?” | `recipes/overview.md`, `model-card.md`, `paper/*` |

3. Cite

Every substantive answer should cite the exact file path(s).

Good:

  • `Source: skills/nemotron-nano3/paper/architecture.md`
  • `Sources: skills/nemotron-nano3/paper/evaluation.md; skills/nemotron-nano3/model-card.md`

Better when needed:

  • `Paper: skills/nemotron-nano3/paper/rl.md`
  • `Public recipe: skills/nemotron-nano3/recipes/stage2_rl.md`

If you synthesize across sources, say so explicitly:

  • `Synthesis from paper + recipe summary: ...`

---

Progressive Disclosure

Do not dump the whole knowledge base unless asked.

Preferred sequence:

1. `INDEX.md` 2. Frontmatter summary and key facts from one chunk 3. Small table or bullet answer 4. Full chunk excerpt summary only if the user wants detail

When a question spans both “paper” and “how to run it,” answer in two blocks:

1. **Paper answer** 2. **Public recipe / reproduction answer**

---

Cross-Skill Handoff

If the user wants to **implement** something, switch from knowledge to pipeline-building:

  • “build a Nano3 SFT pipeline”
  • “how do I run the RL recipe?”
  • “generate the commands/configs”
  • “customize this for my data”
  • “which steps should I chain?”

Then say:

> “This is now a build/customization task. I should hand off to `/nemotron-customize`.”

Use `skills/nemotron-nano3/context/quick-reference.md` to map:

  • paper concept → public recipe stage
  • public recipe stage → `nemotron-customize` step or Explorer-mode fallback

Important caveat:

  • `nemotron-customize` currently has direct catalog support for **packing, SFT, RL, eval, conversion, curation, translation**
  • **Stage 0 pretraining** does **not** yet have a public catalog step in `src/nemotron/steps/STEPS.md`; route that as an **Explorer-mode** or direct recipe task

---

Calibration Examples

Architecture question

User: > How many parameters are active in Nemotron 3 Nano and why is it faster than similarly sized models?

Answer pattern:

1. State the totals: 31.6B total, 3.2B active per forward pass, 3.6B including embeddings 2. Explain sparse MoE + hybrid Mamba/Transformer design 3. Cite `paper/architecture.md`

Reproduction question

User: > Can I reproduce the paper’s SFT and RL results with the public repo?

Answer pattern:

1. Say **not exactly** 2. Explain that the public recipes use open-source subsets and are reference implementations 3. Point to stage summaries and `recipes/overview.md` 4. If they want commands, hand off to

Read more
Ships withnemotron

Open and efficient models for agentic AI. Training recipes, deployment guides, and use-case examples for the Nemotron family.

Get the whole plugin
Stats
2,137
Stars
433
Forks
Active
Maintenance
Jupyter Notebook
Language
Apache-2.0
License
1d ago
Last commit
1y ago
Created
18h ago
Added

Repo: nvidia-nemo/nemotron

Other skills on nemotron.