Open and efficient models for agentic AI. Training recipes, deployment guides, and use-case examples for the Nemotron family.
> /plugin marketplace add nvidia-nemo/nemotron> /plugin install nemotron-customize@nvidia-nemotron
Repo: nvidia-nemo/nemotron
What's inside
Open and efficient models for agentic AI. Training recipes, deployment guides, and use-case examples for the Nemotron family.
πNemotron 3.5 Lightning is now released β a 30B-A3B hybrid Mamba-Transformer MoE with Multi-Token Prediction, built for the high-volume execution layer of long-running agents. See the release blog, the training recipe, and the model weights.
πNemotron 3 Ultra was announced at GTC San Jose 2026. The model is open-source on Hugging Face, and the training recipe is now available in this repo. To learn more, see the usage guide!
πNemotron 3 Nano Omni is now released β a 30B-A3B hybrid Mamba-Transformer MoE with native text, image, video, and audio support, designed as a multimodal perception sub-agent for agentic AI. See the release blog, the training recipe, and the model weights.
| Open Models | Fully transparent training data, techniques, and weights for community innovation |
| Compute Efficiency | Model pruning and optimization enabling higher throughput via TensorRT-LLM |
| High Accuracy | Built on frontier open models with human-aligned reasoning for agentic workflows |
| Flexible Deployment | Deploy anywhere: edge, single GPU, or data center with NIM microservices |
This repo ships a Claude Code plugin called nemotron-customize that turns the step catalog under src/nemotron/steps/ into a guided, repo-native pipeline builder.
Install once:
/plugin marketplace add NVIDIA/Nemotron
/plugin install nemotron-customize@nvidia-nemotron
Then, start Claude Code from the repo root and invoke the skill:
cd /path/to/Nemotron # repo root: must contain pyproject.toml and src/nemotron/steps/
claude
/nemotron-customize
The skill resolves all file paths against your current working directory, so it must be invoked from the Nemotron checkout root. Running it from a subdirectory will cause file reads to fail.
The skill plans the step DAG, validates artifact wiring, and emits the YAML configs needed to run the requested pipeline. See skills/nemotron-customize/SKILL.md for the full contract.
The marketplace installs only
nemotron-customize. The other folders underskills/(model knowledge bases, contributor add-*skills) stay on disk for repo browsing but are not loaded as plugins.
nemotron/
β
βββ src/nemotron/steps/ Modular building blocks for training, eval, SDG, and more
β
βββ src/nemotron/recipes/ Training recipes (complete, reproducible pipelines)
β
βββ usage-cookbook/ Usage cookbooks (deployment and model usage guides)
β
βββ use-case-examples/ Examples of leveraging Nemotron in agentic workflows
| Nemotron Steps | Training Recipes | Usage Cookbooks | Use Case Examples | |
|---|---|---|---|---|
| Purpose | Full lifecycle building blocks, chain data prep, training, eval and other steps | Reproduce full training pipelines from raw data to model | Deploy and use trained models | Build end-to-end applications |
| Format | The nemotron steps CLI and YAML configs | Python packages with configs, scripts, and evaluation | Jupyter notebooks with step-by-step guides | Jupyter notebooks and scripts |
| When to use | You want to run one stage in isolation or compose a custom pipeline | You want to train, fine-tune, or understand how a model was built | You have a model and want to deploy or run inference | You want to build an application (RAG, agents, tool use) |
| Location | src/nemotron/steps/ | src/nemotron/recipes/ | usage-cookbook/ | use-case-examples/ |
NVIDIA Nemotron is a family of open, high-efficiency multimodal models purpose-built for agentic AI.
Model Tiers:
Nemotron models excel at coding, math, scientific reasoning, tool calling, instruction following, and visual reasoning. Deploy across edge, single GPU, or data center environments with support for NeMo, TensorRT-LLM, vLLM, SGLang, and NIM microservices.
A Nemotron step is a named, reusable unit of work that you invoke with the nemotron steps CLI.
Each step packages a description of the work it performs, the artifacts it consumes and produces, and one or more named configurations that supply parameter values.
Steps live under src/nemotron/steps/, and the CLI discovers them at startup.
The training recipes in the next section are composed from these steps. Run a step on its own when you want one stage, or chain steps together when you need a different pipeline shape than the published recipes.
The catalog covers the full training lifecycle.
curate/* and data_prep/*.sdg/*.translate/*.byob/*.pretrain/*, sft/*, peft/*, and rl/*.convert/* and optimize/*.eval/*.env/*.The Nemotron repository provides reproducible training pipelines from raw data to deployment-ready models. These implementations reflect how large language models are actually trained: careful experimentation, validation gates, and systematic optimization.
Training a production model involves interconnected components. Isolated examples miss how stages interact. Complete pipelines show:
Because these are complete systems, you can extract specific techniques with confidence. Each component has been proven to work in context.
| Model | Description | Stages | Guide |
|---|---|---|---|
| Nemotron 3 Ultra | 550B total / 55B active hybrid Mamba-Attention LatentMoE Transformer with MTP and 1M context β NVIDIA's largest Nemotron 3 model for datacenter-scale agentic reasoning | Pretrain β SFT β RLVR β MOPD | Training Guide |
| Nemotron 3 Super | 120.6B total / 12.7B active Hybrid Mamba Latent MoE Transformer for frontier reasoning, coding, and agentic tasks | Pretrain β SFT β RL | Training Guide |
| Nemotron 3 Nano | 31.6B total / 3.6B active MoE Hybrid Mamba-Transformer for agentic reasoning | Pretrain β SFT β RL | Training Guide |
| Nemotron 3.5 Lightning | 30B total / 3B active hybrid Mamba-Transformer MoE with Multi-Token Prediction | Pretrain β SFT β RL β Quantization | Training Guide |
| Nemotron 3 Nano Omni | 30B total / 3B active hybrid Mamba-Transformer MoE β native text, image, video, and audio for agentic multimodal perception | SFT β RL (MPO / text / vision) β Eval | Training Guide |
A training recipe for NVIDIA's largest Nemotron 3 model β a 550B-A55B hybrid Mamba-Attention Mixture-of-Experts Transformer with LatentMoE and multi-token prediction (MTP), pretrained in NVFP4 and extended to 1M-token context for datacenter-scale agentic reasoning.
Open-Source Data Only: These recipes train exclusively on the open-sourced subset of training data. Results will differ from the tech report benchmarks, which used additional proprietary data. Use these recipes as reference implementations to apply the methodology with your own data.
Model Specifications:
What You Can Extract:
Showing a partial view of a very large repo.
FAQ
nemotron is a Claude Code plugin with 10 hand-picked skills for machine learning work, indexed on Flowy. Install it with the command on its page. It includes nemotron-add-model, nemotron-add-pattern, nemotron-add-step. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it