/nemo-automodel-distributed-training
Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
$ npx -y skills add NVIDIA/skills --skill nemo-automodel-distributed-training --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/nemo-automodel-distributed-training
Context preview
The summary Claude sees to decide when to auto-load this skill.
Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
SKILL.md
nemo-automodel-distributed-training.SKILL.mdname: nemo-automodel-distributed-training
description: Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
when_to_use: Adding or modifying distributed training strategies (FSDP2, HSDP, DDP), debugging multi-GPU or multi-node failures, configuring context or tensor parallelism, or tuning sharding settings.
license: Apache-2.0
metadata:
author: NVIDIA
tags:
- nemo-automodel
- distributed-trainingDistributed Training in NeMo AutoModel
Purpose
NeMo AutoModel uses PyTorch-native distributed training. All parallelism is orchestrated through a single `MeshContext` object that holds device meshes, strategy configs, and axis names. <!-- NVSkills catalog signing requested after PR #2937 (2026-07-31). -->
Instructions
For conceptual distributed-training questions, answer directly from the quick patterns in this skill without inspecting the repository. Start with the strategy choice, then list only the YAML fields and constraints relevant to the question.
Use direct action verbs in the final answer: recommend the strategy, show the minimal YAML, state the sizing constraint, and name the unsupported strategies. Do not discuss model onboarding, recipes, Slurm, SkyPilot, or checkpointing unless the user asks.
Examples
TP plus PP for a large multi-node model
Recommend `strategy: fsdp2`. Mention `tp_size`, `pp_size`, `cp_size`, `ep_size`, and the `pipeline` sub-config. State that `dp_size` is inferred from `world_size / (tp_size * pp_size * cp_size)`.
distributed:
strategy: fsdp2
tp_size: 8
pp_size: 4
cp_size: 1
ep_size: 1
pipeline:
pp_schedule: interleaved1f1b
pp_microbatch_size: 1MoE expert parallelism
Recommend `strategy: fsdp2` with `ep_size > 1`. Say this creates a separate `moe_mesh`; include the `moe` sub-config when relevant; state that `ep_size` must divide `dp_size * cp_size`. Do not recommend `megatron_fsdp` or `ddp`.
distributed:
strategy: fsdp2
ep_size: 8
moe:
reshard_after_forward: falseMegatronFSDP limitations
Say no for pipeline parallelism, expert parallelism, and `sequence_parallel`. Recommend `fsdp2` for PP, EP, or `sequence_parallel`; mention that DDP is only simple data parallelism.
Strategy Selection
Three strategies are available, selected via the `distributed.strategy` YAML key:
| Strategy | YAML value | Best for | |---|---|---| | FSDP2 | `fsdp2` | General use, recommended default. Supports TP, PP, CP, EP, HSDP. | | MegatronFSDP | `megatron_fsdp` | NVIDIA Megatron-style FSDP. No PP, no EP, no sequence_parallel. | | DDP | `ddp` | Simple data parallelism only. No TP, PP, CP, or EP. |
Decision tree:
- Single GPU: no distributed config needed (FSDP2Manager skips parallelization when world_size=1).
- Multi-GPU single node: `fsdp2` (default). Use `ddp` only if you need the simplest possible setup.
- Multi-node: `fsdp2` with appropriate TP/PP sizing.
- MoE models with expert parallelism: `fsdp2` with `ep_size > 1` (creates a separate `moe_mesh`).
- Large models (70B+): `fsdp2` with PP + TP.
- Long sequences (8K+): add CP (`cp_size > 1`).
When answering strategy-selection questions, state the chosen `distributed.strategy` first, then enumerate the YAML fields the user must set.
Quick TP + PP answer:
- Use `strategy: fsdp2`; do not use `megatron_fsdp` when pipeline parallelism is required.
- Set `tp_size` for tensor parallelism and `pp_size` for pipeline parallelism.
- Add a `pipeline:` sub-config with `pp_schedule` and `pp_microbatch_size`.
- Leave `dp_size` unset or `none`; it is inferred as `world_size / (tp_size * pp_size * cp_size)`.
- Keep TP inside a fast intra-node domain when possible, and use PP across model depth for 70B+ models.
Quick MoE expert-parallel answer:
- Start with `strategy: fsdp2` and `ep_size > 1`.
- Include a `moe:` sub-config only when `ep_size > 1`; it maps to `MoEParallelizerConfig`.
- Expect a separate `moe_mesh` for expert parallelism in addition to the main `device_mesh`.
- Do not recommend `megatron_fsdp` or `ddp` for expert parallelism; `megatron_fsdp` has no EP support.
- Before finishing an MoE EP answer, explicitly state that `ep_size` must divide `dp_size * cp_size` and that `megatron_fsdp` does not support EP, PP, or `sequence_parallel`.
YAML Config Structure
The `distributed` section in the recipe YAML maps directly to `parse_distributed_section()` in `recipes/_dist_utils.py`:
distributed:
strategy: fsdp2 # fsdp2 | megatron_fsdp | ddp
dp_size: none # auto-calculated from world_size / (tp * pp * cp)
dp_replicate_size: none # FSDP2-only, for HSDP
tp_size: 1
pp_size: 1
cp_size: 1
ep_size: 1
# Strategy-specific flags (forwarded to the strategy dataclass):
sequence_parallel: false
activation_checkpointing: false
defer_fsdp_grad_sync: true # FSDP2 only
# Sub-configs (optional):
pipeline:
pp_schedule: 1f1b
pp_microbatch_size: 1
# ... see PipelineConfig fields
moe:
reshard_after_forward: false
# ... see MoEParallelizerConfig fieldsThe `dp_size` is always inferred:
dp_size = world_size / (tp_size * pp_size * cp_size)
Infrastructure Flow
initialize_distributed() [components/distributed/init_utils.py]
-> initializes torch.distributed process group and returns DistInfo
YAML distributed section + DistInfo.world_size
-> parse_distributed_section() [recipes/_dist_utils.py]
-> create_distributed_setup_from_config() [recipes/_dist_utils.py]
-> DistributedSetup.build() [components/distributed/config.py]
-> instantiate_infrastructure() [_transformers/infrastructure.py]
-> _instantiate_distributed() -> FSDP2Manager / MegatronFSDPManager / DDPManager
-> _instantiate_pipeline() -> AutoPipeline (if pp_size > 1)
-> paRead more
name: nemo-automodel-distributed-training
description: Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
when_to_use: Adding or modifying distributed training strategies (FSDP2, HSDP, DDP), debugging multi-GPU or multi-node failures, configuring context or tensor parallelism, or tuning sharding settings.
license: Apache-2.0
metadata:
author: NVIDIA
tags:
- nemo-automodel
- distributed-trainingDistributed Training in NeMo AutoModel
Purpose
NeMo AutoModel uses PyTorch-native distributed training. All parallelism is orchestrated through a single `MeshContext` object that holds device meshes, strategy configs, and axis names. <!-- NVSkills catalog signing requested after PR #2937 (2026-07-31). -->
Instructions
For conceptual distributed-training questions, answer directly from the quick patterns in this skill without inspecting the repository. Start with the strategy choice, then list only the YAML fields and constraints relevant to the question.
Use direct action verbs in the final answer: recommend the strategy, show the minimal YAML, state the sizing constraint, and name the unsupported strategies. Do not discuss model onboarding, recipes, Slurm, SkyPilot, or checkpointing unless the user asks.
Examples
TP plus PP for a large multi-node model
Recommend `strategy: fsdp2`. Mention `tp_size`, `pp_size`, `cp_size`, `ep_size`, and the `pipeline` sub-config. State that `dp_size` is inferred from `world_size / (tp_size * pp_size * cp_size)`.
distributed:
strategy: fsdp2
tp_size: 8
pp_size: 4
cp_size: 1
ep_size: 1
pipeline:
pp_schedule: interleaved1f1b
pp_microbatch_size: 1MoE expert parallelism
Recommend `strategy: fsdp2` with `ep_size > 1`. Say this creates a separate `moe_mesh`; include the `moe` sub-config when relevant; state that `ep_size` must divide `dp_size * cp_size`. Do not recommend `megatron_fsdp` or `ddp`.
distributed:
strategy: fsdp2
ep_size: 8
moe:
reshard_after_forward: falseMegatronFSDP limitations
Say no for pipeline parallelism, expert parallelism, and `sequence_parallel`. Recommend `fsdp2` for PP, EP, or `sequence_parallel`; mention that DDP is only simple data parallelism.
Strategy Selection
Three strategies are available, selected via the `distributed.strategy` YAML key:
| Strategy | YAML value | Best for | |---|---|---| | FSDP2 | `fsdp2` | General use, recommended default. Supports TP, PP, CP, EP, HSDP. | | MegatronFSDP | `megatron_fsdp` | NVIDIA Megatron-style FSDP. No PP, no EP, no sequence_parallel. | | DDP | `ddp` | Simple data parallelism only. No TP, PP, CP, or EP. |
Decision tree:
- Single GPU: no distributed config needed (FSDP2Manager skips parallelization when world_size=1).
- Multi-GPU single node: `fsdp2` (default). Use `ddp` only if you need the simplest possible setup.
- Multi-node: `fsdp2` with appropriate TP/PP sizing.
- MoE models with expert parallelism: `fsdp2` with `ep_size > 1` (creates a separate `moe_mesh`).
- Large models (70B+): `fsdp2` with PP + TP.
- Long sequences (8K+): add CP (`cp_size > 1`).
When answering strategy-selection questions, state the chosen `distributed.strategy` first, then enumerate the YAML fields the user must set.
Quick TP + PP answer:
- Use `strategy: fsdp2`; do not use `megatron_fsdp` when pipeline parallelism is required.
- Set `tp_size` for tensor parallelism and `pp_size` for pipeline parallelism.
- Add a `pipeline:` sub-config with `pp_schedule` and `pp_microbatch_size`.
- Leave `dp_size` unset or `none`; it is inferred as `world_size / (tp_size * pp_size * cp_size)`.
- Keep TP inside a fast intra-node domain when possible, and use PP across model depth for 70B+ models.
Quick MoE expert-parallel answer:
- Start with `strategy: fsdp2` and `ep_size > 1`.
- Include a `moe:` sub-config only when `ep_size > 1`; it maps to `MoEParallelizerConfig`.
- Expect a separate `moe_mesh` for expert parallelism in addition to the main `device_mesh`.
- Do not recommend `megatron_fsdp` or `ddp` for expert parallelism; `megatron_fsdp` has no EP support.
- Before finishing an MoE EP answer, explicitly state that `ep_size` must divide `dp_size * cp_size` and that `megatron_fsdp` does not support EP, PP, or `sequence_parallel`.
YAML Config Structure
The `distributed` section in the recipe YAML maps directly to `parse_distributed_section()` in `recipes/_dist_utils.py`:
distributed:
strategy: fsdp2 # fsdp2 | megatron_fsdp | ddp
dp_size: none # auto-calculated from world_size / (tp * pp * cp)
dp_replicate_size: none # FSDP2-only, for HSDP
tp_size: 1
pp_size: 1
cp_size: 1
ep_size: 1
# Strategy-specific flags (forwarded to the strategy dataclass):
sequence_parallel: false
activation_checkpointing: false
defer_fsdp_grad_sync: true # FSDP2 only
# Sub-configs (optional):
pipeline:
pp_schedule: 1f1b
pp_microbatch_size: 1
# ... see PipelineConfig fields
moe:
reshard_after_forward: false
# ... see MoEParallelizerConfig fieldsThe `dp_size` is always inferred:
dp_size = world_size / (tp_size * pp_size * cp_size)
Infrastructure Flow
initialize_distributed() [components/distributed/init_utils.py]
-> initializes torch.distributed process group and returns DistInfo
YAML distributed section + DistInfo.world_size
-> parse_distributed_section() [recipes/_dist_utils.py]
-> create_distributed_setup_from_config() [recipes/_dist_utils.py]
-> DistributedSetup.build() [components/distributed/config.py]
-> instantiate_infrastructure() [_transformers/infrastructure.py]
-> _instantiate_distributed() -> FSDP2Manager / MegatronFSDPManager / DDPManager
-> _instantiate_pipeline() -> AutoPipeline (if pp_size > 1)
-> paOfficial, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
Other skills on nvidia-skills.
- /nvidia-skill-finder
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. Trigger on NVIDIA products, hardware, software, SDKs, GPUs, Jetson/JetPack/L4T/BSP/SDK Manager/driver/flashing/setup, CUDA, NIM, NeMo, Omniverse/OpenUSD/SimReady,
Open skill - /accelerated-computing-cudf
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
Open skill - /aiq-deploy
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Open skill - /aiq-research
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Open skill - /amc-run-sample-calibration
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample dataset', 'run sample calibration', 'verify AMC install', or 'launch and test'.
Open skill - /amc-run-video-calibration
Calibrate a new dataset from pre-recorded video files via the AutoMagicCalib REST API. Use when user has local MP4s and says 'calibrate my videos', 'run AMC on these videos', or similar. For RTSP/live streams, use amc-run-rtsp-calibration instead.
Open skill

