Skip to content
Development
Skill

/nemo-automodel-distributed-training

Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.

From plugin
nvidia-skills
2.8k200 skills3 agents
Install
$ npx -y skills add NVIDIA/skills --skill nemo-automodel-distributed-training --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/nemo-automodel-distributed-training

Context preview

The summary Claude sees to decide when to auto-load this skill.

Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.

SKILL.md

nemo-automodel-distributed-training.SKILL.md
name: nemo-automodel-distributed-training
description: Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
when_to_use: Adding or modifying distributed training strategies (FSDP2, HSDP, DDP), debugging multi-GPU or multi-node failures, configuring context or tensor parallelism, or tuning sharding settings.
license: Apache-2.0
metadata:
  author: NVIDIA
  tags:
    - nemo-automodel
    - distributed-training

Distributed Training in NeMo AutoModel

Purpose

NeMo AutoModel uses PyTorch-native distributed training. All parallelism is orchestrated through a single `MeshContext` object that holds device meshes, strategy configs, and axis names. <!-- NVSkills catalog signing requested after PR #2937 (2026-07-31). -->

Instructions

For conceptual distributed-training questions, answer directly from the quick patterns in this skill without inspecting the repository. Start with the strategy choice, then list only the YAML fields and constraints relevant to the question.

Use direct action verbs in the final answer: recommend the strategy, show the minimal YAML, state the sizing constraint, and name the unsupported strategies. Do not discuss model onboarding, recipes, Slurm, SkyPilot, or checkpointing unless the user asks.

Examples

TP plus PP for a large multi-node model

Recommend `strategy: fsdp2`. Mention `tp_size`, `pp_size`, `cp_size`, `ep_size`, and the `pipeline` sub-config. State that `dp_size` is inferred from `world_size / (tp_size * pp_size * cp_size)`.

distributed:
  strategy: fsdp2
  tp_size: 8
  pp_size: 4
  cp_size: 1
  ep_size: 1
  pipeline:
    pp_schedule: interleaved1f1b
    pp_microbatch_size: 1

MoE expert parallelism

Recommend `strategy: fsdp2` with `ep_size > 1`. Say this creates a separate `moe_mesh`; include the `moe` sub-config when relevant; state that `ep_size` must divide `dp_size * cp_size`. Do not recommend `megatron_fsdp` or `ddp`.

distributed:
  strategy: fsdp2
  ep_size: 8
  moe:
    reshard_after_forward: false

MegatronFSDP limitations

Say no for pipeline parallelism, expert parallelism, and `sequence_parallel`. Recommend `fsdp2` for PP, EP, or `sequence_parallel`; mention that DDP is only simple data parallelism.

Strategy Selection

Three strategies are available, selected via the `distributed.strategy` YAML key:

| Strategy | YAML value | Best for | |---|---|---| | FSDP2 | `fsdp2` | General use, recommended default. Supports TP, PP, CP, EP, HSDP. | | MegatronFSDP | `megatron_fsdp` | NVIDIA Megatron-style FSDP. No PP, no EP, no sequence_parallel. | | DDP | `ddp` | Simple data parallelism only. No TP, PP, CP, or EP. |

Decision tree:

  • Single GPU: no distributed config needed (FSDP2Manager skips parallelization when world_size=1).
  • Multi-GPU single node: `fsdp2` (default). Use `ddp` only if you need the simplest possible setup.
  • Multi-node: `fsdp2` with appropriate TP/PP sizing.
  • MoE models with expert parallelism: `fsdp2` with `ep_size > 1` (creates a separate `moe_mesh`).
  • Large models (70B+): `fsdp2` with PP + TP.
  • Long sequences (8K+): add CP (`cp_size > 1`).

When answering strategy-selection questions, state the chosen `distributed.strategy` first, then enumerate the YAML fields the user must set.

Quick TP + PP answer:

  • Use `strategy: fsdp2`; do not use `megatron_fsdp` when pipeline parallelism is required.
  • Set `tp_size` for tensor parallelism and `pp_size` for pipeline parallelism.
  • Add a `pipeline:` sub-config with `pp_schedule` and `pp_microbatch_size`.
  • Leave `dp_size` unset or `none`; it is inferred as `world_size / (tp_size * pp_size * cp_size)`.
  • Keep TP inside a fast intra-node domain when possible, and use PP across model depth for 70B+ models.

Quick MoE expert-parallel answer:

  • Start with `strategy: fsdp2` and `ep_size > 1`.
  • Include a `moe:` sub-config only when `ep_size > 1`; it maps to `MoEParallelizerConfig`.
  • Expect a separate `moe_mesh` for expert parallelism in addition to the main `device_mesh`.
  • Do not recommend `megatron_fsdp` or `ddp` for expert parallelism; `megatron_fsdp` has no EP support.
  • Before finishing an MoE EP answer, explicitly state that `ep_size` must divide `dp_size * cp_size` and that `megatron_fsdp` does not support EP, PP, or `sequence_parallel`.

YAML Config Structure

The `distributed` section in the recipe YAML maps directly to `parse_distributed_section()` in `recipes/_dist_utils.py`:

distributed:
  strategy: fsdp2           # fsdp2 | megatron_fsdp | ddp
  dp_size: none             # auto-calculated from world_size / (tp * pp * cp)
  dp_replicate_size: none   # FSDP2-only, for HSDP
  tp_size: 1
  pp_size: 1
  cp_size: 1
  ep_size: 1

  # Strategy-specific flags (forwarded to the strategy dataclass):
  sequence_parallel: false
  activation_checkpointing: false
  defer_fsdp_grad_sync: true   # FSDP2 only

  # Sub-configs (optional):
  pipeline:
    pp_schedule: 1f1b
    pp_microbatch_size: 1
    # ... see PipelineConfig fields

  moe:
    reshard_after_forward: false
    # ... see MoEParallelizerConfig fields

The `dp_size` is always inferred:

dp_size = world_size / (tp_size * pp_size * cp_size)

Infrastructure Flow

initialize_distributed()                       [components/distributed/init_utils.py]
    -> initializes torch.distributed process group and returns DistInfo
YAML distributed section + DistInfo.world_size
    -> parse_distributed_section()          [recipes/_dist_utils.py]
    -> create_distributed_setup_from_config()              [recipes/_dist_utils.py]
        -> DistributedSetup.build()         [components/distributed/config.py]
    -> instantiate_infrastructure()         [_transformers/infrastructure.py]
        -> _instantiate_distributed()       -> FSDP2Manager / MegatronFSDPManager / DDPManager
        -> _instantiate_pipeline()          -> AutoPipeline (if pp_size > 1)
        -> pa
Read more
Ships withnvidia-skills

Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.

Get the whole plugin