New to Claude Skills? Learn how to install them →

Tnvidia on GitHub

TP DP PP Communication Overlap

OfficialFree

Optimize Megatron-Bridge for efficient communication overlap.

by nvidia2.8k stars on nvidia/skills
Updated Aug 7, 2026
Get this skill

Free · Opens the source repo

What TP DP PP Communication Overlap does

The TP DP PP Communication Overlap skill provides an operational guide for configuring communication overlap in the Megatron-Bridge framework. This skill is particularly useful for developers working with large-scale distributed training models, as it allows for the fine-tuning of tensor model parallelism (TP), data parallelism (DP), and pipeline parallelism (PP) to maximize performance. By leveraging specific configuration settings, users can enable efficient communication overlap, which is crucial for reducing idle time and improving throughput during training.

The skill includes detailed code snippets and configuration examples that demonstrate how to set up communication overlap effectively. Users can adjust parameters such as the tensor model parallel size and sequence parallelism to suit their specific training requirements. Additionally, the skill outlines potential pitfalls to avoid, ensuring that users can implement these configurations without running into common issues that may arise during setup.

Verification of the configuration can be performed through unit tests provided in the package, allowing users to confirm that their settings are functioning as intended. This skill is ideal for machine learning engineers and researchers who are looking to enhance the performance of their distributed training setups in Megatron-Bridge, particularly when working with large models that require efficient resource management.

Overall, this skill serves as a comprehensive resource for anyone involved in optimizing communication strategies within the Megatron-Bridge framework, making it easier to achieve high-performance training outcomes.

When to use it

Use this skill when configuring Megatron-Bridge for distributed training to ensure optimal communication overlap settings are applied.

When not to use it

This skill may not be suitable for users not working with Megatron-Bridge or those who do not require advanced configuration for communication overlap.

What you can build with it

Configuring a New Model

When starting a new distributed training project with Megatron-Bridge, use this skill to set up communication overlap from the beginning.

Optimizing Existing Training Jobs

If you have ongoing training jobs that are underperforming, apply the configurations from this skill to enhance communication efficiency.

Verifying Training Setup

After configuring your training environment, utilize the verification steps provided in this skill to ensure everything is set up correctly.

How to install TP DP PP Communication Overlap

View source

1. Install with the skills CLI

npx skills add nvidia/skills/nemo-mbridge-perf-tp-dp-comm-overlap --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by nvidia

TP / DP / PP Communication Overlap Skill

For stable background and recommendation level, see:

  • @docs/training/communication-overlap.md

Enablement

Minimal Bridge override:

from megatron.bridge.training.comm_overlap import CommOverlapConfig

cfg.model.tensor_model_parallel_size = 4
cfg.model.sequence_parallel = True
cfg.model.pipeline_model_parallel_size = 4
cfg.model.virtual_pipeline_model_parallel_size = 2

cfg.comm_overlap = CommOverlapConfig(
    tp_comm_overlap=True,
)

cfg.ddp.use_distributed_optimizer = True
cfg.ddp.overlap_grad_reduce = True
cfg.ddp.overlap_param_gather = True

Optional TP preset:

from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048

cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048

Precision knobs belong to mixed precision:

cfg.mixed_precision.grad_reduce_in_fp32 = False
cfg.mixed_precision.fp8_param_gather = False

Code Anchors

Bridge overlap gating:

if self.user_comm_overlap_cfg.tp_comm_overlap is True:
    if model_cfg.tensor_model_parallel_size < 2:
        ...
    elif not model_cfg.sequence_parallel:
        ...
    elif not HAVE_TE:
        ...

PP overlap selection:

if model_cfg.pipeline_model_parallel_size > 1:
    if vp_size > 1:
        comm_overlap_cfg.overlap_p2p_comm = True
        comm_overlap_cfg.batch_p2p_comm = False
    else:
        comm_overlap_cfg.overlap_p2p_comm = False
        comm_overlap_cfg.batch_p2p_comm = True

DP overlap defaults:

if self.data_parallel_size > 1:
    comm_overlap_cfg.bucket_size = 128 * 1024 * 1024
    comm_overlap_cfg.overlap_grad_reduce = True
    comm_overlap_cfg.overlap_param_gather = True

Launch-time env tuning:

executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)
...
executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)

Pitfalls

  1. TP overlap silently disables itself if sequence_parallel=False or Transformer Engine is unavailable.
  2. PP overlap is not enabled for all PP cases. Bridge only auto-selects overlap_p2p_comm=True when PP > 1 and VPP > 1.
  3. bucket_size is a parameter-count knob, not a byte-size knob.
  4. grad_reduce_in_fp32 and fp8_param_gather should be set through mixed precision, not as standalone DDP tuning first.
  5. CUDA_DEVICE_MAX_CONNECTIONS and LayerNorm SM margin are launch-time plugin settings, not CommOverlapConfig fields.

Verification

Use the checked-in overlap unit coverage first:

uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q

Optional second check if nemo_run is available:

uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q

Success criteria:

  • first command reports 26 passed
  • second command validates plugin-owned env wiring when not skipped

Frequently asked questions about TP DP PP Communication Overlap

Similar skills