
TP DP PP Communication Overlap
OfficialFreeOptimize Megatron-Bridge for efficient communication overlap.
Free · Opens the source repo
What TP DP PP Communication Overlap does
The TP DP PP Communication Overlap skill provides an operational guide for configuring communication overlap in the Megatron-Bridge framework. This skill is particularly useful for developers working with large-scale distributed training models, as it allows for the fine-tuning of tensor model parallelism (TP), data parallelism (DP), and pipeline parallelism (PP) to maximize performance. By leveraging specific configuration settings, users can enable efficient communication overlap, which is crucial for reducing idle time and improving throughput during training.
The skill includes detailed code snippets and configuration examples that demonstrate how to set up communication overlap effectively. Users can adjust parameters such as the tensor model parallel size and sequence parallelism to suit their specific training requirements. Additionally, the skill outlines potential pitfalls to avoid, ensuring that users can implement these configurations without running into common issues that may arise during setup.
Verification of the configuration can be performed through unit tests provided in the package, allowing users to confirm that their settings are functioning as intended. This skill is ideal for machine learning engineers and researchers who are looking to enhance the performance of their distributed training setups in Megatron-Bridge, particularly when working with large models that require efficient resource management.
Overall, this skill serves as a comprehensive resource for anyone involved in optimizing communication strategies within the Megatron-Bridge framework, making it easier to achieve high-performance training outcomes.
When to use it
Use this skill when configuring Megatron-Bridge for distributed training to ensure optimal communication overlap settings are applied.
When not to use it
This skill may not be suitable for users not working with Megatron-Bridge or those who do not require advanced configuration for communication overlap.
What you can build with it
Configuring a New Model
When starting a new distributed training project with Megatron-Bridge, use this skill to set up communication overlap from the beginning.
Optimizing Existing Training Jobs
If you have ongoing training jobs that are underperforming, apply the configurations from this skill to enhance communication efficiency.
Verifying Training Setup
After configuring your training environment, utilize the verification steps provided in this skill to ensure everything is set up correctly.
How to install TP DP PP Communication Overlap
View source1. Install with the skills CLI
npx skills add nvidia/skills/nemo-mbridge-perf-tp-dp-comm-overlap --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by nvidiaTP / DP / PP Communication Overlap Skill
For stable background and recommendation level, see:
- @docs/training/communication-overlap.md
Enablement
Minimal Bridge override:
from megatron.bridge.training.comm_overlap import CommOverlapConfig
cfg.model.tensor_model_parallel_size = 4
cfg.model.sequence_parallel = True
cfg.model.pipeline_model_parallel_size = 4
cfg.model.virtual_pipeline_model_parallel_size = 2
cfg.comm_overlap = CommOverlapConfig(
tp_comm_overlap=True,
)
cfg.ddp.use_distributed_optimizer = True
cfg.ddp.overlap_grad_reduce = True
cfg.ddp.overlap_param_gather = True
Optional TP preset:
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
Precision knobs belong to mixed precision:
cfg.mixed_precision.grad_reduce_in_fp32 = False
cfg.mixed_precision.fp8_param_gather = False
Code Anchors
Bridge overlap gating:
if self.user_comm_overlap_cfg.tp_comm_overlap is True:
if model_cfg.tensor_model_parallel_size < 2:
...
elif not model_cfg.sequence_parallel:
...
elif not HAVE_TE:
...
PP overlap selection:
if model_cfg.pipeline_model_parallel_size > 1:
if vp_size > 1:
comm_overlap_cfg.overlap_p2p_comm = True
comm_overlap_cfg.batch_p2p_comm = False
else:
comm_overlap_cfg.overlap_p2p_comm = False
comm_overlap_cfg.batch_p2p_comm = True
DP overlap defaults:
if self.data_parallel_size > 1:
comm_overlap_cfg.bucket_size = 128 * 1024 * 1024
comm_overlap_cfg.overlap_grad_reduce = True
comm_overlap_cfg.overlap_param_gather = True
Launch-time env tuning:
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)
...
executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
Pitfalls
- TP overlap silently disables itself if
sequence_parallel=Falseor Transformer Engine is unavailable. - PP overlap is not enabled for all PP cases. Bridge only auto-selects
overlap_p2p_comm=TruewhenPP > 1andVPP > 1. bucket_sizeis a parameter-count knob, not a byte-size knob.grad_reduce_in_fp32andfp8_param_gathershould be set through mixed precision, not as standalone DDP tuning first.CUDA_DEVICE_MAX_CONNECTIONSand LayerNorm SM margin are launch-time plugin settings, notCommOverlapConfigfields.
Verification
Use the checked-in overlap unit coverage first:
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q
Optional second check if nemo_run is available:
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q
Success criteria:
- first command reports
26 passed - second command validates plugin-owned env wiring when not skipped
Frequently asked questions about TP DP PP Communication Overlap
Similar skills
Heap Snapshot Analysis
Investigate V8 heap snapshots for memory issues.
VS Code Performance Workflow
Automate performance investigations in VS Code.
Memory Leak Audit
Prevent memory leaks with effective coding patterns.
CPU Profile Analysis
Analyze V8 and Chrome performance profiles for optimization.
Chat Performance Testing
Benchmark and validate chat UI performance in VS Code.
Vercel React Best Practices
Optimize your React and Next.js applications for performance.
