
MoE Optimization Workflow
OfficialFreeStreamline MoE model training with systematic guidance.
Free · Opens the source repo
What MoE Optimization Workflow does
The MoE Optimization Workflow skill provides a structured approach to optimizing Mixture of Experts (MoE) training within the Megatron Bridge framework. It is based on the principles outlined in the Megatron-Core MoE paper, specifically focusing on the Three Walls framework: memory, communication, and compute/host overhead. This skill is ideal for developers and researchers working with large-scale machine learning models who need to enhance their training efficiency and resource utilization.
This skill emphasizes an iterative tuning process that starts by making the model memory-feasible, followed by scaling, profiling, and retuning. The workflow guides users through critical phases, helping them prioritize actions based on the current bottleneck. For instance, it recommends using selective recompute to manage memory effectively before resorting to more resource-intensive methods. Each phase is detailed with specific strategies, including parallelism choices and profiling techniques, enabling users to systematically address performance issues.
Additionally, the skill provides concrete recommendations for dispatcher choices and CUDA graph usage, which are essential for optimizing communication and compute efficiency. It includes practical examples of parallel folding configurations and default mappings tailored for different hardware setups, ensuring users can adapt the workflow to their specific environments. This structured approach not only improves model training times but also reduces resource waste, making it a valuable tool for anyone involved in developing or optimizing MoE models.
In summary, the MoE Optimization Workflow skill is designed for those looking to enhance their model training processes through a systematic and well-documented methodology. By following the outlined steps and leveraging the provided guidance, users can effectively navigate the complexities of MoE training optimization.
When to use it
Use this skill when working on MoE model training in Megatron Bridge and needing to optimize performance.
When not to use it
This skill is not suitable for users unfamiliar with MoE concepts or those not using the Megatron framework.
What you can build with it
Optimizing Large-Scale MoE Models
When training large MoE models, use this skill to systematically address memory and communication bottlenecks.
Improving Training Efficiency
Leverage the structured workflow to enhance training times and resource utilization during model development.
Adapting to Hardware Changes
Utilize the provided default mappings and parallel folding configurations to optimize performance on various hardware setups.
How to install MoE Optimization Workflow
View source1. Install with the skills CLI
npx skills add nvidia/skills/nemo-mbridge-perf-moe-optimization-workflow --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by nvidiaMoE Training Optimization Workflow
Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-optimization-workflow/card.yaml Source: Scalable Training of MoE Models with Megatron Core
Quick Reference
Think in terms of the paper's Three Walls:
- memory wall
- communication wall
- compute and host-overhead wall
MoE tuning is iterative. Fixing one wall usually exposes the next one, so the best workflow is: fit first, scale second, profile third, then retune.
First Answer Checklist
For MoE optimization workflow prompts, present the response in this order:
- Fit: make the model memory-feasible first. Use the smallest model
parallelism that fits, prefer selective recompute before full recompute, add
offloading only after recompute and parallelism are insufficient, and use
--fake-init-process-groupto sanity-check large layouts. - Scale: maximize DP after the model fits, keep hot communication inside the fastest interconnect, use PP plus VPP for multi-node scaling, prefer EP over extra TP for expert layers, and add CP when long context makes attention memory dominant.
- Profile: identify the dominant wall: memory, communication, host overhead, or compute.
- Retune: change dispatcher, overlap, FP8 mode, CUDA graphs, or recompute based on the profiled bottleneck.
- Include the exact Parallel Folding meshes:
Attention: TP x CP x DP x PPandMoE: ETP x EP x EDP x PP. - Include the default mappings:
alltoallfor safe bring-up,flex+deepepfor H100/B200-style systems,flex+hybridepfor GB200/GB300/NVL72 systems, Hopper to FP8 blockwise, Blackwell to MXFP8, and dropless MoE TE-scoped CUDA graphs overattn,moe_router, andmoe_preprocess.
Phase 1: Make The Run Memory-Feasible
Start with a configuration that fits reliably before chasing throughput.
Recommended order:
- Use the smallest amount of model parallelism that still fits.
- Turn on selective recompute before falling back to full recompute.
- Add offloading only when recompute and parallelism are still insufficient.
- Use
--fake-init-process-groupto sanity-check large parallel layouts on a single GPU before burning cluster time.
Recompute guidance
Prefer selective recompute for MoE runs:
- good first choices:
layernorm,core_attn,moe_act,mlp, or model-specific modules (shared_experts,mla_up_proj) - use full recompute only when the run still does not fit
- revisit recompute after enabling CUDA graphs, because some graph scopes and full recompute paths do not mix well
As a rule of thumb, fine-grained recompute often recovers most of the needed memory while keeping throughput much closer to the non-recompute baseline than full-layer recompute does.
Phase 2: Choose Parallelism For Scale
Priority order:
- Maximize DP once the model fits.
- Keep the hot communication path inside the fast interconnect when possible.
- Use PP, plus VPP if needed, for multi-node scaling.
- Prefer EP over extra TP for expert layers.
- Add CP for long context once sequence length makes attention memory dominant.
Parallel Folding
Parallel Folding decouples attention and MoE parallelism so you do not have to pick a single compromise layout:
Attention: TP × CP × DP × PP
MoE: ETP × EP × EDP × PP
Key knobs:
--expert-model-parallel-size--expert-tensor-parallel-size
Use it when attention prefers some TP or CP, but expert layers benefit from a larger EP degree than the dense layers can tolerate.
Phase 3: Profile The Dominant Bottleneck
| Bottleneck | What it looks like | Primary fixes |
|---|---|---|
| Memory | Run fits only with aggressive full recompute or OOMs during warmup | selective recompute, FP8, offloading, better PP layout |
| Communication | Nsight shows large all-to-all or collective blocks | DeepEP or HybridEP, EP overlap, DP/TP overlap, better PP layout |
| Host overhead | GPU gaps, launch-bound traces, Python overhead | CUDA graphs, --manual-gc, higher MBS, CPU affinity tuning |
| Compute | Low SM utilization after comm and host issues are addressed | grouped GEMM, fusion work, FP8, dispatcher-specific kernel tuning |
Dispatcher And Overlap Guidance
Use dispatcher choice as a bottleneck fix, not as the first tuning knob.
moe_token_dispatcher_type="alltoall": safest bring-up path, fine for smaller EP sizesmoe_token_dispatcher_type="flex"+moe_flex_dispatcher_backend="deepep": strong default for H100 and B200 style deploymentsmoe_token_dispatcher_type="flex"+moe_flex_dispatcher_backend="hybridep": strongest starting point on GB200 or GB300 NVL72 systems
If the all-to-all path is visible in profiles, combine dispatcher tuning with:
--overlap-moe-expert-parallel-comm--overlap-grad-reduce--tp-comm-overlap
FP8 Recipe Quick Decision
| Platform | Recommended starting recipe |
|---|---|
| Hopper | FP8 blockwise |
| Blackwell | MXFP8 |
| Blackwell, speed-first exploration | NVFP4 after the BF16 or FP8 path is stable |
Keep the router in FP32. The largest wins usually come from expert GEMMs and other heavy matrix math, not from trying to quantize every small MoE component.
CUDA Graphs For MoE
For dropless MoE, start with partial TE-scoped graphs:
attnmoe_routermoe_preprocess
That path usually gives a meaningful step-time win while keeping the dynamic expert work outside the graph. Expect a moderate speedup when launch overhead is visible, but budget several extra GB of memory and verify that shapes remain static.
Use full-iteration graphs only for graph-friendly workloads such as drop-and-pad or tightly controlled static-shape experiments.
Related references:
- @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md
- @docs/training/cuda-graphs.md
- @docs/training/activation-recomputation.md
Pitfalls
-
Do not optimize in the wrong order: fitting the model and selecting sane parallelism matter more than micro-optimizations.
-
Platform changes the limiting wall: H100-class runs often feel more communication-bound, while GB200 or GB300 runs often expose CPU or launch overhead earlier.
-
FP8 MFU can look misleadingly low: compare absolute throughput as well as MFU when switching precision modes.
-
CUDA graphs and recompute interact: TE-scoped graphs are usually paired with selective recompute, not blanket full recompute.
-
Parallel Folding is not optional at large scale: once attention and expert layers want clearly different layouts, a single shared TP or EP plan becomes a tax on both.
Frequently asked questions about MoE Optimization Workflow
Similar skills
Heap Snapshot Analysis
Investigate V8 heap snapshots for memory issues.
VS Code Performance Workflow
Automate performance investigations in VS Code.
Memory Leak Audit
Prevent memory leaks with effective coding patterns.
CPU Profile Analysis
Analyze V8 and Chrome performance profiles for optimization.
Chat Performance Testing
Benchmark and validate chat UI performance in VS Code.
Vercel React Best Practices
Optimize your React and Next.js applications for performance.
