
NeMo AutoModel Onboarding
OfficialFreeStreamline the integration of new model architectures.
Free · Opens the source repo
What NeMo AutoModel Onboarding does
The NeMo AutoModel Onboarding skill provides a structured guide for developers looking to integrate new model architectures into the NeMo AutoModel framework. This skill outlines a clear five-phase onboarding process, ensuring that users can efficiently classify and implement new models. By following the instructions, users can navigate through the complexities of architecture discovery, implementation patterns, registration, and validation without confusion.
This skill is particularly useful for machine learning engineers and researchers who are working with various model architectures, especially those sourced from the Hugging Face Hub. The onboarding process includes specific steps such as fetching the model's configuration, identifying the model type, and checking for existing similar architectures. Each step is designed to minimize errors and streamline the incorporation of new models into existing workflows.
The skill emphasizes direct action verbs to guide users through the onboarding process. For example, users will classify the model, name the necessary implementation files, and register the model in the appropriate registry. This hands-on approach allows for a practical understanding of the integration process, making it easier for users to adapt to the NeMo AutoModel framework.
In addition to practical onboarding steps, the skill also provides reference patterns for common model types, such as Dense LLMs, MoE LLMs, and VLMs. This ensures that users have access to the necessary resources to handle specific architecture requirements effectively. Overall, this skill is an essential tool for any developer or researcher aiming to enhance their model integration capabilities within NeMo AutoModel.
When to use it
Use this skill when you need to add or modify support for model architectures in NeMo AutoModel, such as integrating new models from Hugging Face.
When not to use it
This skill is not suitable for standalone training recipe questions or general configuration issues unrelated to model architecture onboarding.
What you can build with it
Integrating a New Causal LM
Use this skill to add support for a new Hugging Face causal language model by following the structured onboarding steps.
Mapping MoE Router Weights
Utilize the skill to accurately map MoE router and expert weights from a Hugging Face checkpoint during the onboarding process.
Registering a New Model Class
Leverage this skill to register a new model class in NeMo AutoModel, ensuring all necessary components are included.
How to install NeMo AutoModel Onboarding
View source1. Install with the skills CLI
npx skills add nvidia/skills/nemo-automodel-model-onboarding --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by nvidiaAdding Model Support to NeMo AutoModel
Purpose
This skill guides implementation of new model architectures in NeMo AutoModel. Follow the five phases in order.
<!-- NVSkills signature refresh requested after PR #2998 (2026-07-31). -->Instructions
When answering an onboarding question, keep the response in this order:
- Classify the architecture from
config.json. - Name the exact implementation files under
components/models/<name>/. - Identify registry and optional custom-config updates.
- State the validation tests that must be added before full checkpoint use.
For conceptual onboarding questions, answer from this skill without opening the pattern files unless the user asks you to edit code. Mention pattern filenames as references, then give the direct checklist.
Use direct action verbs: classify the model, name the files, map the weights, register the class, and add tests. Do not discuss distributed strategy, launcher configuration, or general recipe authoring unless the user explicitly connects it to onboarding a new architecture.
Examples
Use these compact answer patterns for common questions:
- Dense causal LM: classify as dense only when
architecturescontains aForCausalLMclass and expert fields such asnum_local_experts,n_routed_experts, ornum_experts_per_tokare absent. Createcomponents/models/<name>/model.py,state_dict_adapter.py,__init__.py, and optionalconfig.py, registerMODEL_ARCH_MAPPINGin_transformers/registry.py, add example YAML, and add tiny-config unit tests plus layer-equivalence tests for rewritten layers. - MoE state dict: identify expert fields in
config.json, referencemoe-patterns.md, map router tensors separately, preserve routed-expert index order, map routed experts, shared experts, and gate/up/down projections, add adapter key-map tests and tiny-config numerical equivalence tests, and do not rely only onfrom_pretrained()or silent tensor reshapes. - VLM onboarding: classify as VLM only when
vision_config,text_config, and aForConditionalGenerationarchitecture are present. Referencevlm-patterns.mdand existing VLM implementations such asmistral4,kimivl, orkimi_k25_vl; check text backbone, vision tower, projector, processor assumptions, text and visionstate_dict_adapter.pymappings, registry registration, and tiny image-text tests before full checkpoints. Do not treat VLM onboarding as a pure causal-LM path or skip processor/image tests.
For MoE state-dict and VLM questions, apply the checklists in Sections 2.4 and 2.5.
Routing Boundary
Use this skill only when the user is adding or modifying model architecture support: model files, custom layers, state-dict adapters, Hugging Face config mapping, registry entries, or model capability flags.
Do not use this skill for standalone training recipe YAML questions about optimizers, datasets, schedulers, validation datasets, or trainer wiring unless they are explicitly part of onboarding a new model architecture. Those recipe questions belong to the nemo-automodel-recipe-development skill.
In-scope examples:
- "Add support for a new Hugging Face causal LM architecture."
- "Map MoE router and expert weights from a Hugging Face checkpoint."
- "Register a new model class in NeMo AutoModel."
Out-of-scope examples:
- "Write a finetuning recipe YAML with optimizer and dataset sections."
- "Choose FSDP2, DDP, tensor parallel, or context parallel settings."
- "Configure Slurm, SkyPilot, containers, mounts, or launch dispatch."
Phase 1: Discovery
Before writing code, gather information about the target model.
1.1 Fetch HuggingFace config.json
Download the model's config.json from the HuggingFace Hub (or use AutoConfig.from_pretrained). Key fields to extract:
architectures-- determines the class name and registration key (e.g.,"LlamaForCausalLM","Qwen3MoeForCausalLM","Mistral3ForConditionalGeneration")model_type-- used for custom config registration in_CUSTOM_CONFIG_REGISTRATIONSif HF does not have a built-in config classhidden_size,intermediate_size,num_hidden_layers,num_attention_heads,num_key_value_heads-- sizingvocab_size-- needed for tiny test configstie_word_embeddings-- the saved setting in each supported checkpoint; do not infer it from a bare config constructorhidden_act-- activation function (e.g.,"silu"for SwiGLU)
1.2 Determine model type
| Type | Indicators | Pattern file |
|---|---|---|
| Dense LLM | ForCausalLM in architectures, no expert fields | llm-patterns.md |
| MoE LLM | n_routed_experts, num_local_experts, num_experts_per_tok in config | moe-patterns.md |
| VLM | ForConditionalGeneration in architectures, has vision_config + text_config | vlm-patterns.md |
1.3 Check for existing similar architectures
Look in components/models/ for architectures with similar attention or MLP patterns:
components/models/
llama/ # Standard GQA + SwiGLU (CombinedQKV + CombinedGateUpMLP)
qwen2/ # Same as Llama but with attention bias + QKV bias
baichuan/ # ALiBi attention variant
deepseek_v3/ # MLA attention + MoE (DeepSeek-style grouped experts)
mistral4/ # MLA + MoE + VLM (Pixtral vision)
kimivl/ # DeepSeek-V3 backbone + MoonVit vision
kimi_k25_vl/ # Updated KimiVL with different projector
qwen3_moe/ # Qwen3 with MoE layers
nemotron_v3/ # Hybrid mamba-attention
1.4 Identify custom components
Check whether the model needs:
- Custom attention: GQA (standard), MLA (DeepSeek/Mistral4), sliding window, bidirectional
- Custom RoPE: Standard (Llama), YaRN scaling, NTK-aware, complex-number (DeepSeek)
- Custom normalization: RMSNorm (standard), LayerNorm, different eps values
- Custom MLP: SwiGLU (standard), GeGLU, ReLU-squared, MoE routing
- Custom config class: Needed only if HF
AutoConfigcannot parse the model'sconfig.json(checkauto_mapfield)
1.5 Note dimensions for test config
For unit tests, create a tiny config. Target: ~1M parameters or less.
# Example tiny config for a Llama-like model:
tiny_config = LlamaConfig(
hidden_size=64,
intermediate_size=128,
num_hidden_layers=2,
num_attention_heads=4,
num_key_value_heads=2,
vocab_size=256,
max_position_embeddings=128,
)
Phase 2: Implementation
2.1 Create directory structure
components/models/<name>/
__init__.py
model.py
state_dict_adapter.py
config.py # Only if HF config is insufficient
layers.py # Only for MoE / MLA / other non-standard layers
rope_utils.py # Only for custom RoPE
2.2 Implementation order
Implement files in dependency order:
- config.py (if needed) -- Custom
PretrainedConfigsubclass - rope_utils.py (if needed) -- RoPE implementation
- layers.py (if needed) -- Attention, MLP, decoder block classes
- model.py -- The main
ForCausalLM(orForConditionalGeneration) class - state_dict_adapter.py -- HF weight conversion
- init.py -- Re-export the main model class
See the pattern files for detailed implementation guidance:
- Dense LLM: llm-patterns.md
- MoE: moe-patterns.md
- VLM: vlm-patterns.md
- Capabilities and fp32 precision: capabilities-and-precision.md
2.3 Causal LM weight tying
Every registered model class with a causal lm_head must:
- Declare
tie_word_embeddings_support: TieSupportasBOTH,TIED_ONLY, orUNTIED_ONLY. - Call
reject_unsupported_tie_word_embeddings(type(self), config)at the top of__init__, using the original config before unwrappingtext_configorthinker_config.
Only classes with no causal LM head may be explicitly exempted from the registry test.
Choose the policy from the implementation and the actual supported checkpoint configs, not from a bare config constructor:
BOTH: tied and untied configurations are both supported.TIED_ONLY: only a tied configuration is supported.UNTIED_ONLY: only an untied configuration is supported.
Runtime helpers must treat TIED_ONLY and UNTIED_ONLY as authoritative and
only resolve a per-checkpoint config flag for BOTH. All current BOTH VLMs
honor the outer tie_word_embeddings flag, so do not add a model-specific
resolver until a supported BOTH model actually requires another config path.
For BOTH and TIED_ONLY, always declare _tied_weights_keys and implement
tie_weights() with the actual lm_head and input-embedding FQNs. Do not rely
on inherited Hugging Face tying, and re-tie after any language-model swap.
Add policy-specific tests:
BOTH: tied aliases; untied does not alias.TIED_ONLY: tied aliases; untied is rejected.UNTIED_ONLY: weights stay separate; tied is rejected.
Do not tie architectures with intentionally separate heads, asymmetric vocab sizes, or stages that do not own both tensors.
For from_pretrained, the checkpoint's saved tie_word_embeddings value is
authoritative, even for BOTH. The NeMoAuto* bridge rejects flips in either
direction. A model-owned from_pretrained that bypasses that bridge must call
reject_tie_word_embeddings_flip(checkpoint_config, requested_config, model_class_name).
2.4 MoE state-dict adapter checklist
For MoE models, do not stop at generic loading. The adapter must explicitly map:
- Router weights, including gate bias or correction-bias tensors when the Hugging Face model has them.
- Expert weights, preserving expert index order across local and routed experts.
- Gate/up/down projections, including combined or split projection layouts.
- Shared experts separately from routed experts when the architecture has both.
Add tests that assert expected key mappings and run numerical equivalence with tiny configs before trying full checkpoints.
Do not use these shortcuts:
- Do not validate the adapter only by calling
from_pretrained(). - Do not accept missing or extra expert keys without an explicit mapping reason.
- Do not change dtype, transpose dimensions, or reshape tensors unless the HF and NeMo layouts require it and a test proves the conversion is reversible.
- Do not skip router or shared-expert tests because dense-layer tests pass.
2.5 VLM onboarding checklist
For VLMs, confirm the Hugging Face config has vision_config and text_config
and that architectures points to a conditional-generation class. Start from
the closest VLM pattern file, usually vlm-patterns.md, and
compare existing implementations such as mistral4, kimivl, or
kimi_k25_vl.
The implementation should explicitly cover:
- Text backbone, vision tower, projector, and processor or image preprocessing assumptions.
- Weight mapping for both text and vision modules in
state_dict_adapter.py. - Registration of the
ForConditionalGenerationclass in_transformers/registry.py. - Tiny tests that exercise image-text inputs and verify the adapter round-trip.
2.6 Register in registry
Add the model to MODEL_ARCH_MAPPING in _transformers/registry.py:
# In _transformers/registry.py
MODEL_ARCH_MAPPING = OrderedDict([
# ... existing entries ...
(
"NewModelForCausalLM",
("nemo_automodel.components.models.new_model.model", "NewModelForCausalLM"),
),
])
If the model has a custom config class with auto_map in its config.json, also register in _CUSTOM_CONFIG_REGISTRATIONS:
_CUSTOM_CONFIG_REGISTRATIONS: Dict[str, Tuple[str, str]] = {
# ... existing entries ...
"new_model": ("nemo_automodel.components.models.new_model.configuration", "NewModelConfig"),
}
2.7 Declare capabilities and precision-sensitive params
Every class registered in MODEL_ARCH_MAPPING must declare parallelism
capabilities, either with a static nested ModelCapabilities dataclass or a
variant-aware get_capabilities(cls, config) method. Pick exactly one pattern.
Capabilities should reflect recipe YAMLs that have been validated end to end.
If the model has precision-sensitive parameters such as Mamba A_log /
dt_bias, MoE sigmoid gate bias, attention-sink bias, or per-head scale,
declare _keep_in_fp32_modules_strict so sharding keeps those params in fp32
compute. See capabilities-and-precision.md
for examples, variant dispatch rules, and frozen-submodule dtype guidance.
Phase 3: Onboarding Example Config
This phase is only for adding a minimal example config that proves the newly onboarded architecture can load and run. Use nemo-automodel-recipe-development for general recipe authoring or existing recipe modifications.
3.1 Create example YAML config
Create an example config under examples/llm_finetune/<name>/ (or examples/vlm_finetune/<name>/):
model:
_target_: nemo_automodel.NeMoAutoModelForCausalLM.from_pretrained
pretrained_model_name_or_path: <org>/<model-name>
trainer:
max_steps: 100
gradient_clip_val: 1.0
accumulate_grad_batches: 1
# ... data, optimizer config ...
3.2 Verify model loads
Test that the model loads from a HuggingFace checkpoint:
from nemo_automodel import NeMoAutoModelForCausalLM
model = NeMoAutoModelForCausalLM.from_pretrained("<org>/<model-name>")
3.3 Test with tiny config first
Before using full-size models, verify with a tiny config (1-2 layers, small hidden dim) to catch shape mismatches early.
Phase 4: Tests
Create tests/unit_tests/models/<name>/ and cover the checks below before
loading full checkpoints:
- Forward-shape smoke test with a tiny config.
- State-dict adapter round-trip:
from_hf -> to_hfpreserves mapped names, shapes, dtypes, and values. - Layer equivalence tests for every rewritten attention, MLP, normalization,
RoPE, or MoE layer. Use the model dtype from config, identical seeded weights,
identical inputs, and dtype-appropriate
torch.allclosetolerances. - Short functional test that verifies loss decreases over a few training steps.
Phase 5: Documentation
5.1 Update model coverage page
Edit the appropriate file in docs/model-coverage/:
- LLM/MoE:
docs/model-coverage/llm/index.md - VLM:
docs/model-coverage/vlm/index.md
Add a row with the model name, supported features (TP, PP, FSDP, LoRA, QLoRA), and any limitations.
Phase 6: Parity Testing
After implementation and unit tests are complete, run the full parity-testing workflow to verify that the new model produces numerically equivalent results to the reference HuggingFace implementation.
Run three levels of comparison:
- State-dict round-trip: load a reference HuggingFace checkpoint, convert it into the NeMo AutoModel layout, export it back, and verify that all mapped tensors match the reference names, shapes, dtypes, and values within the expected tolerance.
- Component-level parity: compare rewritten attention, MLP, normalization, RoPE, and MoE components against the HuggingFace implementation with fixed seeds and identical dtype.
- End-to-end forward pass: run the full NeMo AutoModel and HuggingFace model on the same tokenized input and compare logits, hidden states, and loss.
Do not skip this phase. A model that passes unit tests can still diverge from HF due to subtle weight-conversion bugs, backend differences, or RoPE mismatches that only surface in a full parity comparison.
Key Files Reference
| File | Purpose |
|---|---|
_transformers/registry.py | MODEL_ARCH_MAPPING and _CUSTOM_CONFIG_REGISTRATIONS |
components/models/common/__init__.py | Exports CombinedQKVAttentionMixin, CombinedGateUpMLP, BackendConfig, HFCheckpointingMixin, etc. |
components/models/common/combined_projection/combined_qkv.py | CombinedQKVAttentionMixin with setup_qkv_projection() and compute_qkv() |
components/models/common/combined_projection/combined_mlp.py | CombinedGateUpMLP with interleaved gate/up layout |
components/models/common/combined_projection/state_dict_adapter.py | CombinedProjectionStateDictAdapter base class |
components/models/common/hf_checkpointing_mixin.py | HFCheckpointingMixin for save/load |
components/models/common/utils.py | BackendConfig, initialize_rms_norm_module, initialize_linear_module, get_rope_config |
components/moe/config.py | MoEConfig dataclass |
components/moe/fsdp_mixin.py | MoEFSDPSyncMixin for distributed expert handling |
components/moe/layers.py | MoE layer, MLP (dense) for MoE blocks |
components/moe/experts.py | GroupedExperts, GroupedExpertsDeepEP, GroupedExpertsTE |
Checklist
- Fetched and analyzed
config.jsonfrom HuggingFace - Determined model type (dense LLM / MoE / VLM)
- Identified custom components (attention, RoPE, normalization, MLP)
- Created
components/models/<name>/directory - Implemented config.py (if custom config needed)
- Implemented layers.py (if custom layers needed)
- Implemented rope_utils.py (if custom RoPE needed)
- Implemented model.py with
HFCheckpointingMixin - Implemented state_dict_adapter.py
- Implemented init.py with re-export
- Registered in
MODEL_ARCH_MAPPINGin_transformers/registry.py - Registered custom config in
_CUSTOM_CONFIG_REGISTRATIONS(if applicable) - Declared
ModelCapabilitiesnested dataclass (static) ORget_capabilities(cls, config)classmethod (variant dispatch, e.g. ERNIE-4.5 MoE vs dense) — never both, never neither - Declared
TieSupportand called the constructor guard for every class with a causallm_head(or added an explicit no-head exemption) -- see §2.3 - Added explicit
_tied_weights_keysandtie_weights()forBOTH/TIED_ONLY, plus policy-specific alias and rejection tests -- see §2.3 - Guarded any model-owned
from_pretrainedthat bypasses theNeMoAuto*bridge against checkpoint flips -- see §2.3 - Created example YAML config
- Verified model loads via
NeMoAutoModelForCausalLM.from_pretrained() - Created unit tests (forward shape, state_dict round-trip)
- Declared
_keep_in_fp32_modules_strictfor every intrinsically-fp32 param (SSMA_log/dt_bias, MambaDwhen reference-fp32, MoE gate bias, attention-sink bias,scale, …) — see §2.7 - Created layer equivalence tests for every rewritten layer (matching model dtype)
- Created functional tests (training loss decreases)
- Updated docs/model-coverage page
- Ran state-dict round-trip, component parity, and E2E forward-pass parity checks
- Set
ModelClass = <Name>ForCausalLMat module bottom
Frequently asked questions about NeMo AutoModel Onboarding
Similar skills
WinMD API Search
Easily find and explore Windows desktop APIs.
WebMCPify
Transform any web app into an agent-ready platform.
Phoenix Tracing
Instrument LLM applications with OpenInference tracing.
Foundry Hosted Agent CopilotKit
Guidance for developing agentic web apps on Azure.
Power Automate Foundation
Connect AI agents to Power Automate seamlessly.
Power Automate Flow Builder
Efficiently build and deploy Power Automate flows programmatically.
