New to Claude Skills? Learn how to install them →

Swshobson on GitHub

Spark Training Gotchas

Free

Diagnose and preempt common ML training issues on DGX Spark.

Get this skill

Free · Opens the source repo

What Spark Training Gotchas does

The Spark Training Gotchas skill is designed to assist developers and data scientists who are working with NVIDIA DGX Spark, particularly on the GB10 chip. This skill provides a structured approach to diagnosing and resolving ten known failure modes that can occur during machine learning training runs. These issues can manifest as import errors, out-of-memory (OOM) errors, performance degradation, and more, which can be critical when managing long training jobs. By utilizing this skill, users can ensure their training runs are set up for success before they commence, potentially saving significant time and resources.

The skill includes a comprehensive reference guide that categorizes these failure modes, labeled G1 through G10, each corresponding to specific symptoms and recommended fixes. For example, G1 addresses CUDA ABI mismatches that can lead to segfaults, while G3 tackles OOM errors that may occur despite available memory. Each gotcha is linked to a detailed check in the accompanying reference documentation, allowing users to quickly identify and rectify issues related to their training environments.

This skill is particularly valuable for those who frequently utilize DGX Spark for deep learning tasks, as it provides proactive checks and solutions tailored to the unique challenges of this hardware. It is an essential tool for ensuring that training jobs are not only initiated correctly but also run efficiently, minimizing the risk of costly interruptions.

In summary, the Spark Training Gotchas skill equips users with the knowledge and tools necessary to navigate the complexities of training on DGX Spark, making it an indispensable resource for anyone involved in high-performance machine learning workflows.

When to use it

Use this skill when preparing for a long training run on DGX Spark or when encountering issues like OOM errors or performance drops during training.

When not to use it

This skill is not suitable for environments outside of NVIDIA DGX Spark or for troubleshooting issues unrelated to the ten documented failure modes.

What you can build with it

Pre-Training Checks

Before starting a multi-hour training job, use this skill to check for potential issues that could cause failures.

Diagnosing OOM Errors

If a training run fails with an OOM error despite available memory, this skill can guide you through the troubleshooting process.

Performance Optimization

When experiencing throughput drops mid-training, refer to the skill for insights on bandwidth and thermal issues.

How to install Spark Training Gotchas

View source

1. Install with the skills CLI

npx skills add wshobson/agents/spark-training-gotchas --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by wshobson

Spark Training Gotchas

DGX Spark's GB10 chip (Grace Blackwell, SM121, 128GB unified memory, aarch64) has ten recurring failure modes across launch, memory, thermals, bandwidth, and precision. Each is named G1–G10 so it can be checked by number — the numbering is load-bearing for tooling that runs these checks. Read this before a long run, not after hour six.

When to Use This Skill

  • A training run fails to start, with an import error or a segfault that doesn't point at the real cause.
  • A run OOMs while nvidia-smi still shows headroom.
  • Throughput degrades partway through a run that started fine.
  • Before any multi-hour or multi-epoch job on GB10.
  • Wiring two Sparks together, before picking a parallelism strategy.
  • Choosing between FP8 and NVFP4 for a Spark-hosted run.

Common Issues Quick Reference

#SymptomFix
G1undefined symbol / segfaultcu130 wheel or container
G2flash-attn wrong backend usedskip pip build; monkeypatch on NGC
G3OOM despite headroomdrop page cache
G4throughput drop / rebootexpect ~100W sustained cap
G5memory-bound step slowbudget 180–192 GB/s
G6cache evicted mid-runone GPU server at a time
G7NVFP4 slower than FP8stay FP8 unless sm_121a
G8playbook fails outrightcheck upstream issues
G9env breaks after installuse a container
G102-Spark TP hangsDDP/FSDP only, never TP

The Ten Gotchas

G1: CUDA 12/13 ABI Mismatch

  • SYMPTOM: ImportError: undefined symbol naming a CUDA function, or a segfault on the first .cuda() call.
  • CAUSE: most PyPI wheels link libcudart.so.12; Spark ships CUDA 13. pip never checks CUDA ABI, so it surfaces only at import or first kernel launch.
  • CHECK: references/gotcha-checks.md G1 — the wheel's CUDA build tag.
  • FIX: reinstall from download.pytorch.org/whl/cu130 or use a matched container.

G2: flash-attn — Skip the pip Build, Watch Unsloth's Auto-Detect

  • SYMPTOM: pip install flash-attn still fails/hangs. Unsloth may also silently train flash-attn over an explicitly requested SDPA.
  • CAUSE: no aarch64/sm_121 wheel for bare pip — but NGC containers ship a working SM121 flash-attn, and Unsloth auto-prefers it, dropping attn_implementation="sdpa".
  • CHECK: references/gotcha-checks.md G2 — is flash-attn already present and working.
  • FIX: bare pip — skip flash-attn, use SDPA (unchanged). On NGC — the only reliable override is the monkeypatch in references/gotcha-checks.md G2.

G3: UMA OOM Below 128GB

  • SYMPTOM: OOM during model load/training while nvidia-smi still reports free memory under the 128GB cap — or, on some setups, [N/A] outright instead of a number.
  • CAUSE: mmap and the CUDA allocator double-count pages during safetensors load; QLoRA can OOM earlier than bf16 since dequantization adds transient allocs.
  • CHECK: references/gotcha-checks.md G3 — read free -g and /proc/meminfo, not nvidia-smi.
  • FIX: drop the page cache with sync; echo 3 > /proc/sys/vm/drop_caches — needs root, a between-run reset, not a mid-training step.

G4: Thermal Throttling

  • SYMPTOM: throughput drops partway through a multi-hour run, or the box spontaneously reboots under sustained load.
  • CAUSE: sustained power draw caps around 100W versus the 240W rated figure; long runs push into that ceiling and throttle or, sometimes, reboot.
  • CHECK: references/gotcha-checks.md G4 — sample nvidia-smi --query-gpu=temperature.gpu,power.draw.
  • FIX: if power plateaus under 240W while temperature climbs, treat throttling as the cause; improve cooling or cap run length.

G5: Bandwidth Ceiling

  • SYMPTOM: memory-bound workloads, decode-heavy RL loops especially, plateau well below expected throughput.
  • CAUSE: 273 GB/s is a spec ceiling, not sustained; measured bandwidth runs 180–192 GB/s.
  • CHECK: references/gotcha-checks.md G5 — observed step time vs. the measured range, not spec.
  • FIX: budget throughput from 180–192 GB/s; revise a plan built on the 273 GB/s figure.

G6: Global UMA Resource Contention

  • SYMPTOM: a process's KV cache/weights get evicted mid-run silently, no OOM in its own logs.
  • CAUSE: unified memory is one global pool; an uncapped or near-capacity process competes with anything else and can evict it. A small, bounded workload doesn't — a <4GB LoRA coexists fine alongside vLLM capped at gpu-memory-utilization<=0.5.
  • CHECK: references/gotcha-checks.md G6 — other GPU-resident processes and whether capped.
  • FIX: the one-heavy-job rule applies to uncapped or near-capacity workloads — cap or stop unrelated servers first. A small, capped workload need not stop.

G7: NVFP4 Slower Than FP8 on SM121

  • SYMPTOM: switching an inference workload from FP8 to NVFP4 on Spark makes it slower, not faster.
  • CAUSE: SM121 lacks cvt.e2m1x2 unless kernels target sm_121a; NVFP4 runs ~32% slower without it.
  • CHECK: references/gotcha-checks.md G7 — capability reports (12, 1); does the build target sm_121a?
  • FIX: stay on FP8 unless the build targets sm_121a.

G8: Stale Official Playbooks

  • SYMPTOM: following an official DGX Spark playbook still fails, with no local misconfiguration explaining it.
  • CAUSE: official playbooks have shipped broken before; the stack moves faster than the docs.
  • CHECK: references/gotcha-checks.md G8 — the playbook repo's recent issues.
  • FIX: check github.com/NVIDIA/dgx-spark-playbooks issues before trusting a recipe for an expensive run.

G9: Container-First, Not Bare Pip

  • SYMPTOM: a bare-pip environment that worked yesterday breaks after an unrelated pip install, or two "identical" environments behave differently.
  • CAUSE: bare pip lets Triton, xformers, and transformers drift independently; nothing pins them to GB10's SM121 target.
  • CHECK: references/gotcha-checks.md G9 — container or bare pip?
  • FIX: prefer an NGC container (see spark-environment-setup for tag guidance) or Unsloth's container. If bare pip is unavoidable, follow the NVIDIA install order, including --no-deps on Unsloth.

G10: Dual-Spark Is DDP/FSDP Only

  • SYMPTOM: a tensor-parallel launch across two Sparks hangs, runs far slower than single-Spark, or errors out.
  • CAUSE: ConnectX-7 is fast enough for gradient/parameter sync (DDP, FSDP) but too thin for TP's fine-grained traffic.
  • CHECK: references/gotcha-checks.md G10 — the configured parallelism strategy.
  • FIX: on a two-Spark setup, choose DDP or FSDP, never tensor parallelism — TP is single-node only here.

Fast Triage

The cheapest checks to run before anything else:

python3 -c "import torch; print(torch.version.cuda)"  # expect 13.x (G1); NGC builds have no +cu130 tag — that's not a failure
import torch; print(torch.cuda.get_device_capability())  # expect (12, 1) (G7)
{ [ -f /.dockerenv -o -f /run/.containerenv ] || grep -qE 'docker|containerd' /proc/1/cgroup; } 2>/dev/null && echo container || echo unknown  # G9

assets/preflight.sh runs G1, G3, G4, G7, G9 and produces one output line per gotcha in a fixed format: G-number first, then PASS/FAIL/WARN where automatable, SKIP when unavailable, or INFO: for a raw reading (G3, G4). Full commands: references/gotcha-checks.md. See also spark-environment-setup for the environment assumed working.

Frequently asked questions about Spark Training Gotchas

Similar skills