
NVIDIA Docker Skill
OfficialFreeRun NVIDIA GPU workloads in Docker containers seamlessly.
Free · Opens the source repo
What NVIDIA Docker Skill does
The NVIDIA Docker Skill provides essential conventions and commands for running GPU container workloads using Docker. It is specifically designed for developers and data scientists who need to leverage NVIDIA's GPU capabilities within their Docker environments. This skill covers critical aspects of setting up and executing Docker commands that comply with NVIDIA's requirements, ensuring that users can effectively manage their GPU resources while working with containerized applications.
This skill is particularly useful when working with NVIDIA GPU containers hosted on the NVIDIA GPU Cloud (NGC). It includes detailed instructions on how to authenticate with the NGC, utilize the --gpus flag, and manage environment variables for seamless integration with other skills that require Docker commands. By following the conventions outlined in this skill, users can avoid common pitfalls and errors associated with running GPU workloads in Docker, such as container name collisions and improper resource allocation.
The skill also provides practical examples of Docker commands tailored for NVIDIA workloads, including best practices for data mounting, container inspection, and managing shared memory. These examples help users optimize their workflows and ensure that their applications run efficiently on GPU hardware. Whether you are deploying machine learning models or conducting data analysis, this skill equips you with the necessary tools to harness the power of NVIDIA GPUs within Docker containers.
In summary, the NVIDIA Docker Skill is an invaluable resource for anyone looking to run GPU-accelerated applications in a Docker environment. It simplifies the complexities of GPU management and ensures that users can focus on their core tasks without being bogged down by technical challenges.
When to use it
Use this skill when you need to run Docker containers that require NVIDIA GPU support, especially in machine learning or data processing tasks.
When not to use it
This skill is not suitable for non-GPU workloads or if your Docker environment does not have the necessary NVIDIA drivers and toolkit installed.
What you can build with it
Running a Machine Learning Model
Utilize this skill to set up and run a Docker container that executes a machine learning model leveraging NVIDIA GPUs.
Data Processing with GPU Acceleration
Use the skill to configure Docker for data processing tasks that require high-performance GPU resources.
Container Management for Multi-Step Workflows
Implement the detached and exec pattern to manage complex workflows across multiple Docker containers efficiently.
How to install NVIDIA Docker Skill
View source1. Install with the skills CLI
npx skills add nvidia/skills/tao-run-on-docker --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by nvidiaDocker for NVIDIA GPU Workloads
This skill documents the generic Docker conventions that GPU container workloads rely on. Model and data skills specify what image and what command to run; this skill covers how to run docker in a way that satisfies GPU + NVIDIA container requirements.
Sources: official Docker CLI reference (https://docs.docker.com/reference/cli/docker/) and NVIDIA Container Toolkit docs.
Prerequisites
- Host GPU runtime — NVIDIA driver branch 580, CUDA Toolkit 13.0, and NVIDIA Container Toolkit 1.19.0. Check with the
tao-setup-nvidia-gpu-hostskill before any GPU workflow starts. - Docker —
docker --versionmust return ≥ 20.10. Install: https://docs.docker.com/engine/install/. - NGC API key for
nvcr.io/*pulls. Get from https://ngc.nvidia.com/.
TAO_SKILL_BANK_ROOT="${TAO_SKILL_BANK_ROOT:-$PWD}"
SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT}/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"
bash "$SETUP_SCRIPT" --backend docker --check-only || {
echo "MISSING: TAO GPU host runtime is not ready."
echo "After user approval, run (append --yes for non-interactive agent runs):"
echo " bash \"$SETUP_SCRIPT\" --backend docker --install"
exit 1
}
docker --version
docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
[ -n "$NGC_KEY" ] || echo "NGC_KEY unset — cannot pull nvcr.io images"
NGC authentication
echo "$NGC_KEY" | docker login nvcr.io -u '$oauthtoken' --password-stdin
Persists in ~/.docker/config.json across reboots. Re-run on unauthorized errors.
docker run — canonical flags
docker run \
--gpus all \ # all GPUs (requires nvidia-container-toolkit)
--rm \ # delete container after exit (image is preserved)
--shm-size=8g \ # shared mem for torchrun / DataLoader
-v /host/data:/data \ # bind-mount input
-v /host/results:/results \ # bind-mount output
-e HF_TOKEN -e NGC_KEY \ # env-var passthrough (values from parent shell)
<image> \
<command>
Notes:
--gpus '"device=0,1"'— specific GPUs (double-quote-escaped). Without nvidia-container-toolkit:could not select device driver "" with capabilities: [[gpu]].--rm— clean up the container at exit; omit when you wantdocker logsafter exit.--shm-size=8g— torchrun + PyTorch DataLoaders exhaust the default 64 MB/dev/shmotherwise; size it for multi-GPU training and raise (e.g.16g) if you still hitBus error.-v host:container— bind mount; the command references container paths only.-e VAR— passthrough from parent shell (no value needed if already set). Use this form for secrets.
Container name collision
docker run --name X fails if a container named X already exists. Defensive pattern before reusing a name:
docker stop my-worker 2>/dev/null; docker rm my-worker 2>/dev/null
docker run --name my-worker ...
Detached + exec pattern
For multi-step workflows on the same container (download → run → post-process), avoid restart cost:
docker run -d --name <worker> \
--gpus all --shm-size=8g \
-v <mounts...> -e <envs...> \
--entrypoint sh \
<image> -c "tail -f /dev/null"
docker exec <worker> <step_1>
docker exec <worker> <step_2>
docker stop <worker> && docker rm <worker>
Pull-if-missing idiom
docker image inspect <image> >/dev/null 2>&1 || docker pull <image>
Labels for discovery
Tag containers for filtered listing later:
docker run --label tao-toolkit ...
docker ps --filter 'label=tao-toolkit'
Mount patterns
The container expects its data at conventional paths defined by the image (often /data, /results, /workspace/checkpoints). The host side is arbitrary. The command inside docker run references container paths only.
Env-var conventions
Common passthrough vars for TAO-style workloads (the calling skill declares which it needs):
NGC_KEY—nvcr.iopulls; some runtimes also read at runtimeHF_TOKEN— gated HuggingFace model downloadsAWS_ACCESS_KEY_ID,AWS_SECRET_ACCESS_KEY,AWS_ENDPOINT_URL— S3 I/O inside the containerWANDB_API_KEY— optional W&B logging
Use -e VAR (no =value) when the var is in the parent shell. Avoid placing secrets on the command line.
Alternative GPU selection: -e NVIDIA_VISIBLE_DEVICES=0,1 (or all) and -e NVIDIA_DRIVER_CAPABILITIES=all instead of --gpus. The --gpus flag is preferred on standard x86 hosts; the env-var form is older and is what runtime=nvidia (Tegra/Jetson) requires.
Container inspection
docker ps # running containers only
docker ps -a # all containers, including exited
docker ps --filter status=running --format '{{.Names}} {{.Image}}'
docker logs <name_or_id> # stdout/stderr
docker logs -f <name_or_id> # follow (tail -f equivalent)
docker logs --tail 100 <name_or_id> # last N lines
docker inspect <name_or_id> # full config, mounts, env, network, state (JSON)
docker inspect --format '{{.State.Status}}' <name_or_id>
docker stats # live CPU/mem/network/block I/O
docker stats --no-stream # one snapshot, non-interactive
docker inspect is the canonical source of truth for a container's mounts, env, cmd, network, and exit code. Use it to debug why a container isn't behaving as expected.
Image management
docker pull <image>
docker image ls
docker system df # disk usage
docker system prune -a --volumes # reclaim space — destructive, removes unused images + volumes
Pull once per host; docker run reuses cached image. NVIDIA images are typically 5-40GB.
Split-disk data-root relocation
Some cloud GPU providers ship with a small root volume + larger ephemeral. Docker writes to /var/lib/docker on root by default — large images fill it. Check:
df -h / # root volume size/free
lsblk # all block devices and mount points
If / is smaller than your total image footprint and there's a larger disk mounted elsewhere, relocate before pulling images:
sudo systemctl stop docker
sudo mkdir -p <large_volume_path>/docker
sudo rsync -aP /var/lib/docker/ <large_volume_path>/docker/
sudo mv /var/lib/docker /var/lib/docker.old
sudo tee /etc/docker/daemon.json <<'EOF'
{ "data-root": "<large_volume_path>/docker" }
EOF
sudo systemctl start docker
docker info | grep 'Docker Root Dir'
sudo rm -rf /var/lib/docker.old
Networks (multi-container patterns)
For microservice containers that talk to each other by name, create a docker network and attach containers:
docker network create tao-net
docker run --network tao-net --name api ...
docker run --network tao-net --name worker ... # can resolve `api` by name
Most TAO training workloads don't need this — single container per job.
Common error modes
could not select device driver "" with capabilities: [[gpu]] — NVIDIA Container Toolkit missing or Docker is not configured for the NVIDIA runtime. Run tao-setup-nvidia-gpu-host with --backend docker --install after user approval (append --yes for a non-interactive agent run), then restart Docker.
unauthorized: authentication required on docker pull — NGC key invalid/missing. Re-run docker login nvcr.io.
no space left on device — root volume full. docker system df to inspect; relocate data-root (above) or docker system prune -a --volumes.
Bus error / DataLoader worker exited unexpectedly — /dev/shm too small. Increase shared memory with --shm-size (e.g. --shm-size=16g).
permission denied on bind-mounted paths — container UID ≠ host UID. Either -u $(id -u):$(id -g), or pre-create host files owned by the host user, or chmod 777 (dev only).
Error: No such container: <name> after docker run -d — container crashed on startup. docker ps -a shows exited; docker logs <name> for cause. Drop --rm while debugging.
Scope boundary
This skill covers the how of running docker on a GPU host. Platform-specific layering (how to get onto the host, dispatch via a CLI wrapper) lives in:
tao-skill-bank:tao-run-on-brev— running docker viabrev execon a Brev instancetao-skill-bank:tao-run-platform— optional Python layer wrapping docker invocations with Job handles, state persistence, and S3 I/O
Model and data skills specify what image and command; they defer to this skill for the how.
Frequently asked questions about NVIDIA Docker Skill
Similar skills
Turborepo
Optimized build system for JavaScript/TypeScript monorepos.
Azure Pipelines Validation
Streamline your Azure DevOps pipeline changes locally.
Azure Developer CLI
Streamline your Azure project workflows with best practices.
Azure Container Registry CLI
Manage Azure Container Registry resources with ease.
Aspire
Build and orchestrate polyglot distributed applications seamlessly.
Vercel CLI
Manage and deploy Vercel projects from the command line.
