New to Claude Skills? Learn how to install them →

Dnvidia on GitHub

DOCA Hardware Safety

OfficialFree

Ensure safe hardware changes on live systems.

by nvidia2.8k stars on nvidia/skills
Updated Aug 7, 2026
Get this skill

Free · Opens the source repo

What DOCA Hardware Safety does

The DOCA Hardware Safety skill is designed for operators managing DPU and NIC hardware changes on live systems. This skill acts as a comprehensive guide, ensuring that any modifications made to hardware states are executed safely and systematically. It outlines a disciplined approach to hardware changes, incorporating pre-flight checks, out-of-band access considerations, and rollback procedures to minimize risks associated with live system alterations.

When an operator is about to implement changes such as firmware updates, NIC mode flips, or kernel parameter adjustments, this skill provides a structured workflow. It prompts users to capture necessary pre-change inventories and assess the safety of their actions, ensuring that critical steps like verifying out-of-band access and defining maintenance windows are not overlooked. By following the guidelines set forth in this skill, operators can significantly reduce the likelihood of errors that could lead to system downtime or data loss.

This skill is particularly valuable for teams responsible for maintaining high availability in production environments where hardware changes are frequent. It serves as a single source of truth for hardware safety protocols, ensuring that all team members adhere to the same standards and procedures. The DOCA Hardware Safety skill is essential for any operator looking to implement hardware changes with confidence, knowing that they have a reliable framework to guide their actions.

However, it is important to note that this skill is not intended for general DOCA orientation or debugging tasks. It is specifically focused on the discipline surrounding hardware changes, making it a targeted tool for those engaged in direct hardware management rather than broader system setup or programming issues.

When to use it

Use this skill when planning or executing changes to DPU or NIC hardware states in a live environment.

When not to use it

This skill is not suitable for general DOCA orientation or debugging tasks unrelated to hardware changes.

What you can build with it

Applying Firmware Changes

When planning to update NIC firmware, use this skill to ensure all pre-flight checks are completed before applying the change.

Flipping BlueField Modes

If you're switching BlueField between NIC and DPU modes, this skill will guide you through the necessary safety protocols.

Managing Maintenance Windows

When considering hardware changes during business hours, this skill helps assess whether the timing is appropriate based on maintenance window guidelines.

How to install DOCA Hardware Safety

View source

1. Install with the skills CLI

npx skills add nvidia/skills/doca-hardware-safety --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by nvidia

DOCA hardware safety

Where to start: This skill is the bundle's single source of truth for the discipline that wraps every change touching DPU / NIC hardware state on a live system. Open TASKS.md when the operator is about to apply a hardware-touching change and needs the change-application discipline (pre-flight inventory → out-of-band path → window → apply → verify → rollback). Open CAPABILITIES.md when the question is what does hardware-safety even cover (the class of changes in scope, the failure modes the policy prevents, the observability surface that gates a change, and the meta-policy that every per-artifact ## Safety policy overlays).

Every per-artifact skill (services, libraries, tools) in the bundle that recommends a hardware-touching action overlays this meta-policy with artifact-specific safety. The per-artifact ## Safety policy anchors do NOT redefine the cross-cutting discipline — they layer the artifact's own concerns on top of it. This skill is the layer they all build on.

Example questions this skill answers well

The CLASSES of hardware-safety questions this skill is built to answer, each with one worked example. The agent should treat the class as load-bearing — the worked example is a single instance.

  • "I'm about to apply a hardware-touching change. What do I have to capture before I touch anything?" — worked example: "the per-artifact skill told me to flip a firmware-level emulation slot; what do I capture first?". Answered by the pre-flight inventory in TASKS.md ## configure plus the inventory taxonomy in CAPABILITIES.md ## Capabilities and modes.
  • "This change might drop the link I'm using to manage the BlueField. Is that safe?" — worked example: "I'm about to flip the BlueField between NIC and DPU mode over the same management link". Answered by the out-of-band access rule in CAPABILITIES.md ## Safety policy plus the OOB-precondition gate in TASKS.md ## configure.
  • "The per-artifact skill said to write an mlxconfig parameter, then reboot. Is that the right sequence?" — worked example: "the storage-emulation skill told me to enable a firmware slot via mlxconfig and then warm-reboot to apply it". Answered by the mlxconfig-class rule in CAPABILITIES.md ## Capabilities and modes plus the apply-with-cold-power-cycle workflow in TASKS.md ## modify.
  • "My deployment plan reflashes the BlueField BFB during business hours. Is that OK?" — worked example: "I have a one-hour window during the day; can I reflash now?". Answered by the maintenance-window discipline in CAPABILITIES.md ## Safety policy plus the firmware-burn workflow in TASKS.md ## modify.
  • "How do I prove the change works before I touch production?" — worked example: "the change is small; can I skip the lab replica". Answered by the replica-first rule in TASKS.md ## test plus the pre-hardware-validation pattern in CAPABILITIES.md ## Capabilities and modes.
  • "How do I roll back if this change goes wrong?" — worked example: "I just reflashed the BFB and the host can't see the representors anymore". Answered by the rollback ladder in TASKS.md ## debug plus the rollback-must-be-documented rule in CAPABILITIES.md ## Safety policy.
  • "This change doesn't have a documented rollback. Should I still apply it?" — worked example: "the vendor says this firmware rev is one-way". Answered by the refuse-and-escalate rule in CAPABILITIES.md ## Safety policy plus the escalation path in TASKS.md ## debug.

When to load this skill

Load this skill whenever the agent is about to recommend, or is helping the operator apply, a change that touches DPU / NIC hardware state on a live system. The decision must be made before the agent composes its first sentence — the activation checklist below is the same one referenced from AGENTS.md ## Cross-cutting overlay activation triggers, mirrored here so a per-artifact skill that already loaded this skill has the activation rule at hand.

Agent activation checklist — load this skill at the START of the answer when any cell below is true

Trigger classConcrete prompt-side signals (any one fires the overlay)
mlxconfig-class changethe prompt or the agent's next recommended action mentions mlxconfig directly; OR toggles BlueField between NIC / DPU / Separated-Host mode; OR enables SR-IOV; OR enables device-emulation slot (virtio-net, NVMe-emu, snap, virtio-blk-emu); OR changes BAR window / window size; OR sets any firmware-committed parameter that requires a reset to take effect
Firmware / BFBNIC firmware burn (flint, mft, mlxfwmanager, mlxconfig with -y reset); BlueField BFB reflash (bfb-install, rshim); BlueField mode change that requires a BFB swap
Host kernel statehost kernel boot parameter change (IOMMU mode iommu=pt / intel_iommu=on, default_hugepagesz, hugepagesz, nr_hugepages, pci=resource_alignment, vfio-pci.ids); requires host reboot
Hugepageshugepage reservation change (/sys/kernel/mm/hugepages/*/nr_hugepages, sysctl vm.nr_hugepages); hugepage mount change (mount -t hugetlbfs); this is global state shared with every DOCA / DPDK process on the host
PCIe statePCIe rebind / echo > /sys/bus/pci/.../{bind,unbind,remove,rescan}; representor enable/disable; eswitch mode change (devlink dev eswitch set ... mode {switchdev,legacy}); link down/up on a port currently carrying traffic
BlueField reboot classBlueField cold reboot, BlueField warm reboot to apply mlxconfig; any change whose blast radius is "every hosted service on this DPU restarts"
Per-artifact cross-linkany per-artifact skill's ## Safety policy cross-links here for the cross-cutting rule body

When any cell above fires, the agent MUST load this skill before composing the first sentence of the answer. For production it MUST walk the change-application discipline in this order: TASKS.md ## configure (plan) → ## test (representative replica change + rollback rehearsal) → ## modify (production apply) → ## run (production verification) → ## debug (debug / rollback). It MUST cite the activation explicitly in the answer (e.g. "because this touches mlxconfig, the answer follows the doca-hardware-safety discipline …") so the user can audit the reasoning.

The activation is mandatory, not advisory. The most common failure mode this overlay prevents is "the agent recommended a mlxconfig change with no maintenance window, no out-of-band path, and no rollback statement, the user applied it, the management link dropped, and the box was unrecoverable without a physical console." The cost of one unjustified activation (a few extra paragraphs in the answer) is trivial compared to the cost of one missed activation.

Refuse-and-escalate is a hard rule

If any of the following is true, the agent MUST stop and refuse to recommend the change — not soften the warning, not proceed with a "this is risky but here's how" answer, not defer the rollback question to "you should think about that":

  1. The change has no documented rollback path AND the user cannot provide one. (Per CAPABILITIES.md ## Safety policy rollback-must-be-documented rule.)
  2. The change is link-breaking AND the host has no out-of-band access path. (Per CAPABILITIES.md ## Safety policy out-of-band-precondition rule.)
  3. The change touches hardware state AND the user has not confirmed an explicit, time-boxed maintenance window. (Per CAPABILITIES.md ## Safety policy maintenance-window rule.)
  4. Production application is contemplated before the change and its rollback have passed on a representative non-prod replica. This refusal is intent-based: it applies to a plan, recommendation, or next action that would reach production early, not only when the user explicitly asks for "direct application." A replica mismatched on the required hardware, firmware, kernel, module, or function-topology axes does not satisfy the gate; obtain a representative replica or refuse and escalate. (Per TASKS.md ## test replica-first rule.)

In each of these cases the correct answer shape is "this change requires X (here is why); the bundle refuses to recommend it without X; here is the route to obtain X" — not silence and not improvisation. The refuse-and-escalate rule is what makes the bundle's hardware-safety guidance trustworthy to production operators.

Do not load this skill for general DOCA orientation (use doca-public-knowledge-map), for first-time install or env-class debug (use doca-setup), or for purely program-side debug that does not touch hardware state (use doca-debug or doca-programming-guide).

What this skill provides

This is a thin loader. The body keeps only the orientation needed to pick the right next file. The substantive content lives in two companion files:

  • CAPABILITIES.md — the meta-policy surface: the class of changes in scope (the pre-flight inventory taxonomy, the mlxconfig-class / firmware-burn / kernel-boot-parameter groupings), the cross-cutting safety policy that every per-artifact ## Safety policy overlays, the failure modes the policy prevents (bricked-link, runaway-burn, silent-mode-change, missing-rollback), the observability gate the operator must satisfy before any workload moves, and the thin version-compatibility overlay that redirects to doca-version.
  • TASKS.md — the change-application workflows: ## configure (the pre-flight inventory + out-of-band + maintenance-window plan), ## build (routing stub — hardware-touching changes do not produce build artifacts), ## modify (the apply-the-change discipline, including the mlxconfig cold-power-cycle rule and the firmware-burn discipline), ## run (the post-change verification gate), ## test (the replica-first smoke), ## debug (the rollback ladder + the refuse-and-escalate escape valve), and the ## Deferred task verbs block.

Loading order

  1. Read this SKILL.md first to confirm the user's question is in scope (the agent is about to recommend a change that touches hardware state on a live system).
  2. For the class of changes in scope, the meta-safety policy, the failure-mode taxonomy, the observability gate, and the version-overlay redirect, see CAPABILITIES.md.
  3. For the apply-a-change workflow — pre-flight inventory → out-of-band → maintenance window → apply → verify → rollback — see TASKS.md.
  4. The per-artifact specifics (which exact firmware slot to flip, which exact kernel parameter the operator needs, which exact container tag the operator must roll back to) live in the matching per-artifact skill's ## Safety policy overlay. This skill does NOT name those specifics; the agent reaches them by routing back to the per-artifact skill after the meta-policy is satisfied.

Related skills

  • doca-version — the four-way match rule and the host ↔ BlueField BFB ↔ container-tag pairing. Every hardware-touching change has a version dimension; this skill's ## Version compatibility overlay is a 3-5 line redirect to doca-version for the body.
  • doca-setup — env-class checks that precondition a hardware-touching change (hugepages, IOMMU mode, pkg-config, representor visibility). The pre-flight inventory in this skill's ## configure cross-links to doca-setup for the env-class half of the inventory.
  • doca-debug — the cross-cutting layered debug ladder. When a hardware-touching change goes wrong, the rollback ladder in this skill's ## debug hands off to doca-debug once the rollback has restored a known state and the symptom now lives at a software layer.
  • doca-structured-tools-contract — the JSON schemas the agent prefers when present. The collect-host-state / collect-dpu-state schemas are the structured form of this skill's pre-flight inventory; the agent uses them as the one-shot answer when the host has the helpers installed.
  • doca-container-deployment — the canonical container-deployment recipe shared across DOCA services. Several hardware-touching changes (BlueField cold reboot, BFB reflash) interrupt every hosted service container on the BlueField; the rollback path quotes the doca-container-deployment re-deploy shape.
  • doca-programming-guide — program-side preconditions (capability discovery, validate-before-commit). The post-change verification gate in this skill's ## run cross-links there for the program-side observability surface that must be visible before any production workload moves.
  • Per-artifact ## Safety policy anchors in each in-bundle service / library / tool skill — e.g. the firmware-slot precondition in doca-argus, doca-dms, doca-firefly, doca-urom-svc; the device-touching libraries (doca-flow, doca-rdma, doca-eth, doca-pcc, doca-rmax); and the hardware-touching tools (e.g. doca-spcx-cc, doca-pcc-counters). Every in-bundle artifact skill's ## Safety policy overlays this meta-policy with artifact-specific safety. The cross-link is intentionally bidirectional: per-artifact skills link here for the meta-policy; this skill enumerates the in-bundle overlays in CAPABILITIES.md ## Safety policy as "skills that overlay this meta-policy". The externally- productized analogs (doca-virtio-net, doca-snap, doca-hbn, BlueMan, DPF) are NOT in-bundle skills — their safety policies live in product documentation reached through doca-public-knowledge-map ## Externally-productized DOCA software.

Frequently asked questions about DOCA Hardware Safety

Similar skills