nvidia-skill-finder
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. Trigger on NVIDIA products, hardware, software,…
Use this skill for launching, supervising, debugging, OR platform lifecycle on a BlueField — BFB install, RShim/TMFIFO, host PF rebind, post-BFB recovery — taking a DOCA-linked binary to a healthy run directly on hardware (host x86 + BlueField NIC over PCIe, or BlueField Arm
$ npx -y skills add NVIDIA/skills --skill doca-bare-metal-deployment --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/doca-bare-metal-deploymentContext preview
The summary Claude sees to decide when to auto-load this skill.
Use this skill for launching, supervising, debugging, OR platform lifecycle on a BlueField — BFB install, RShim/TMFIFO, host PF rebind, post-BFB recovery — taking a DOCA-linked binary to a healthy run directly on hardware (host x86 + BlueField NIC over PCIe, or BlueField Arm
license: Apache-2.0 name: doca-bare-metal-deployment description: > Use this skill for launching, supervising, debugging, OR platform lifecycle on a BlueField — BFB install, RShim/TMFIFO, host PF rebind, post-BFB recovery — taking a DOCA-linked binary to a healthy run directly on hardware (host x86 + BlueField NIC over PCIe, or BlueField Arm bare-metal). No container, no kubelet. Covers launch mode (direct, tmux, systemd), PCI/NUMA/ CPU/IRQ binding, co-tenant isolation (cgroup-v2/netns/numactl), a seven-layer error taxonomy, and a six-state BlueField lifecycle classifier. Trigger even when user does not say "bare-metal" — implicit phrasings include "binary exits 1 right after launch", "systemd keeps restarting it", "no matching device on the BF", "bfb-install exited 0 but DPU is dead", "ping 192.168.100.2 works but ssh fails", "host PFs aren't showing netdevs". Destructive firmware burn / mlxconfig set requires explicit confirmation via doca-hardware-safety; containers, library APIs, env prep, and build use other skills. metadata: kind: library compatibility: > No DOCA install required to read this skill (it is an overlay loaded against any DOCA artifact skill); the validation steps within this skill require a live DOCA install at /opt/mellanox/doca on a host or BlueField with a built DOCA-linked binary.
**Where to start:** This skill is the bundle's home for *operating* a DOCA-linked application binary **directly on hardware** — no container, no kubelet, no static-pod manifest. It is the parallel of [`doca-container-deployment`](../doca-container-deployment/SKILL.md) for the non-container path. If the user has a DOCA-linked binary they built (per the canonical workflow in [`doca-programming-guide`](../doca-programming-guide/SKILL.md)) and they want to know *how to actually run it on the host or on the BlueField Arm cores correctly*, open [`TASKS.md`](TASKS.md) and start at [`## configure`](TASKS.md#configure). If the question is *what shape does the bare-metal runtime even have and what is the deployment contract*, start at [`CAPABILITIES.md`](CAPABILITIES.md). If the user is not yet sure whether their target system shape is the container path or the bare-metal path, route the recognition step to [`doca-setup`](../doca-setup/SKILL.md) first; only return here once *bare-metal* is the confirmed shape.
This skill serves **external DOCA developers and operators who have a DOCA-linked application binary they built and want to run it directly on hardware** — i.e., people who already have:
[`doca-programming-guide ## build`](../doca-programming-guide/TASKS.md#build),
x86** path — DOCA host install on the host talks to the BlueField NIC over PCIe), OR a BlueField with a console or SSH to the Arm side (the **BlueField Arm bare-metal** path — DOCA installed on the DPU Arm cores; the binary runs there directly), and
inside a kubelet-standalone-managed container.
It is **not** for:
BlueField OS,
DOCA tree, not to a bare-metal deployment),
BlueFields (the bundle covers [`doca-container-deployment`](../doca-container-deployment/SKILL.md) for the single-host kubelet-standalone shape; **fleet/production-scale deployment is fleet-orchestration scope** — route to the orchestration entry-point in [`doca-public-knowledge-map ## Deploying DOCA services at scale`](../doca-public-knowledge-map/references/map.md#deploying-doca-services-at-scale--orchestration-entry-point-personascale-routing) (DPF / Network Operator / Launch Kit), not hand-rolled static-pod loops),
belong on [`doca-setup ## no-install`](../doca-setup/TASKS.md#no-install).
The skill teaches the agent the bare-metal-deployment *procedure* and the rules for quoting documented commands from the public DOCA Programming Guide and the public BlueField / DPU User Manual via [`doca-public-knowledge-map`](../doca-public-knowledge-map/SKILL.md); it does not invent flag names, PCI BDFs, NUMA numbers, devlink paths, representor strings, or systemd `Restart=` mode names from memory.
Load this skill when the user is doing **hands-on bare-metal deployment of a DOCA-linked application binary** on either of the two supported host modes (host x86 or BlueField Arm), or asking a cross-cutting bare-metal question that is not specific to one library's API. Concretely:
with a BlueField NIC in a PCIe slot, with DOCA installed on the host.
directly (BlueField Arm bare-metal mode), with DOCA installed on the Arm side per the BlueField OS image.
interactive debug; tmux/screen for long-running with manual reattach; systemd-supervised for restart-after-reboot, journald-integrated logs, and Restart= policy).
representor, the right NUMA node, and the right CPU set — and pinning IRQs to match — without inventing the addresses or the flag names.
controllers, network namespaces for multi-tenant deployments, `numactl` / `taskset` for CPU + NUMA binding) so multiple DOCA processes co-tenant on the same BlueField without crushing each other.
start, starts and exits immediately, runs but can't find the device, attache
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
Use for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. Trigger on NVIDIA products, hardware, software,…
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and…
Use when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA AI-Q Blueprint infrastructure.
Use when asked to run deep research or AI-Q research through a reachable NVIDIA AI-Q Blueprint backend.
Calibrate a new dataset from live RTSP camera streams via the AutoMagicCalib REST API. Use when the user provides RTSP URLs or asks to calibrate live cameras;…
Run end-to-end calibration on the shipped sample dataset (sdg_08_2_sample_data_010926.zip) against a running AMC microservice. Use when user says 'test sample…