/cuda-profiling
CUDA profiling skill for NVIDIA GPU performance analysis. Use when profiling kernels with Nsight Systems or Nsight Compute, interpreting roofline models, diagnosing memory-bound vs compute-bound kernels, or annotating code with NVTX ranges. Activates on queries about Nsight,
$ npx -y skills add mohitmishra786/low-level-dev-skills --skill cuda-profiling --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/cuda-profiling
Context preview
The summary Claude sees to decide when to auto-load this skill.
CUDA profiling skill for NVIDIA GPU performance analysis. Use when profiling kernels with Nsight Systems or Nsight Compute, interpreting roofline models, diagnosing memory-bound vs compute-bound kernels, or annotating code with NVTX ranges. Activates on queries about Nsight,
SKILL.md
cuda-profiling.SKILL.mdname: cuda-profiling
description: CUDA profiling skill for NVIDIA GPU performance analysis. Use when profiling kernels with Nsight Systems or Nsight Compute, interpreting roofline models, diagnosing memory-bound vs compute-bound kernels, or annotating code with NVTX ranges. Activates on queries about Nsight, NCU, ncu CLI, GPU roofline, occupancy metrics, or CUDA profiling workflow.
CUDA Profiling
Purpose
Guide agents through profiling CUDA applications with Nsight Systems (timeline-level) and Nsight Compute (kernel-level metrics), using the NCU CLI for automated metric collection, interpreting roofline models, and diagnosing whether kernels are memory-bound or compute-bound.
When to Use
- A CUDA kernel is slower than expected and you need bottleneck identification
- Comparing kernel variants (tiling strategies, block sizes)
- Building CI performance regression checks with `ncu` metrics
- Correlating CPU and GPU activity in multi-stream pipelines
- Annotating application phases with NVTX for timeline visibility
- Interpreting occupancy, memory throughput, and SM utilization metrics
Workflow
1. Choose profiling tool
What do you need?
├── System-wide timeline (CPU+GPU+CUDA API) → Nsight Systems (nsys)
├── Per-kernel deep metrics (occupancy, memory) → Nsight Compute (ncu)
└── Quick metric from CLI in CI → ncu --metrics ...
2. Nsight Systems — timeline profiling
# Profile entire application
nsys profile --trace=cuda,nvtx,osrt --output=report ./my_cuda_app
# Open report
nsys-ui report.nsys-rep
# CLI summary
nsys stats report.nsys-rep
What to look for in the timeline:
- Gaps between kernel launches (CPU bottleneck or sync points)
- `cudaDeviceSynchronize` stalls
- Overlap between H2D copies and kernel execution across streams
- CUDA API call overhead
# Capture with CUDA graph info
nsys profile --capture-range=cudaProfilerApi ./my_cuda_app
3. NVTX range annotations
#include <nvtx3/nvToolsExt.h>
void pipeline(void) {
nvtxRangePushA("H2D copy");
cudaMemcpyAsync(d_in, h_in, size, cudaMemcpyHostToDevice, stream);
nvtxRangePop();
nvtxRangePushA("kernel");
my_kernel<<<grid, block, 0, stream>>>(d_in, d_out, n);
nvtxRangePop();
nvtxRangePushA("D2H copy");
cudaMemcpyAsync(h_out, d_out, size, cudaMemcpyDeviceToHost, stream);
nvtxRangePop();
}Compile with `-lnvToolsExt` or link nvtx3 header-only. Ranges appear as colored bands in Nsight Systems.
4. Nsight Compute — kernel analysis
# Profile all kernels, save report
ncu -o kernel_report ./my_cuda_app
# Profile specific kernel by name
ncu --kernel-name regex:matmul_tiled ./my_cuda_app
# Launch UI
ncu-ui kernel_report.ncu-rep
Key sections in NCU report:
- **Speed of Light**: SM throughput vs memory throughput vs peak
- **Occupancy**: Active warps vs hardware limit
- **Memory Workload Analysis**: L1/L2 hit rates, coalescing efficiency
- **Warp State Statistics**: Stall reasons (memory, barrier, dispatch)
5. NCU CLI metrics
# Essential metrics set
ncu --metrics \
sm__throughput.avg.pct_of_peak_sustained_elapsed,\
dram__throughput.avg.pct_of_peak_sustained_elapsed,\
sm__warps_active.avg.pct_of_peak_sustained_active,\
l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum,\
smsp__sass_thread_inst_executed_op_ffma_pred_on.sum \
./my_cuda_app
# CSV export for CI
ncu --csv --metrics dram__bytes_read.sum,dram__bytes_write.sum ./my_cuda_app
# Set kernel replay mode for accurate counters
ncu --kernel-replay-mode application ./my_cuda_app
6. Memory-bound vs compute-bound diagnosis
Roofline interpretation
├── dram__throughput near peak AND sm__throughput low → memory-bound
│ └── Fix: coalescing, shared mem tiling, reduce traffic
├── sm__throughput near peak AND dram low → compute-bound
│ └── Fix: tensor cores, loop unrolling, ILP
└── Both low → launch config, occupancy, or sync overhead
Roofline model (conceptual):
Performance (GFLOP/s)
| /\ compute roof
| / \
| / \____ memory roof (bandwidth-limited region)
| /
+------------------ Arithmetic Intensity (FLOP/byte)Measure arithmetic intensity: `smsp__sass_thread_inst_executed_op_ffma_pred_on.sum * 2 / dram__bytes.sum`
7. Occupancy analysis
ncu --metrics sm__warps_active.avg.pct_of_peak_sustained_active,\
launch__occupancy_limit_registers,\
launch__occupancy_limit_shared_mem,\
launch__occupancy_limit_block_size \
./my_cuda_app
| Limiting factor | Typical fix | |-----------------|-------------| | Registers | `-maxrregcount`, simplify kernel | | Shared memory | Reduce tile size, split phases | | Block size | Try 128 or 256 instead of 512+ |
8. Profiling workflow checklist
# 1. Build with line info (not -G unless debugging)
nvcc -lineinfo -O3 -arch=sm_80 -o app main.cu
# 2. Timeline first
nsys profile --trace=cuda,nvtx -o timeline ./app
# 3. Deep dive on hot kernel
ncu --kernel-name regex:hot_kernel --set full ./app
# 4. Compare before/after
ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v1 > v1.csv
ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v2 > v2.csv
Common Problems
| Symptom | Cause | Fix | |---------|-------|-----| | `ERR_NVGPUCTRPERM` | Insufficient profiling permissions | Run with sudo or set `NVreg_RestrictProfilingToAdminUsers=0` | | All metrics show zero | Profiling disabled or wrong GPU | Check `CUDA_VISIBLE_DEVICES`; use `--target-processes all` | | NCU report empty | Kernel too short or not launched | Increase workload; verify `cudaGetLastError()` | | Huge profiling overhead | Full metric sets on many kernels | Use `--kernel-name` filter; `--launch-skip` | | Timeline shows no overlap | Single default stream | Create multiple streams; use async copies | | Occupancy looks fine but kernel slow | Memory latency not hidden | Check memory coalescing; increase active warps |
R
Read more
name: cuda-profiling description: CUDA profiling skill for NVIDIA GPU performance analysis. Use when profiling kernels with Nsight Systems or Nsight Compute, interpreting roofline models, diagnosing memory-bound vs compute-bound kernels, or annotating code with NVTX ranges. Activates on queries about Nsight, NCU, ncu CLI, GPU roofline, occupancy metrics, or CUDA profiling workflow.
CUDA Profiling
Purpose
Guide agents through profiling CUDA applications with Nsight Systems (timeline-level) and Nsight Compute (kernel-level metrics), using the NCU CLI for automated metric collection, interpreting roofline models, and diagnosing whether kernels are memory-bound or compute-bound.
When to Use
- A CUDA kernel is slower than expected and you need bottleneck identification
- Comparing kernel variants (tiling strategies, block sizes)
- Building CI performance regression checks with `ncu` metrics
- Correlating CPU and GPU activity in multi-stream pipelines
- Annotating application phases with NVTX for timeline visibility
- Interpreting occupancy, memory throughput, and SM utilization metrics
Workflow
1. Choose profiling tool
What do you need? ├── System-wide timeline (CPU+GPU+CUDA API) → Nsight Systems (nsys) ├── Per-kernel deep metrics (occupancy, memory) → Nsight Compute (ncu) └── Quick metric from CLI in CI → ncu --metrics ...
2. Nsight Systems — timeline profiling
# Profile entire application nsys profile --trace=cuda,nvtx,osrt --output=report ./my_cuda_app # Open report nsys-ui report.nsys-rep # CLI summary nsys stats report.nsys-rep
What to look for in the timeline:
- Gaps between kernel launches (CPU bottleneck or sync points)
- `cudaDeviceSynchronize` stalls
- Overlap between H2D copies and kernel execution across streams
- CUDA API call overhead
# Capture with CUDA graph info nsys profile --capture-range=cudaProfilerApi ./my_cuda_app
3. NVTX range annotations
#include <nvtx3/nvToolsExt.h>
void pipeline(void) {
nvtxRangePushA("H2D copy");
cudaMemcpyAsync(d_in, h_in, size, cudaMemcpyHostToDevice, stream);
nvtxRangePop();
nvtxRangePushA("kernel");
my_kernel<<<grid, block, 0, stream>>>(d_in, d_out, n);
nvtxRangePop();
nvtxRangePushA("D2H copy");
cudaMemcpyAsync(h_out, d_out, size, cudaMemcpyDeviceToHost, stream);
nvtxRangePop();
}Compile with `-lnvToolsExt` or link nvtx3 header-only. Ranges appear as colored bands in Nsight Systems.
4. Nsight Compute — kernel analysis
# Profile all kernels, save report ncu -o kernel_report ./my_cuda_app # Profile specific kernel by name ncu --kernel-name regex:matmul_tiled ./my_cuda_app # Launch UI ncu-ui kernel_report.ncu-rep
Key sections in NCU report:
- **Speed of Light**: SM throughput vs memory throughput vs peak
- **Occupancy**: Active warps vs hardware limit
- **Memory Workload Analysis**: L1/L2 hit rates, coalescing efficiency
- **Warp State Statistics**: Stall reasons (memory, barrier, dispatch)
5. NCU CLI metrics
# Essential metrics set ncu --metrics \ sm__throughput.avg.pct_of_peak_sustained_elapsed,\ dram__throughput.avg.pct_of_peak_sustained_elapsed,\ sm__warps_active.avg.pct_of_peak_sustained_active,\ l1tex__t_sectors_pipe_lsu_mem_global_op_ld.sum,\ smsp__sass_thread_inst_executed_op_ffma_pred_on.sum \ ./my_cuda_app # CSV export for CI ncu --csv --metrics dram__bytes_read.sum,dram__bytes_write.sum ./my_cuda_app # Set kernel replay mode for accurate counters ncu --kernel-replay-mode application ./my_cuda_app
6. Memory-bound vs compute-bound diagnosis
Roofline interpretation ├── dram__throughput near peak AND sm__throughput low → memory-bound │ └── Fix: coalescing, shared mem tiling, reduce traffic ├── sm__throughput near peak AND dram low → compute-bound │ └── Fix: tensor cores, loop unrolling, ILP └── Both low → launch config, occupancy, or sync overhead
Roofline model (conceptual):
Performance (GFLOP/s)
| /\ compute roof
| / \
| / \____ memory roof (bandwidth-limited region)
| /
+------------------ Arithmetic Intensity (FLOP/byte)Measure arithmetic intensity: `smsp__sass_thread_inst_executed_op_ffma_pred_on.sum * 2 / dram__bytes.sum`
7. Occupancy analysis
ncu --metrics sm__warps_active.avg.pct_of_peak_sustained_active,\ launch__occupancy_limit_registers,\ launch__occupancy_limit_shared_mem,\ launch__occupancy_limit_block_size \ ./my_cuda_app
| Limiting factor | Typical fix | |-----------------|-------------| | Registers | `-maxrregcount`, simplify kernel | | Shared memory | Reduce tile size, split phases | | Block size | Try 128 or 256 instead of 512+ |
8. Profiling workflow checklist
# 1. Build with line info (not -G unless debugging) nvcc -lineinfo -O3 -arch=sm_80 -o app main.cu # 2. Timeline first nsys profile --trace=cuda,nvtx -o timeline ./app # 3. Deep dive on hot kernel ncu --kernel-name regex:hot_kernel --set full ./app # 4. Compare before/after ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v1 > v1.csv ncu --csv --metrics sm__throughput.avg.pct_of_peak_sustained_elapsed ./app_v2 > v2.csv
Common Problems
| Symptom | Cause | Fix | |---------|-------|-----| | `ERR_NVGPUCTRPERM` | Insufficient profiling permissions | Run with sudo or set `NVreg_RestrictProfilingToAdminUsers=0` | | All metrics show zero | Profiling disabled or wrong GPU | Check `CUDA_VISIBLE_DEVICES`; use `--target-processes all` | | NCU report empty | Kernel too short or not launched | Increase workload; verify `cudaGetLastError()` | | Huge profiling overhead | Full metric sets on many kernels | Use `--kernel-name` filter; `--launch-skip` | | Timeline shows no overlap | Single default stream | Create multiple streams; use async copies | | Occupancy looks fine but kernel slow | Memory latency not hidden | Check memory coalescing; increase active warps |
R
A curated suite of AI agent skills for systems and low-level programming — C/C++, Rust, Zig, GPU, bare-metal firmware, Linux kernel/driver development, computer architecture, compiler internals, HPC, and more.
Repo: mohitmishra786/low-level-dev-skills
Other skills on low-level-dev-skills.
- /custom-allocators
Custom allocator skill for memory allocation strategies. Use when implementing pool/slab/arena allocators, tuning jemalloc/mimalloc, writing Rust GlobalAlloc, or benchmarking allocator performance. Activates on queries about jemalloc, mimalloc, tcmalloc, arena allocator,
Open skill - /numa-programming
NUMA programming skill for multi-socket memory locality. Use when detecting NUMA topology, binding processes with numactl, using libnuma API, building NUMA-aware data structures, or measuring remote access penalties. Activates on queries about numactl, libnuma, NUMA topology,
Open skill - /af-xdp
AF_XDP skill for high-performance XDP sockets. Use when creating AF_XDP sockets, configuring UMEM and XSK rings, XDP_REDIRECT programs, copy vs zero-copy mode, or comparing with DPDK. Activates on queries about AF_XDP, xsk_umem, XDP_REDIRECT, libbpf xsk, or zero-copy XDP.
Open skill - /dpdk
DPDK skill for userspace packet I/O. Use when initializing EAL, configuring PMD drivers, using mbuf pools and rte_ring, setting up huge pages, RSS, or testpmd validation. Activates on queries about DPDK, EAL, rte_eth_rx_burst, hugepages, PMD, or testpmd.
Open skill - /io-uring
io_uring skill for Linux async I/O. Use when building high-performance servers with liburing, multi-shot operations, provided buffers, fixed files, zero-copy send, or tokio-uring. Activates on queries about io_uring, SQE/CQE, liburing, IORING_OP_PROVIDE_BUFFERS, or io_uring vs
Open skill - /adc-dac-baremetal
Bare-metal ADC and DAC skill. Use when configuring analog sampling, DMA-driven ADC, calibration, or DAC output on MCUs. Activates on queries about ADC bare-metal, sampling time, DMA ADC, or DAC channel setup.
Open skill

