custom-allocators
Custom allocator skill for memory allocation strategies. Use when implementing…
Flamegraph generation and interpretation skill. Use when converting perf, Valgrind Callgrind, or other profiler output into SVG flamegraphs using Brendan Gregg's FlameGraph tools, or when reading flamegraphs to identify performance bottlenecks. Activates on queries about
$ npx -y skills add mohitmishra786/low-level-dev-skills --skill flamegraphs --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/flamegraphsContext preview
The summary Claude sees to decide when to auto-load this skill.
Flamegraph generation and interpretation skill. Use when converting perf, Valgrind Callgrind, or other profiler output into SVG flamegraphs using Brendan Gregg's FlameGraph tools, or when reading flamegraphs to identify performance bottlenecks. Activates on queries about
name: flamegraphs description: Flamegraph generation and interpretation skill. Use when converting perf, Valgrind Callgrind, or other profiler output into SVG flamegraphs using Brendan Gregg's FlameGraph tools, or when reading flamegraphs to identify performance bottlenecks. Activates on queries about flamegraphs, stackcollapse, flamegraph.svg, identifying hot frames, wide vs tall frames, or performance visualisation.
Guide agents through the pipeline from profiler data to SVG flamegraph, and teach interpretation of flamegraphs to drive concrete optimisation decisions.
git clone https://github.com/brendangregg/FlameGraph # No install needed; scripts are in the repo export PATH=$PATH:/path/to/FlameGraph
# Step 1: record perf record -F 999 -g -o perf.data ./prog # Step 2: generate script output perf script -i perf.data > out.perf # Step 3: collapse stacks stackcollapse-perf.pl out.perf > out.folded # Step 4: generate SVG flamegraph.pl out.folded > flamegraph.svg # Step 5: view xdg-open flamegraph.svg # Linux open flamegraph.svg # macOS
One-liner:
perf record -F 999 -g ./prog && perf script | stackcollapse-perf.pl | flamegraph.pl > fg.svg
# Collect two profiles perf record -g -o before.data ./prog_old perf record -g -o after.data ./prog_new # Collapse perf script -i before.data | stackcollapse-perf.pl > before.folded perf script -i after.data | stackcollapse-perf.pl > after.folded # Diff (red = regressed, blue = improved) difffolded.pl before.folded after.folded | flamegraph.pl > diff.svg
valgrind --tool=callgrind --callgrind-out-file=cg.out ./prog stackcollapse-callgrind.pl cg.out | flamegraph.pl > fg.svg
# Go pprof
go tool pprof -raw -output=prof.txt prog
stackcollapse-go.pl prof.txt | flamegraph.pl > fg.svg
# DTrace
dtrace -x ustackframes=100 -n 'profile-99 /execname=="prog"/ { @[ustack()] = count(); }' \
-o out.stacks sleep 10
stackcollapse.pl out.stacks | flamegraph.pl > fg.svg
# Java (async-profiler)
async-profiler -d 30 -f out.collapsed PID
flamegraph.pl out.collapsed > fg.svgA flamegraph is a call-stack visualisation:
**What to look for:**
| Pattern | Meaning | Action | |---------|---------|--------| | Wide frame near bottom | Function itself is hot | Optimise that function | | Wide frame with tall narrow towers | Calling many different callees | Hot dispatch; reduce call overhead | | Very tall stack with wide base | Deep recursion | Check recursion depth; consider iterative approach | | Plateau at the top | Leaf function with no callees | This leaf is the actual hotspot | | Many narrow identical stacks | Many threads doing the same work | Consider parallelism or batching |
**Identifying the actionable hotspot:**
1. Find the widest top frame (a frame with no or narrow children above it) 2. That is where CPU time is actually spent 3. Trace down to understand what called it and why
**Differential flamegraph:**
flamegraph.pl --title "My App" \
--subtitle "Release build, workload X" \
--width 1600 \
--height 16 \
--minwidth 0.5 \
--colors java \
out.folded > fg.svg| Option | Effect | |--------|--------| | `--title` | SVG title | | `--width` | Width in pixels | | `--height` | Frame height in pixels | | `--minwidth` | Omit frames < N% (reduces clutter) | | `--colors` | Palette: `hot` (default), `mem`, `io`, `java`, `js`, `perl`, `red`, `green`, `blue` | | `--inverted` | Icicle chart (roots at top) | | `--reverse` | Reverse stacks | | `--cp` | Consistent palette (same frame = same color across SVGs) |
For tool installation, stackcollapse scripts, and palette options, see [references/tools.md](references/tools.md).
A curated suite of AI agent skills for systems and low-level programming — C/C++, Rust, Zig, GPU, bare-metal firmware, Linux kernel/driver development, computer architecture, compiler internals, HPC, and more.
Repo: mohitmishra786/low-level-dev-skills
Custom allocator skill for memory allocation strategies. Use when implementing…
NUMA programming skill for multi-socket memory locality. Use when detecting NUMA topology,…
AF_XDP skill for high-performance XDP sockets. Use when creating AF_XDP sockets, configuring…
DPDK skill for userspace packet I/O. Use when initializing EAL, configuring PMD drivers,…
io_uring skill for Linux async I/O. Use when building high-performance servers with liburing,…
Bare-metal ADC and DAC skill. Use when configuring analog sampling, DMA-driven ADC,…