custom-allocators
Custom allocator skill for memory allocation strategies. Use when implementing…
Deep compiler optimizations skill for RA, ISel, and PGO. Use when explaining register allocation, instruction selection, LICM, vectorization limits, or profile-guided optimization beyond -O3. Activates on queries about register allocation, instruction selection, LICM,
$ npx -y skills add mohitmishra786/low-level-dev-skills --skill compiler-optimizations-deep --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/compiler-optimizations-deepContext preview
The summary Claude sees to decide when to auto-load this skill.
Deep compiler optimizations skill for RA, ISel, and PGO. Use when explaining register allocation, instruction selection, LICM, vectorization limits, or profile-guided optimization beyond -O3. Activates on queries about register allocation, instruction selection, LICM,
name: compiler-optimizations-deep description: Deep compiler optimizations skill for RA, ISel, and PGO. Use when explaining register allocation, instruction selection, LICM, vectorization limits, or profile-guided optimization beyond -O3. Activates on queries about register allocation, instruction selection, LICM, auto-vectorization failure, PGO, or BOLT.
Explain optimization phases beyond flags: mid-level IR opts, register allocation, instruction selection/scheduling, vectorization boundaries, PGO, and post-link BOLT — bridging `skills/compilers/pgo` and LLVM/GCC internals.
Frontend → LLVM IR / GCC GIMPLE ├── Mid-level: DCE, GVN, LICM, inlining ├── Loop opts: unroll, vectorize ├── Codegen prep: legalize types ├── Instruction selection (DAG → machine ops) ├── Register allocation (greedy, linear scan) └── Peephole / scheduling
clang -O3 -Rpass=loop-vectorize -Rpass-missed=loop-vectorize foo.c
| Miss reason | Typical fix | |-------------|-------------| | Unknown trip count | peel loop; assert count | | Dependence | reorder / separate accumulators | | Function call in loop | inline or outline | | Alignment unknown | `__builtin_assume_aligned` |
When live ranges exceed physical registers, the allocator **spills** to stack slots — costly loads/stores. Reducing live ranges (splitting variables, rematerialization) helps.
GCC/LLVM both use graph coloring variants (LLVM "greedy regalloc").
clang -fprofile-instr-generate -O2 -o app foo.c ./app # training workload llvm-profdata merge default.profraw -o default.profdata clang -fprofile-instr-use=default.profdata -O2 -o app_pgo foo.c
Improves branch layout, inlining, and vectorization thresholds.
See `skills/compilers/pgo` for GCC and BOLT.
llvm-bolt -instrument app -o app.inst ./app.inst llvm-bolt -data=perf.fdata -reorder-blocks=+ -o app.bolt app
Optimizes layout after linker — needs relocations (`-Wl,--emit-relocs`).
Loop-invariant code motion hoists `x * scale` out of inner loop when legal — reduces work per iteration.
/compiler-optimizations-deep Why did LLVM fail to vectorize this reduction loop?
| Symptom | Cause | Fix | |---------|-------|-----| | PGO no gain | Unrepresentative training | Match production input | | BOLT crash | Stripped binary | Keep symbols + relocs | | Spills in asm | Register pressure | Simplify live ranges | | `-O3` slower | Code bloat / cache | Try `-O2` or PGO | | Different GCC/Clang | Pass ordering differs | Compare IR + asm |
A curated suite of AI agent skills for systems and low-level programming — C/C++, Rust, Zig, GPU, bare-metal firmware, Linux kernel/driver development, computer architecture, compiler internals, HPC, and more.
Repo: mohitmishra786/low-level-dev-skills
Custom allocator skill for memory allocation strategies. Use when implementing…
NUMA programming skill for multi-socket memory locality. Use when detecting NUMA topology,…
AF_XDP skill for high-performance XDP sockets. Use when creating AF_XDP sockets, configuring…
DPDK skill for userspace packet I/O. Use when initializing EAL, configuring PMD drivers,…
io_uring skill for Linux async I/O. Use when building high-performance servers with liburing,…
Bare-metal ADC and DAC skill. Use when configuring analog sampling, DMA-driven ADC,…