custom-allocators
Custom allocator skill for memory allocation strategies. Use when implementing…
CPU pipeline skill for hazards, forwarding, and stalls. Use when explaining pipeline stages, data/control hazards, forwarding paths, or branch stalls in performance analysis. Activates on queries about pipeline hazard, data hazard, control hazard, forwarding, pipeline stall, or
$ npx -y skills add mohitmishra786/low-level-dev-skills --skill cpu-pipelines-and-hazards --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/cpu-pipelines-and-hazardsContext preview
The summary Claude sees to decide when to auto-load this skill.
CPU pipeline skill for hazards, forwarding, and stalls. Use when explaining pipeline stages, data/control hazards, forwarding paths, or branch stalls in performance analysis. Activates on queries about pipeline hazard, data hazard, control hazard, forwarding, pipeline stall, or
name: cpu-pipelines-and-hazards description: CPU pipeline skill for hazards, forwarding, and stalls. Use when explaining pipeline stages, data/control hazards, forwarding paths, or branch stalls in performance analysis. Activates on queries about pipeline hazard, data hazard, control hazard, forwarding, pipeline stall, or superscalar basics.
Explain classic and modern CPU pipeline concepts: stages, data and control hazards, forwarding/bypassing, stalls, and branch handling — foundational for optimization and understanding microarchitecture counters.
IF → ID → EX → MEM → WB
Overlapped execution: instruction N in EX while N+1 in ID.
| Type | Example | Mitigation | |------|---------|------------| | RAW (true) | `add r1,r2,r3` then `sub r4,r1,r5` | Forwarding from EX/MEM/WB | | WAR / WAW | Rare in in-order; relevant in OoO rename | Register renaming |
Without forwarding:
stall until writeback completes
Branch in ID → target unknown until EX ├── Predict taken/not-taken (static or dynamic) ├── Flush wrong-path instructions on mispredict └── Penalty = pipeline depth (varies by CPU)
See `skills/computer-architecture/branch-prediction-and-speculation`.
Limited functional units (single memory port) cause stalls even without dependencies.
/* Bad — tight dependency chain */
for (int i = 0; i < n; i++)
acc = acc + data[i]; /* each iter waits on acc */
/* Better — multiple accumulators (ILP) */
acc0 = acc1 = 0;
for (int i = 0; i < n; i += 2) {
acc0 += data[i];
acc1 += data[i+1];
}
acc = acc0 + acc1;Pair with `skills/low-level-programming/cpu-cache-opt` — memory often dominates.
perf stat -e instructions,cycles,stalls-frontend,stalls-backend ./app
VTune "Microarchitecture Exploration" maps to pipeline slots.
/cpu-pipelines-and-hazards Explain RAW hazard in this ARM assembly loop and how to break it
| Symptom | Cause | Fix | |---------|-------|-----| | High `stalls-frontend` | I-cache misses / branch mispredict | Align hot loop; see branch skill | | No speedup from unroll | Memory bound | Profile loads; prefetch | | Wrong cycle model | Ignored OoO execution | Use perf hardware counters | | "NOP fixes it" | Timing-sensitive MMIO | Never tune device delays by NOP |
A curated suite of AI agent skills for systems and low-level programming — C/C++, Rust, Zig, GPU, bare-metal firmware, Linux kernel/driver development, computer architecture, compiler internals, HPC, and more.
Repo: mohitmishra786/low-level-dev-skills
Custom allocator skill for memory allocation strategies. Use when implementing…
NUMA programming skill for multi-socket memory locality. Use when detecting NUMA topology,…
AF_XDP skill for high-performance XDP sockets. Use when creating AF_XDP sockets, configuring…
DPDK skill for userspace packet I/O. Use when initializing EAL, configuring PMD drivers,…
io_uring skill for Linux async I/O. Use when building high-performance servers with liburing,…
Bare-metal ADC and DAC skill. Use when configuring analog sampling, DMA-driven ADC,…