custom-allocators
Custom allocator skill for memory allocation strategies. Use when implementing…
Rust build time optimization skill for reducing slow compilation. Use when using cargo-timings to profile builds, configuring sccache for Rust, using the Cranelift backend, splitting workspaces for parallelism, choosing between thin LTO and fat LTO, or using the mold linker with
$ npx -y skills add mohitmishra786/low-level-dev-skills --skill rust-build-times --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/rust-build-timesContext preview
The summary Claude sees to decide when to auto-load this skill.
Rust build time optimization skill for reducing slow compilation. Use when using cargo-timings to profile builds, configuring sccache for Rust, using the Cranelift backend, splitting workspaces for parallelism, choosing between thin LTO and fat LTO, or using the mold linker with
name: rust-build-times description: Rust build time optimization skill for reducing slow compilation. Use when using cargo-timings to profile builds, configuring sccache for Rust, using the Cranelift backend, splitting workspaces for parallelism, choosing between thin LTO and fat LTO, or using the mold linker with Rust. Activates on queries about slow Rust compilation, cargo-timings, sccache Rust, cranelift backend, Rust workspace splitting, LTO tradeoffs, or mold linker with Rust.
Guide agents through diagnosing and improving Rust compilation speed: `cargo-timings` for build profiling, `sccache` for caching, the Cranelift codegen backend for faster dev builds, workspace crate splitting, LTO configuration trade-offs, and fast linkers (mold/lld).
# Build with timing report cargo build --timings # Opens build/cargo-timings/cargo-timing.html # Shows: crate compilation timeline, parallelism, bottlenecks # For release builds cargo build --release --timings # Key things to look for in the timing report: # - Long sequential chains (no parallelism) # - Individual crates taking > 10s (candidates for optimization) # - Proc-macro crates blocking everything downstream
# cargo-llvm-lines — count LLVM IR lines per function (monomorphization) cargo install cargo-llvm-lines cargo llvm-lines --release | head -20 # Shows functions generating the most LLVM IR (template explosion)
# Install cargo install sccache # or: brew install sccache # Configure for Rust builds export RUSTC_WRAPPER=sccache # Add to .cargo/config.toml (project or global) # ~/.cargo/config.toml [build] rustc-wrapper = "sccache" # Check cache stats sccache --show-stats # S3 backend for CI teams export SCCACHE_BUCKET=my-rust-cache export SCCACHE_REGION=us-east-1 export AWS_ACCESS_KEY_ID=xxx export AWS_SECRET_ACCESS_KEY=yyy sccache --start-server # GitHub Actions with sccache # - uses: mozilla-actions/sccache-action@v0.0.4
Cranelift is a fast codegen backend (vs LLVM) — produces slower code but compiles much faster. Ideal for development builds:
# Install nightly (Cranelift requires nightly for now) rustup toolchain install nightly rustup component add rustc-codegen-cranelift-preview --toolchain nightly # Use Cranelift for dev builds only # .cargo/config.toml [unstable] codegen-backend = true [profile.dev] codegen-backend = "cranelift"
# Use per-build CARGO_PROFILE_DEV_CODEGEN_BACKEND=cranelift \ RUSTFLAGS="-Zunstable-options" \ cargo +nightly build
Cranelift vs LLVM trade-off:
A single large crate compiles sequentially. Split into smaller crates to enable Cargo parallelism:
# Before: one giant crate
[package]
name = "monolith" # everything in one crate = sequential compile
# After: workspace with parallel crates
[workspace]
members = [
"core", # compiled in parallel
"networking", # no deps on ui → parallel with ui
"ui", # no deps on networking → parallel
"server", # depends on core + networking
"cli", # depends on core + ui
]# Visualize dependency graph cargo tree | head -30 cargo tree --graph | dot -Tsvg > deps.svg # visual graph # Check how many crates compile in parallel cargo build -j$(nproc) --timings # maximize parallelism
Rules for effective workspace splitting:
LTO improves runtime performance but increases link time:
# Cargo.toml profile configuration [profile.release] lto = "thin" # thin LTO: good performance, much faster than "fat" codegen-units = 1 # needed for best optimization (but disables parallelism) [profile.release-fast] inherits = "release" lto = "fat" # full LTO: maximum performance, very slow link [profile.dev] lto = "off" # never use LTO in dev (compilation speed) codegen-units = 16 # maximize parallel codegen in dev
LTO comparison:
| Setting | Link time | Runtime perf | Use when | |---------|-----------|-------------|---------| | `lto = false` | Fast | Baseline | Dev builds | | `lto = "thin"` | Moderate | +5–15% | Most release builds | | `lto = "fat"` | Slow | +15–30% | Maximum performance | | `codegen-units = 1` | Slowest | Best | With LTO for release |
The linker is often the bottleneck for large Rust projects:
# mold — fastest general-purpose linker (Linux) sudo apt-get install mold # .cargo/config.toml [target.x86_64-unknown-linux-gnu] linker = "clang" rustflags = ["-C", "link-arg=-fuse-ld=mold"] # Or use cargo-zigbuild (uses zig cc as linker) cargo install cargo-zigbuild cargo zigbuild --release # lld — LLVM's linker (faster than GNU ld, available everywhere) # .cargo/config.toml [target.x86_64-unknown-linux-gnu] rustflags = ["-C", "link-arg=-fuse-ld=lld"] # On macOS: zld or the default lld [target.x86_64-apple-darwin] rustflags = ["-C", "link-arg=-fuse-ld=/usr/local/bin/zld"]
Linker speed comparison (large project, typical):
# Reduce debug info level (faster but less debuggable) # Ca
A curated suite of AI agent skills for systems and low-level programming — C/C++, Rust, Zig, GPU, bare-metal firmware, Linux kernel/driver development, computer architecture, compiler internals, HPC, and more.
Repo: mohitmishra786/low-level-dev-skills
Custom allocator skill for memory allocation strategies. Use when implementing…
NUMA programming skill for multi-socket memory locality. Use when detecting NUMA topology,…
AF_XDP skill for high-performance XDP sockets. Use when creating AF_XDP sockets, configuring…
DPDK skill for userspace packet I/O. Use when initializing EAL, configuring PMD drivers,…
io_uring skill for Linux async I/O. Use when building high-performance servers with liburing,…
Bare-metal ADC and DAC skill. Use when configuring analog sampling, DMA-driven ADC,…