Skip to content
Development
Skill

/performance

CLI performance optimization - startup time, memory usage, token savings benchmarking

From plugin
rtk
75k12 skills6 agents9 commands
Install
$ npx -y skills add rtk-ai/rtk --skill performance --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/performance

Context preview

The summary Claude sees to decide when to auto-load this skill.

CLI performance optimization - startup time, memory usage, token savings benchmarking

SKILL.md

performance.SKILL.md
description: CLI performance optimization - startup time, memory usage, token savings benchmarking

Performance Optimization Skill

Systematic performance analysis and optimization for RTK CLI tool, focusing on **startup time (<10ms)**, **memory usage (<5MB)**, and **token savings (60-90%)**.

When to Use

  • **Automatically triggered**: After filter changes, regex modifications, or dependency additions
  • **Manual invocation**: When performance degradation suspected or before release
  • **Proactive**: After any code change that could impact startup time or memory

RTK Performance Targets

| Metric | Target | Verification Method | Failure Threshold | |--------|--------|---------------------|-------------------| | **Startup time** | <10ms | `hyperfine 'rtk <cmd>'` | >15ms = blocker | | **Memory usage** | <5MB resident | `/usr/bin/time -l rtk <cmd>` (macOS) | >7MB = blocker | | **Token savings** | 60-90% | Tests with `count_tokens()` | <60% = blocker | | **Binary size** | <5MB stripped | `ls -lh target/release/rtk` | >8MB = investigate |

Performance Analysis Workflow

1. Establish Baseline

Before making any changes, capture current performance:

# Startup time baseline
hyperfine 'rtk git status' --warmup 3 --export-json /tmp/baseline_startup.json

# Memory usage baseline (macOS)
/usr/bin/time -l rtk git status 2>&1 | grep "maximum resident set size" > /tmp/baseline_memory.txt

# Memory usage baseline (Linux)
/usr/bin/time -v rtk git status 2>&1 | grep "Maximum resident set size" > /tmp/baseline_memory.txt

# Binary size baseline
ls -lh target/release/rtk | tee /tmp/baseline_binary_size.txt

2. Make Changes

Implement optimization or feature changes.

3. Rebuild and Measure

# Rebuild with optimizations
cargo build --release

# Measure startup time
hyperfine 'target/release/rtk git status' --warmup 3 --export-json /tmp/after_startup.json

# Measure memory usage
/usr/bin/time -l target/release/rtk git status 2>&1 | grep "maximum resident set size" > /tmp/after_memory.txt

# Check binary size
ls -lh target/release/rtk | tee /tmp/after_binary_size.txt

4. Compare Results

# Startup time comparison
hyperfine 'rtk git status' 'target/release/rtk git status' --warmup 3

# Example output:
#   Benchmark 1: rtk git status
#     Time (mean ± σ):       6.2 ms ±   0.3 ms    [User: 4.1 ms, System: 1.8 ms]
#   Benchmark 2: target/release/rtk git status
#     Time (mean ± σ):       7.8 ms ±   0.4 ms    [User: 5.2 ms, System: 2.1 ms]
#
#   Summary
#     'rtk git status' ran 1.26 times faster than 'target/release/rtk git status'

# Memory comparison
diff /tmp/baseline_memory.txt /tmp/after_memory.txt

# Binary size comparison
diff /tmp/baseline_binary_size.txt /tmp/after_binary_size.txt

5. Identify Regressions

**Startup time regression** (>15% increase or >2ms absolute):

# Profile with flamegraph
cargo install flamegraph
cargo flamegraph -- target/release/rtk git status

# Open flamegraph.svg
open flamegraph.svg
# Look for:
# - Repeated fixed-regex compilation (should be in LazyLock init)
# - Excessive allocations
# - File I/O on startup (should be zero)

**Memory regression** (>20% increase or >1MB absolute):

# Profile allocations (requires nightly)
cargo +nightly build --release -Z build-std
RUSTFLAGS="-C link-arg=-fuse-ld=lld" cargo +nightly build --release

# Use DHAT for heap profiling
cargo install dhat
# Add to main.rs:
# #[global_allocator]
# static ALLOC: dhat::Alloc = dhat::Alloc;

**Token savings regression** (<60% savings):

# Run token accuracy tests
cargo test test_token_savings

# Example failure output:
# Git log filter: expected ≥60% savings, got 52.3%

# Fix: Improve filter condensation logic

Common Performance Issues

Issue 1: Regex Recompilation

**Symptom**: Startup time >20ms, flamegraph shows regex compilation in hot path

**Detection**:

# Flamegraph shows Regex::new() calls during execution
cargo flamegraph -- target/release/rtk git log -10
# Check whether fixed patterns are compiled outside LazyLock statics

**Fix**:

// ❌ WRONG: Recompiled on every call
fn filter_line(line: &str) -> Option<&str> {
    let re = Regex::new(r"pattern").unwrap(); // RECOMPILED!
    re.find(line).map(|m| m.as_str())
}

// ✅ RIGHT: Compiled once with LazyLock
use std::sync::LazyLock;

static LINE_PATTERN: LazyLock<Regex> =
    LazyLock::new(|| Regex::new(r"pattern").unwrap());

fn filter_line(line: &str) -> Option<&str> {
    LINE_PATTERN.find(line).map(|m| m.as_str())
}

Issue 2: Excessive Allocations

**Symptom**: Memory usage >5MB, many small allocations in flamegraph

**Detection**:

# DHAT heap profiling
cargo +nightly build --release
valgrind --tool=dhat target/release/rtk git status

**Fix**:

// ❌ WRONG: Allocates Vec for every line
fn filter_lines(input: &str) -> String {
    input.lines()
        .map(|line| line.to_string()) // Allocates String
        .collect::<Vec<_>>()
        .join("\n")
}

// ✅ RIGHT: Borrow slices, single allocation
fn filter_lines(input: &str) -> String {
    input.lines()
        .collect::<Vec<_>>() // Vec of &str (no String allocation)
        .join("\n")
}

Issue 3: Startup I/O

**Symptom**: Startup time varies wildly (5ms to 50ms), flamegraph shows file reads

**Detection**:

# strace on Linux
strace -c target/release/rtk git status 2>&1 | grep -E "open|read"

# dtrace on macOS (requires SIP disabled)
sudo dtrace -n 'syscall::open*:entry { @[execname] = count(); }' &
target/release/rtk git status
sudo pkill dtrace

**Fix**:

// ❌ WRONG: File I/O on startup
fn main() {
    let config = load_config().unwrap(); // Reads ~/.config/rtk/config.toml
    // ...
}

// ✅ RIGHT: Lazy config loading (only if needed)
fn main() {
    // No I/O on startup
    // Config loaded on-demand when first accessed
}

Issue 4: Dependency Bloat

**Symptom**: Binary size >5MB, many unused depende

Read more
Ships withrtk

CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies

Get the whole plugin