Skip to content
Development
Skill

/cpu-gpu-performance

Establishes CPU/GPU baselines before resource-intensive operations. Use before builds, training runs, or any task that pins cores or GPUs for over a minute.

From plugin
claude-night-market
337200 skills59 agents162 commands1 MCP
Install
$ npx -y skills add athola/claude-night-market --skill cpu-gpu-performance --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/cpu-gpu-performance

Context preview

The summary Claude sees to decide when to auto-load this skill.

Establishes CPU/GPU baselines before resource-intensive operations. Use before builds, training runs, or any task that pins cores or GPUs for over a minute.

SKILL.md

cpu-gpu-performance.SKILL.md
name: cpu-gpu-performance
description: Establishes CPU/GPU baselines before resource-intensive operations. Use before builds, training runs, or any task that pins cores or GPUs for over a minute.
alwaysApply: false
progressive_loading: true
dependencies:
  hub:
  - token-conservation
  modules: []
model_hint: standard

Table of Contents

  • [When to Use](#when-to-use)
  • [Required TodoWrite Items](#required-todowrite-items)
  • [Step 1: Establish Current Baseline](#step-1-establish-current-baseline)
  • [Step 2: Narrow the Scope](#step-2-narrow-the-scope)
  • [Step 3: Instrument Before You Optimize](#step-3-instrument-before-you-optimize)
  • [Step 4: Throttle and Sequence Work](#step-4-throttle-and-sequence-work)
  • [Step 5: Log Decisions and Next Steps](#step-5-log-decisions-and-next-steps)
  • [Output Expectations](#output-expectations)

CPU/GPU Performance Discipline

When To Use

  • At the beginning of every session (auto-load alongside `token-conservation`).
  • Whenever you plan to build, train, or test anything that could pin CPU cores

or GPUs for more than a minute.

  • Before retrying a failing command that previously consumed significant resources.

When NOT To Use

  • Simple operations with no resource impact
  • Quick single-file operations

Required TodoWrite Items

1. `cpu-gpu-performance:baseline` 2. `cpu-gpu-performance:scope` 3. `cpu-gpu-performance:instrument` 4. `cpu-gpu-performance:throttle` 5. `cpu-gpu-performance:log`

Step 1: Establish Current Baseline

  • Capture current utilization:
  • `uptime`
  • `ps -eo pcpu,cmd | head`
  • `nvidia-smi --query-gpu=utilization.gpu,memory.used --format=csv`

Note which hosts/GPUs are already busy.

  • Record any CI/cluster budgets (time quotas, GPU hours) before launching work.
  • Set a per-task CPU minute / GPU minute budget that respects those limits.

Step 2: Narrow the Scope

  • Avoid running "whole world" jobs after a small fix. Prefer diff-based

or tag-based selective testing:

  • `pytest -k`
  • Bazel target patterns
  • `cargo test <module>`
  • Batch low-level fixes so you can validate multiple changes with a single targeted command.
  • For GPU jobs, favor unit-scale smoke inputs or lower epoch counts before

scheduling the full training/eval sweep.

Step 3: Instrument Before You Optimize

  • Pick the right profiler/monitor:
  • CPU work:
  • `perf`
  • `intel vtune`
  • `cargo flamegraph`
  • language-specific profilers
  • GPU work:
  • `nvidia-smi dmon`
  • `nsys`
  • `nvprof`
  • DLProf
  • framework timeline tracers
  • Capture kernel/ops timelines, memory footprints, and data pipeline latency

so you have evidence when throttling or parallelizing.

  • Record hot paths and I/O bottlenecks in notes so future reruns can jump straight to the culprit.

Step 4: Throttle and Sequence Work

  • Use `nice`, `ionice`, or Kubernetes/Slurm quotas to prevent starvation of shared nodes.
  • Chain heavy tasks with guardrails:
  • Rerun only the failed test/module
  • Then (optionally) escalate to the next-wider shard
  • Reserve the full suite for the final gate
  • Stagger GPU kernels (smaller batch sizes or gradient accumulation) when memory

pressure risks eviction; prefer checkpoint/restore over restarts.

Step 5: Log Decisions and Next Steps

Conclude by documenting the commands that were run and their resource cost (duration, CPU%, GPU%), confirming whether they remained within the per-task budget. If a full suite or long training run was necessary, justify why selective or staged approaches were not feasible. Capture any follow-up tasks, such as adding a new test marker or profiling documentation, to simplify future sessions.

Output Expectations

  • Brief summary covering:
  • baseline metrics
  • scope chosen
  • instrumentation captured
  • throttling tactics
  • follow-up items
  • Concrete example(s) of what ran (e.g.):
  • "reran `pytest tests/test_orders.py -k test_refund` instead of `pytest -m slow`"
  • "profiled `nvidia-smi dmon` output to prove GPU idle time before scaling"

Exit Criteria

  • [ ] `uptime` and `ps` baseline captured and recorded before any

build, training run, or test suite starts

  • [ ] Scope narrowed to diff-based or tag-based targets (e.g.,

`pytest -k`, `cargo test <module>`); full-suite justification documented if selective approach was not feasible

  • [ ] Output summary includes: duration, CPU% or GPU% consumed, and

whether the run stayed within the per-task budget

  • [ ] Any follow-up tasks (new test markers, profiling docs) written

to a todo or issue so they survive the session

Read more
Ships withclaude-night-market

A plugin marketplace for Claude Code. Install only the plugins you need to run git workflows, code review, spec-driven development, and autonomous agents from inside your Claude Code session.

Get the whole plugin

Other skills on claude-night-market.