Skip to content
Development
Skill

/cost-local

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a hosted model. Use for local-vs-API break-even questions.

BOOST
From plugin
ruvnet-ruflo
74k197 skills164 agents196 commands3 MCP
Install
$ npx -y skills add ruvnet/ruflo --skill cost-local --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/cost-local

Context preview

The summary Claude sees to decide when to auto-load this skill.

Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a hosted model. Use for local-vs-API break-even questions.

SKILL.md

cost-local.SKILL.md
name: cost-local
description: Cost per million tokens on hardware you own (Ollama, llama.cpp, vLLM, LM Studio) from watts, electricity price, hardware price and measured tokens/second, and the utilisation at which local beats a hosted model. Use for local-vs-API break-even questions.
argument-hint: "--tok-per-s <n> [--busy 0.25] [--compare <model>] [--format json|markdown]"
allowed-tools: Bash

Cost Local

node ${CLAUDE_PLUGIN_ROOT}/scripts/local-cost.mjs --tok-per-s 80 --busy 0.25 --compare claude-haiku-4-5

Inputs (all the user's): `--watts-idle`, `--watts-active`, `--kwh`, `--hw-price`, `--life-years`, `--resale`. `--tok-per-s` is required and must be a **measured** generation speed.

Rules

  • Cost = **fixed** (amortisation + idle power, paid busy or not) + **marginal** (extra watts while generating). Utilisation decides the answer: always show the table at 5/25/50/100% busy.
  • It is an estimate from the inputs; say so. Nothing is measured.
  • Compare against a same-capability hosted model, not a frontier one; price alone is a poor reason to buy a GPU for sporadic use.
  • Read real counts from the server (Ollama `eval_count`/`prompt_eval_count` on the native API, llama.cpp `timings`, vLLM `/metrics`), not the lossy OpenAI-compatible `usage`, and count model reloads (`load_duration`) as overhead.
Ships withruvnet-ruflo

An agent meta-harness for Claude Code and Codex. 📖 RuFlo Explained — Build an AI Team That Plans, Remembers, Tests, and Improves A 14-chapter guide: from the basic idea to a first useful task, then memory, agent teams, plugins, cost and verification.

Get the whole plugin

Other skills on ruvnet-ruflo.