cavecrew
When to delegate to `cavecrew-investigator` (locate code), `cavecrew-builder` (1-2 file edit) or `cavecrew-reviewer` (diff review) instead of working inline or…
Wire a repository through the Caveman Cloud gateway so every LLM request is measured, with no behavior change. Use for "set up caveman" or adding LLM spend observability.
$ npx -y skills add JuliusBrussee/caveman --skill caveman-setup --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/caveman-setupContext preview
The summary Claude sees to decide when to auto-load this skill.
Wire a repository through the Caveman Cloud gateway so every LLM request is measured, with no behavior change. Use for "set up caveman" or adding LLM spend observability.
name: caveman-setup description: > Wire a repository through the Caveman Cloud gateway so every LLM request is measured, with no behavior change. Use for "set up caveman" or adding LLM spend observability.
You are wiring this repository through the Caveman gateway. Caveman is a byte-preserving LLM proxy: in record mode it measures what your app sends and what it costs, and changes nothing else. Your job is a minimal, verified integration — not a refactor.
The prompt that sent you here provides four values. Refer to them as:
If any value is missing, stop and ask for it. Do not guess a URL or mint a key.
1. **Coherent integration.** Wire every live LLM callsite through existing configuration and responsible seams. Touch each layer correctness requires. No drive-by refactors or formatting sweeps; add an abstraction only when it clarifies ownership or lowers lifecycle cost. 2. **Secrets stay in env vars.** `CAVE_API_KEY` goes into the env file the repo already uses (`.env`, `.env.local`, …). If that file isn't gitignored, add it to `.gitignore` and say so. Never hardcode the key in source. 3. **Report only what you observed.** The final report states the HTTP status and usage numbers from the real verification response — never assumed success. If verification fails, report the failure template instead. 4. **Record mode only.** You are adding measurement. You do not enable any optimization, and you do not claim any savings — verified savings are $0 until an optimizer is explicitly turned on and passes its eval gate. 5. **Provider keys are not your business.** With `PROVIDER_KEYS: stored` you never see one. With `byok`, the app's existing provider key stays exactly where it already is.
Read dependency files (`package.json`, `requirements.txt`, `pyproject.toml`, `go.mod`, lockfiles) and search the source for LLM clients:
`@ai-sdk/*` (Vercel), `langchain*`, `litellm`, `google-genai` / `@google/genai`, `crewai`, `pydantic_ai`, `openai-agents` / `agents`
`ANTHROPIC_BASE_URL`, `GEMINI_BASE_URL`, `GOOGLE_GEMINI_BASE_URL`
List what you found (file:line per callsite) before changing anything. If you find **no** LLM callsites, stop and report the "nothing to wire" template at the end of this file — do not invent an integration.
One slug names this app in the gateway path: `GATEWAY/w/<app>`. Derive it from the package/module name (e.g. `support-bot`, `acme-api`). Grammar: lowercase `[a-z0-9]` first, then `[a-z0-9._-]`, max 64 chars. Spend for this whole app groups under that slug on the dashboard.
The pattern is always the same: **base URL → the gateway with `/w/<app>`, plus one auth header.** Gateway auth is `x-cave-api-key: CAVE_API_KEY` (`Authorization: Bearer CAVE_API_KEY` also works where a header is awkward). With `PROVIDER_KEYS: byok`, also send `x-cave-upstream-key: <the provider key the app already uses>`.
Two facts that make the wiring safe (both are gateway-enforced, not hopes): the gateway rebuilds upstream auth headers from scratch, so a client's `Authorization`/`x-api-key` value is never forwarded to the provider; and with `stored`, upstream auth comes from the encrypted connection server-side. So in `stored` mode, where an SDK insists on an api-key parameter, set it to the Cave key — it authenticates the gateway and goes no further.
Exact shapes (use the one matching each callsite — these are the product's published recipes, not suggestions):
**OpenAI SDK (TS)** — Chat Completions and Responses both route through:
const client = new OpenAI({
baseURL: `${process.env.CAVE_GATEWAY_URL}/w/<app>/openai/v1`,
apiKey: process.env.OPENAI_API_KEY, // byok: unchanged · stored: use CAVE_API_KEY
defaultHeaders: {
"x-cave-api-key": process.env.CAVE_API_KEY!,
// byok only:
"x-cave-upstream-key": process.env.OPENAI_API_KEY!,
},
});**OpenAI SDK (Python)** — same shape: `base_url=f"{gw}/w/<app>/openai/v1"`, `default_headers={"x-cave-api-key": ..., "x-cave-upstream-key": ...}`.
**Anthropic SDK (TS/Python)** — the SDK appends `/v1/messages` itself. The `x-cave-api-key` header is required here in both modes (this SDK's own key param rides `x-api-key`, which is not a gateway-auth header):
client = anthropic.Anthropic(
base_url=f"{os.environ['CAVE_GATEWAY_URL']}/w/<app>",
api_key=os.environ["ANTHROPIC_API_KEY"], # byok: unchanged · stored: use CAVE_API_KEY
default_headers={
"x-cave-api-key": os.environ["CAVE_API_KEY"],
# byok only:
"x-cave-upstream-key": os.environ["ANTHROPIC_API_KEY"],
},
)**Vercel AI SDK** — `createOpenAICompatible({ baseURL: `${gw}/w/<app>/openai/v1`, headers: { "x-cave-api-key": ... } })`; Anthropic models via `createAnthropic({ baseURL: `${gw}/w/<app>/v1`, headers: { ... } })`.
**LangChain / LangGraph** — `ChatOpenAI(base_url=f"{gw}/w/<app>/openai/v1", default_headers={...})`; `ChatAnthropic(base_url=f"{gw}/w/<app>", default_headers={...})`. LangGraph inherits whatever model you pass it.
**LiteLLM** — per call `api_base=f"{gw}/w/<app>/openai/v1"` + `extra_headers={...}`, or fleet-wide in the LiteLLM proxy `config.yaml`.
**Raw HTTP / anything else** — s
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
Repo: JuliusBrussee/caveman
When to delegate to `cavecrew-investigator` (locate code), `cavecrew-builder` (1-2 file edit) or `cavecrew-reviewer` (diff review) instead of working inline or…
Compress a memory file such as CLAUDE.md or a todo list into caveman format to save input tokens, keeping a readable backup. Trigger: /caveman-compress.
Show real token usage and estimated savings for the current session, read from the session log. Trigger: /caveman-stats.
Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for…
Write a Conventional Commits message compressed to intent only. Use for "write a commit", "commit message", /commit or /caveman-commit.
Find and label every LLM workflow in the repository so Caveman Cloud groups spend by workflow instead of one bucket. Use for "discover workflows" or breaking…