/exploring-llm-costs
Investigate LLM spend in PostHog — total cost over time, cost by model, provider, user, trace, or custom dimension, token and cache-hit economics, and cost regressions. Use when the user asks "how much are we spending on LLMs?", "which model / user / feature is most expensive?",
$ npx -y skills add posthog/posthog --skill exploring-llm-costs --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/exploring-llm-costs
Context preview
The summary Claude sees to decide when to auto-load this skill.
Investigate LLM spend in PostHog — total cost over time, cost by model, provider, user, trace, or custom dimension, token and cache-hit economics, and cost regressions. Use when the user asks "how much are we spending on LLMs?", "which model / user / feature is most expensive?",
SKILL.md
exploring-llm-costs.SKILL.mdname: exploring-llm-costs
description: >
Investigate LLM spend in PostHog — total cost over time, cost by model,
provider, user, trace, or custom dimension, token and cache-hit economics,
and cost regressions. Use when the user asks "how much are we spending on
LLMs?", "which model / user / feature is most expensive?", "why did cost
spike?", wants to build a cost dashboard or alert, or pastes a trace URL
and asks about its cost.
Exploring LLM costs
PostHog attaches per-call cost metadata to every `$ai_generation` and `$ai_embedding` event at ingestion time. Every cost question reduces to an aggregation over those two event types — the interesting variation is only in how you group, filter, and compare.
This skill covers the common cost investigations: total spend, breakdowns (model, provider, user, trace, custom property), token and cache-hit analysis, regression debugging, and materializing results as insights, dashboards, or alerts.
Tools
| Tool | Purpose | | ------------------------------- | ------------------------------------------------------------------- | | `posthog:execute-sql` | Ad-hoc HogQL for any cost aggregation — the workhorse of this skill | | `posthog:query-llm-traces-list` | List traces with rolled-up cost, token, and error metrics | | `posthog:query-llm-trace` | Cost breakdown of a single trace across all its events | | `posthog:read-data-schema` | Discover which custom properties exist for breakdowns | | `posthog:insight-create` | Materialize a cost chart as a saved insight | | `posthog:dashboard-create` | Bundle cost insights into a dashboard | | `posthog:alert-create` | Alert when cost crosses a threshold |
Core rules
Three rules cover most of what goes wrong:
- **Sum `$ai_total_cost_usd` for rollups, never the components.** Components drop
request and web-search fees. The UI's cost cells sum `$ai_total_cost_usd` over `event IN ('$ai_generation', '$ai_embedding')`; mirror that. Full schema and rationale in [cost properties](./references/cost-properties.md).
- **Always include both `$ai_generation` and `$ai_embedding`** in cost queries
unless the project demonstrably does not use embeddings — missing them silently under-counts. `$ai_trace` and `$ai_span` carry no rollup cost; some SDK wrappers duplicate `$ai_total_cost_usd` onto `$ai_trace` so don't include it in rollups or you'll double-count.
- **Always set a time range.** Cost queries without one scan the full events table.
`$ai_total_cost_usd` is set at ingestion via one of three paths (passthrough, custom pricing, automatic lookup). When a cost looks wrong, read `$ai_cost_model_source` first — see [cost sources](./references/cost-sources.md) for the precedence rules and a diagnostic query.
Cache-hit math depends on whether the provider reports cache tokens inclusively or exclusively of `$ai_input_tokens`. Always branch on the per-event `$ai_cache_reporting_exclusive` flag, never on provider name — see [cache accounting](./references/cache-accounting.md) for the exclusive-vs-inclusive formula.
`distinct_id` is the canonical user dimension. Customers often attach custom properties (`feature`, `tenant_id`, `workflow_name`) — discover them with `posthog:read-data-schema` before grouping. Don't guess names.
Workflow: total spend in a window
posthog:execute-sql
SELECT round(sum(toFloat(properties.$ai_total_cost_usd)), 4) AS total_cost_usd
FROM events
WHERE event IN ('$ai_generation', '$ai_embedding')
AND timestamp >= now() - INTERVAL 30 DAYWorkflow: cost breakdowns
Every cost question is a variation of the same template — group by a dimension, aggregate `$ai_total_cost_usd`. See [breakdown patterns](./references/breakdown-patterns.md) for ready-to-run recipes:
- Cost over time (daily)
- Cost by model
- Cost by user (top spenders)
- Cost by trace (top expensive traces)
- Cost by custom dimension
- Cost-per-call distribution
- Input vs output vs cache economics
Workflow: inspect a single trace's cost
When the user pastes a trace URL and asks about its cost, fetch the trace and surface the per-event breakdown:
posthog:query-llm-trace
{ "traceId": "<trace_id>", "dateRange": {"date_from": "-30d"} }Sum `$ai_total_cost_usd` across the returned events, grouped by span name or model, to show which step(s) drove the cost. The trace response already includes `totalCost` as a convenience.
Workflow: debug a cost regression
"Our LLM bill jumped — why?" is almost always one of: more calls, bigger prompts, a new model, or a change in cache-hit rate. Work through them in order — see [regression debugging](./references/regression-debugging.md) for the 5-step playbook.
Workflow: materialize as an insight, dashboard, or alert
After ad-hoc queries answer the question, persist them as insights, bundle into a dashboard, or wire up alerts. See [materializing](./references/materializing.md) for ready-to-run JSON for `posthog:insight-create`, `posthog:dashboard-create`, and `posthog:alert-create`.
Constructing UI links
- **Dashboard**: `https://app.posthog.com/ai-observability/dashboard`
- **Traces list** (sort by cost): `https://app.posthog.com/ai-observability/traces`
- **Generations list**: `https://app.posthog.com/ai-observability/generations`
- **Users list** (per-user cost): `https://app.posthog.com/ai-observability/users`
- **Single trace**: `https://app.posthog.com/ai-observability/traces/<trace_id>?timestamp=<url_encoded_iso>`
Always surface a UI link so the user can verify visually.
Keeping this skill current
Provider reporting behavior (which tokens are inclusive vs exclusive, which costs show up where) shifts over time and can differ between SDK versions for the sam
Read more
name: exploring-llm-costs description: > Investigate LLM spend in PostHog — total cost over time, cost by model, provider, user, trace, or custom dimension, token and cache-hit economics, and cost regressions. Use when the user asks "how much are we spending on LLMs?", "which model / user / feature is most expensive?", "why did cost spike?", wants to build a cost dashboard or alert, or pastes a trace URL and asks about its cost.
Exploring LLM costs
PostHog attaches per-call cost metadata to every `$ai_generation` and `$ai_embedding` event at ingestion time. Every cost question reduces to an aggregation over those two event types — the interesting variation is only in how you group, filter, and compare.
This skill covers the common cost investigations: total spend, breakdowns (model, provider, user, trace, custom property), token and cache-hit analysis, regression debugging, and materializing results as insights, dashboards, or alerts.
Tools
| Tool | Purpose | | ------------------------------- | ------------------------------------------------------------------- | | `posthog:execute-sql` | Ad-hoc HogQL for any cost aggregation — the workhorse of this skill | | `posthog:query-llm-traces-list` | List traces with rolled-up cost, token, and error metrics | | `posthog:query-llm-trace` | Cost breakdown of a single trace across all its events | | `posthog:read-data-schema` | Discover which custom properties exist for breakdowns | | `posthog:insight-create` | Materialize a cost chart as a saved insight | | `posthog:dashboard-create` | Bundle cost insights into a dashboard | | `posthog:alert-create` | Alert when cost crosses a threshold |
Core rules
Three rules cover most of what goes wrong:
- **Sum `$ai_total_cost_usd` for rollups, never the components.** Components drop
request and web-search fees. The UI's cost cells sum `$ai_total_cost_usd` over `event IN ('$ai_generation', '$ai_embedding')`; mirror that. Full schema and rationale in [cost properties](./references/cost-properties.md).
- **Always include both `$ai_generation` and `$ai_embedding`** in cost queries
unless the project demonstrably does not use embeddings — missing them silently under-counts. `$ai_trace` and `$ai_span` carry no rollup cost; some SDK wrappers duplicate `$ai_total_cost_usd` onto `$ai_trace` so don't include it in rollups or you'll double-count.
- **Always set a time range.** Cost queries without one scan the full events table.
`$ai_total_cost_usd` is set at ingestion via one of three paths (passthrough, custom pricing, automatic lookup). When a cost looks wrong, read `$ai_cost_model_source` first — see [cost sources](./references/cost-sources.md) for the precedence rules and a diagnostic query.
Cache-hit math depends on whether the provider reports cache tokens inclusively or exclusively of `$ai_input_tokens`. Always branch on the per-event `$ai_cache_reporting_exclusive` flag, never on provider name — see [cache accounting](./references/cache-accounting.md) for the exclusive-vs-inclusive formula.
`distinct_id` is the canonical user dimension. Customers often attach custom properties (`feature`, `tenant_id`, `workflow_name`) — discover them with `posthog:read-data-schema` before grouping. Don't guess names.
Workflow: total spend in a window
posthog:execute-sql
SELECT round(sum(toFloat(properties.$ai_total_cost_usd)), 4) AS total_cost_usd
FROM events
WHERE event IN ('$ai_generation', '$ai_embedding')
AND timestamp >= now() - INTERVAL 30 DAYWorkflow: cost breakdowns
Every cost question is a variation of the same template — group by a dimension, aggregate `$ai_total_cost_usd`. See [breakdown patterns](./references/breakdown-patterns.md) for ready-to-run recipes:
- Cost over time (daily)
- Cost by model
- Cost by user (top spenders)
- Cost by trace (top expensive traces)
- Cost by custom dimension
- Cost-per-call distribution
- Input vs output vs cache economics
Workflow: inspect a single trace's cost
When the user pastes a trace URL and asks about its cost, fetch the trace and surface the per-event breakdown:
posthog:query-llm-trace
{ "traceId": "<trace_id>", "dateRange": {"date_from": "-30d"} }Sum `$ai_total_cost_usd` across the returned events, grouped by span name or model, to show which step(s) drove the cost. The trace response already includes `totalCost` as a convenience.
Workflow: debug a cost regression
"Our LLM bill jumped — why?" is almost always one of: more calls, bigger prompts, a new model, or a change in cache-hit rate. Work through them in order — see [regression debugging](./references/regression-debugging.md) for the 5-step playbook.
Workflow: materialize as an insight, dashboard, or alert
After ad-hoc queries answer the question, persist them as insights, bundle into a dashboard, or wire up alerts. See [materializing](./references/materializing.md) for ready-to-run JSON for `posthog:insight-create`, `posthog:dashboard-create`, and `posthog:alert-create`.
Constructing UI links
- **Dashboard**: `https://app.posthog.com/ai-observability/dashboard`
- **Traces list** (sort by cost): `https://app.posthog.com/ai-observability/traces`
- **Generations list**: `https://app.posthog.com/ai-observability/generations`
- **Users list** (per-user cost): `https://app.posthog.com/ai-observability/users`
- **Single trace**: `https://app.posthog.com/ai-observability/traces/<trace_id>?timestamp=<url_encoded_iso>`
Always surface a UI link so the user can verify visually.
Keeping this skill current
Provider reporting behavior (which tokens are inclusive vs exclusive, which costs show up where) shifts over time and can differ between SDK versions for the sam
:hedgehog: PostHog is the leading platform for building self-driving products. Our developer tools – AI observability, analytics, session replay, flags, experiments, error tracking, logs, and more – capture all the context agents need to diagnose problems, uncover opportunities, and ship fixes. Steer it all from Slack, web, desktop, or the MCP.
Repo: posthog/posthog
Other skills on posthog.
- /analyzing-expensive-users
Analyze the most expensive users in AI observability and explain why they cost so much. Use when the user asks about top spenders, expensive users, per-user LLM cost, user-level cost drivers, or patterns behind high AI observability spend.
Open skill - /creating-online-evaluations
Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified. Use when the user wants evaluations that automatically score new generations or whole traces going forward — "create an eval to catch X", "continuously
Open skill - /exploring-ai-failures
Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces. Use when someone wants to understand what's going wrong with an AI feature, find and categorize failure modes, triage errors, or investigate quality issues
Open skill - /exploring-llm-clusters
Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.
Open skill - /exploring-llm-evaluations
Investigate AI observability evaluations — `hog` (deterministic code-based), `llm_judge` (LLM-prompt-based), and `sentiment` (user-message sentiment). Find existing evaluations, inspect their configuration, run them against specific generations, query individual results, and
Open skill - /exploring-llm-traces
ABSOLUTE MUST to debug and inspect LLM/AI agent traces using PostHog's MCP tools. Use when the user pastes a trace or session URL (e.g. /ai-observability/traces/<id> or /ai-observability/sessions/<id>), asks to debug a trace, figure out what went wrong, check if an agent used a
Open skill

