admission-control
Use when the user asks to "write a validator", "add validation", "implement admission…
Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid memory growth. Walks through tsdb status endpoints, per-metric and per-label
$ npx -y skills add grafana/skills --skill prometheus-cardinality-troubleshooter --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/prometheus-cardinality-troubleshooterContext preview
The summary Claude sees to decide when to auto-load this skill.
Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid memory growth. Walks through tsdb status endpoints, per-metric and per-label
name: prometheus-cardinality-troubleshooter license: Apache-2.0 description: > Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid memory growth. Walks through tsdb status endpoints, per-metric and per-label drill-downs, common-culprit galleries, and remediation paths. Use when the user is *currently experiencing* a cardinality fire. For preventing cardinality issues at the source, route to prometheus-label-strategy. For post-ingest aggregation, route to adaptive-metrics. For DPM-specific analysis, route to dpm-finder.
You are an expert in diagnosing live Prometheus cardinality problems. When a user reports a Prometheus performance, memory, or cost issue that smells like cardinality, use this guide to triage systematically.
This skill is **diagnostic and operational**. For schema design and prevention, route to `prometheus-label-strategy`.
---
Under pressure, the tempting move is to `labeldrop` the high-cardinality label at scrape time. **Do not.** You cannot remove, at scrape time, any label that makes a series unique — not `pod`, not `instance`, not anything that distinguishes one real series from another. It looks like it stops the bleeding; it actually **breaks the data**:
The only safe remediations are:
1. **Drop an *entire* unwanted metric** (`action: drop` on `__name__`) — you're discarding the whole metric, not merging distinct series. 2. **Fix the source** — stop the application emitting the bad label (the real fix for unbounded `path`, `user_id`, etc.). 3. **Adaptive Metrics** — for structural cardinality on series you can't fix at the source. It aggregates *correctly* (counter-reset-aware, audited, reversible). This is the right way to reduce the cost of a label like `pod`. Route to `adaptive-metrics`.
Everywhere below that says "drop a label," read it through this rule: drop whole metrics, fix the source, or use Adaptive Metrics — never `labeldrop` a distinguishing label.
---
| Symptom | Likely Cause | First Action | |---|---|---| | Prometheus OOMKilled or memory growing linearly | Active series growth (often from a new bad metric or label) | [Active Series triage](#step-1-active-series-triage) | | Single PromQL query slow or OOMs the querier | One or more metrics in the query have high cardinality | [Per-query drill-down](#step-3-per-metric-drill-down) | | Remote write lagging, WAL growing | Sample throughput spike — series count OR scrape interval changed | [Active Series triage](#step-1-active-series-triage) + check scrape intervals | | `429 Too Many Samples` / `out of bounds` errors | Hitting Mimir/Cortex ingester per-tenant series limit | [Per-metric drill-down](#step-3-per-metric-drill-down), find the new offender | | Grafana Cloud Active Series bill spiked | New metric, new label, or rollout creating churn | [Per-metric drill-down](#step-3-per-metric-drill-down) + churn check | | Grafana Cloud DPM bill spiked but Active Series flat | Scrape interval shortened, OR remote_write sending duplicates | DPM-side issue — route to `dpm-finder` | | `series_limit_per_user` errors after a deploy | Application change introduced a new bad label | [Recent change diff](#step-4-recent-change-diff) | | Series count grows then resets every restart | Series churn from ephemeral label values | [Churn diagnosis](#step-5-churn-diagnosis) |
---
# Total active series in the local Prometheus
prometheus_tsdb_head_series
# Or for Mimir / Grafana Cloud Metrics (per tenant)
cortex_ingester_memory_series{user="<tenant>"}Compare to recent history:
# Growth over the last 7 days deriv(prometheus_tsdb_head_series[7d]) * 86400
A growth rate > a few % per day on a stable application set is a red flag.
Prometheus exposes a built-in cardinality breakdown:
curl -s http://prometheus:9090/api/v1/status/tsdb | jq
Returns:
This is usually the fastest path to "which metric / which label is the problem."
For Grafana Cloud:
# Same endpoint, authenticated against the per-tenant Mimir curl -s -u "<user>:<token>" \ "https://prometheus-prod-XX.grafana.net/api/prom/api/v1/status/tsdb" | jq
---
"seriesCountByMetricName": [
{ "name": "http_request_duration_seconds_bucket", "value": 184320 },
{ "name": "go_gc_duration_seconds", "value": 80 },
...
]**Heuristics**:
Public skills for working with Grafana, Prometheus, Loki, Tempo, Pyroscope, k6, and the broader LGTM observability stack. Compatible with Claude Code, Cursor, Codex, and any tool supporting the Agent Skills open standard.
Repo: grafana/skills
Use when the user asks to "write a validator", "add validation", "implement admission…
Use when starting any grafana-app-sdk work — scaffolding a Grafana app, initializing a…
Author CUE kind definitions for grafana-app-sdk apps - schemas, versioning, field…
Implement reconcilers and watchers for grafana-app-sdk apps — write…
Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics…