/slo-manage
Creates, updates, syncs, and deletes Grafana SLO definitions via gcx with dry-run validation and GitOps pull/push workflows. Use when the user wants to create, update, pull, push, or delete SLO definitions. Trigger on phrases like "create an SLO", "update SLO objective", "push
$ npx -y skills add grafana/gcx --skill slo-manage --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/slo-manage
Context preview
The summary Claude sees to decide when to auto-load this skill.
Creates, updates, syncs, and deletes Grafana SLO definitions via gcx with dry-run validation and GitOps pull/push workflows. Use when the user wants to create, update, pull, push, or delete SLO definitions. Trigger on phrases like "create an SLO", "update SLO objective", "push
SKILL.md
slo-manage.SKILL.mdname: slo-manage
description: Creates, updates, syncs, and deletes Grafana SLO definitions via gcx with dry-run validation and GitOps pull/push workflows. Use when the user wants to create, update, pull, push, or delete SLO definitions. Trigger on phrases like "create an SLO", "update SLO objective", "push SLO", "pull SLOs", "delete SLO", or "GitOps sync SLOs". For checking SLO health or status, use slo-check-status instead. For investigating a breaching SLO, use slo-investigate instead.
allowed-tools: Bash, Read, Write, Edit
SLO Management
Create, update, sync, and delete SLO definitions using gcx.
Core Principles
1. Use gcx commands exclusively — do not call Grafana APIs directly 2. Always run `--dry-run` before any push operation; proceed only if dry-run succeeds 3. Trust the user's expertise — skip explanations of SLO concepts 4. Use `-o json` for agent processing; default table/yaml for user display 5. Auto-resolve datasource UIDs; only ask if auto-discovery fails
Query Type Decision Table
Select query type based on what the user describes:
| User describes | Query type | |----------------|------------| | "percentage of successful requests", "success rate", "error rate" | `ratio` | | "raw PromQL expression", "custom metric formula" | `freeform` | | "metric above/below threshold", "latency under X ms", "availability percentage" | `threshold` |
Metric-Pattern Decision Table
Use the metric name suffix to pick the query type when the user provides a metric name:
| Metric suffix / type | Query type | Rationale | |----------------------|------------|-----------| | `_total` counter | `ratio` | success_total / all_total | | `_bucket` histogram | `threshold` | use le-bound threshold on quantile | | `_gauge` or `up` metric | `threshold` | compare to fixed threshold | | None of the above | `freeform` | last resort only |
**Guardrail:** Freeform is a last resort. Before choosing freeform, verify the SLI cannot be expressed as ratio or threshold.
**Hard requirement:** Freeform queries MUST use `$__rate_interval` in all `rate()`/`increase()` calls. Literal ranges like `[5m]` are rejected by the SLO API.
Workflow 1: Create New SLO
Step 1: Determine query type using the decision table above
Step 2: Resolve destination datasource UID
gcx datasources list --type prometheus
Use the UID from the output. Some stacks do not populate `type` in the list response, so the filter can come back empty even though Prometheus datasources exist — in that case list everything and pick the canonical Grafana Cloud Prometheus entry (name like `grafanacloud-<stack>-prom`):
gcx datasources list
If that leaves zero or multiple plausible Prometheus candidates, ask the user which UID to use rather than guessing (the stack's `default: true` datasource is not necessarily Prometheus).
Step 3: Build YAML from the appropriate template
For complete ratio, freeform, and threshold templates (including alerting, labels, and folder fields), see [references/slo-templates.md](references/slo-templates.md). Key structure:
apiVersion: slo.ext.grafana.app/v1alpha1
kind: SLO
metadata:
name: "" # leave empty for new SLO (server assigns UUID on create)
spec:
name: "my-api-availability"
description: "API availability over 28 days"
query:
type: ratio # freeform | ratio | threshold
ratio: # field matches type
successMetric:
prometheusMetric: http_requests_total{status!~"5.."}
totalMetric:
prometheusMetric: http_requests_total
groupByLabels: [cluster, service]
objectives:
- value: 0.999 # 0.9 to 0.9999 typical range
window: 28d # 7d | 14d | 28d | 30d
destinationDatasource:
uid: <prometheus-uid>Step 4: Validate with dry-run, then push
gcx slo definitions push slo.yaml --dry-run
gcx slo definitions push slo.yaml
**Push semantics:**
- `metadata.name` empty → always creates (server assigns UUID)
- `metadata.name` set to UUID → upsert (updates if exists, creates if not)
After creation, server assigns UUID. Run `gcx slo definitions list` to confirm.
Workflow 2: Update Existing SLO
Step 1: Get current definition
gcx slo definitions get <UUID> -o yaml > slo.yaml
Step 2: Modify the YAML file
Edit the relevant fields (objective value, query, alerting, etc.). Do not modify `metadata.name` (UUID) or `readOnly` fields.
Step 3: Dry-run, then push
gcx slo definitions push slo.yaml --dry-run
gcx slo definitions push slo.yaml
Workflow 3: GitOps Sync (Pull/Push)
Pull all SLOs to disk
gcx slo definitions pull -d ./slos
# Writes to ./slos/SLO/<uuid>.yaml
Push directory of SLOs
gcx slo definitions push ./slos/SLO/*.yaml --dry-run
gcx slo definitions push ./slos/SLO/*.yaml
Workflow 4: Delete SLO
Step 1: Confirm SLO identity
gcx slo definitions list
gcx slo definitions get <UUID>
Confirm the UUID and name with the user before deletion.
Step 2: Delete
gcx slo definitions delete <UUID> --force
Use `--force` to skip the confirmation prompt when running in agent mode (there is no `-f` shorthand on delete).
Configuration Guidance
**Objective values** (stored as 0–1, displayed as percentage):
- Typical range: 0.9 (90%) to 0.9999 (99.99%)
- Common starting points: 0.99 (99%), 0.999 (99.9%), 0.9999 (99.99%)
**Window options:** `7d`, `14d`, `28d`, `30d`
- 28d is most common; matches many SLO frameworks
- Shorter windows (7d) react faster but have higher variance
**Alerting best practices:**
- `fastBurn`: Pages on-call (high burn rate, short window — catches rapid budget consumption)
- `slowBurn`: Creates tickets (low burn rate, long window — catches gradual degradation)
**Labels:** Use consistent label keys (team, service, environment, tier) for filtering and grouping.
**GroupByLabels** (ratio/threshold queries): Add labels like `cluster`, `ser
Read more
name: slo-manage description: Creates, updates, syncs, and deletes Grafana SLO definitions via gcx with dry-run validation and GitOps pull/push workflows. Use when the user wants to create, update, pull, push, or delete SLO definitions. Trigger on phrases like "create an SLO", "update SLO objective", "push SLO", "pull SLOs", "delete SLO", or "GitOps sync SLOs". For checking SLO health or status, use slo-check-status instead. For investigating a breaching SLO, use slo-investigate instead. allowed-tools: Bash, Read, Write, Edit
SLO Management
Create, update, sync, and delete SLO definitions using gcx.
Core Principles
1. Use gcx commands exclusively — do not call Grafana APIs directly 2. Always run `--dry-run` before any push operation; proceed only if dry-run succeeds 3. Trust the user's expertise — skip explanations of SLO concepts 4. Use `-o json` for agent processing; default table/yaml for user display 5. Auto-resolve datasource UIDs; only ask if auto-discovery fails
Query Type Decision Table
Select query type based on what the user describes:
| User describes | Query type | |----------------|------------| | "percentage of successful requests", "success rate", "error rate" | `ratio` | | "raw PromQL expression", "custom metric formula" | `freeform` | | "metric above/below threshold", "latency under X ms", "availability percentage" | `threshold` |
Metric-Pattern Decision Table
Use the metric name suffix to pick the query type when the user provides a metric name:
| Metric suffix / type | Query type | Rationale | |----------------------|------------|-----------| | `_total` counter | `ratio` | success_total / all_total | | `_bucket` histogram | `threshold` | use le-bound threshold on quantile | | `_gauge` or `up` metric | `threshold` | compare to fixed threshold | | None of the above | `freeform` | last resort only |
**Guardrail:** Freeform is a last resort. Before choosing freeform, verify the SLI cannot be expressed as ratio or threshold.
**Hard requirement:** Freeform queries MUST use `$__rate_interval` in all `rate()`/`increase()` calls. Literal ranges like `[5m]` are rejected by the SLO API.
Workflow 1: Create New SLO
Step 1: Determine query type using the decision table above
Step 2: Resolve destination datasource UID
gcx datasources list --type prometheus
Use the UID from the output. Some stacks do not populate `type` in the list response, so the filter can come back empty even though Prometheus datasources exist — in that case list everything and pick the canonical Grafana Cloud Prometheus entry (name like `grafanacloud-<stack>-prom`):
gcx datasources list
If that leaves zero or multiple plausible Prometheus candidates, ask the user which UID to use rather than guessing (the stack's `default: true` datasource is not necessarily Prometheus).
Step 3: Build YAML from the appropriate template
For complete ratio, freeform, and threshold templates (including alerting, labels, and folder fields), see [references/slo-templates.md](references/slo-templates.md). Key structure:
apiVersion: slo.ext.grafana.app/v1alpha1
kind: SLO
metadata:
name: "" # leave empty for new SLO (server assigns UUID on create)
spec:
name: "my-api-availability"
description: "API availability over 28 days"
query:
type: ratio # freeform | ratio | threshold
ratio: # field matches type
successMetric:
prometheusMetric: http_requests_total{status!~"5.."}
totalMetric:
prometheusMetric: http_requests_total
groupByLabels: [cluster, service]
objectives:
- value: 0.999 # 0.9 to 0.9999 typical range
window: 28d # 7d | 14d | 28d | 30d
destinationDatasource:
uid: <prometheus-uid>Step 4: Validate with dry-run, then push
gcx slo definitions push slo.yaml --dry-run gcx slo definitions push slo.yaml
**Push semantics:**
- `metadata.name` empty → always creates (server assigns UUID)
- `metadata.name` set to UUID → upsert (updates if exists, creates if not)
After creation, server assigns UUID. Run `gcx slo definitions list` to confirm.
Workflow 2: Update Existing SLO
Step 1: Get current definition
gcx slo definitions get <UUID> -o yaml > slo.yaml
Step 2: Modify the YAML file
Edit the relevant fields (objective value, query, alerting, etc.). Do not modify `metadata.name` (UUID) or `readOnly` fields.
Step 3: Dry-run, then push
gcx slo definitions push slo.yaml --dry-run gcx slo definitions push slo.yaml
Workflow 3: GitOps Sync (Pull/Push)
Pull all SLOs to disk
gcx slo definitions pull -d ./slos # Writes to ./slos/SLO/<uuid>.yaml
Push directory of SLOs
gcx slo definitions push ./slos/SLO/*.yaml --dry-run gcx slo definitions push ./slos/SLO/*.yaml
Workflow 4: Delete SLO
Step 1: Confirm SLO identity
gcx slo definitions list gcx slo definitions get <UUID>
Confirm the UUID and name with the user before deletion.
Step 2: Delete
gcx slo definitions delete <UUID> --force
Use `--force` to skip the confirmation prompt when running in agent mode (there is no `-f` shorthand on delete).
Configuration Guidance
**Objective values** (stored as 0–1, displayed as percentage):
- Typical range: 0.9 (90%) to 0.9999 (99.99%)
- Common starting points: 0.99 (99%), 0.999 (99.9%), 0.9999 (99.99%)
**Window options:** `7d`, `14d`, `28d`, `30d`
- 28d is most common; matches many SLO frameworks
- Shorter windows (7d) react faster but have higher variance
**Alerting best practices:**
- `fastBurn`: Pages on-call (high burn rate, short window — catches rapid budget consumption)
- `slowBurn`: Creates tickets (low burn rate, long window — catches gradual degradation)
**Labels:** Use consistent label keys (team, service, environment, tier) for filtering and grouping.
**GroupByLabels** (ratio/threshold queries): Add labels like `cluster`, `ser
Grafana — in your terminal and your agentic coding environment. gcx works with Grafana Cloud, Enterprise, and OSS (Grafana 12+). See the compatibility matrix for details. Query production. Investigate alerts. Let the Assistant root-cause issues.
Repo: grafana/gcx
Other skills on gcx.
- /add-datasource
Use when adding a new datasource type to gcx (e.g., Elasticsearch, CloudWatch, InfluxDB), or when the user says "add datasource", "new datasource type", or "integrate [datasource]".
Open skill - /add-provider
Use when adding a new Grafana Cloud product provider to gcx (SLO, OnCall, Synthetic Monitoring, k6, ML, etc.), or when the user says "add provider", "new provider", or "integrate [product]".
Open skill - /generate-slide
Regenerate the gcx marketing bento-box slide (slide.html) with verified commands from the current codebase. Builds a fresh binary and reflects against the actual command tree. Use when the user says "regenerate slide", "update slide", "generate slide", or "/generate-slide".
Open skill - /migrate-provider
Use when porting a Grafana Cloud product from grafana-cloud-cli (gcx) to gcx, when a bead task references gcx provider migration, or when user says "migrate provider", "port from gcx", "port oncall", "port k6". Not for building providers from scratch — use /add-provider for that.
Open skill - /release
Tag and release a new gcx version. Use when the user wants to cut a release, tag a version, run the release process, or says "release patch/minor/major".
Open skill - /agento11y-instrument
Sets up and instruments a developer's own LLM app or agent to send generations and agentic workflow to Grafana Agent Observability (the Agent Observability SDKs) — greenfield setup, fixing broken instrumentation, or filling gaps in existing instrumentation. Uses gcx for the
Open skill

