/ml-ai
Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN outlier detection), Sift (8-analysis automated root-cause), Knowledge Graph + RCA
$ npx -y skills add grafana/skills --skill ml-ai --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/ml-ai
Context preview
The summary Claude sees to decide when to auto-load this skill.
Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN outlier detection), Sift (8-analysis automated root-cause), Knowledge Graph + RCA
SKILL.md
ml-ai.SKILL.mdname: ml-ai
license: Apache-2.0
description: Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN outlier detection), Sift (8-analysis automated root-cause), Knowledge Graph + RCA Workbench, and the LLM Plugin (OpenAI / Anthropic / Azure / Ollama / vLLM / LiteLLM). Use when you want anomaly alerts without static thresholds, natural-language querying, automated incident investigation, dashboards generated from a sentence, or a managed LLM proxy for plugins — even when the user says "alert when something looks weird", "explain this PromQL", "find the root cause", "make this a dashboard", or "wire Claude into Grafana" without naming any of these products.
Grafana Cloud AI & ML
> **Docs**: https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/
ML alerting + automated RCA + LLM-powered Assistant in one Grafana Cloud stack.
Prerequisites
- Grafana Cloud stack (Pro / Advanced — most features GA, some in preview)
- API token with `plugins:write` for ML / Sift / LLM-plugin endpoints
- For Dynamic Alerting: at least 14 days (ideally 90d) of history for the metric you want to forecast
Common Workflows
1. Forecasting alert with Dynamic Alerting
# 1. Create forecast job (Prophet — learns daily/weekly seasonality)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/forecast \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{
"name": "cpu-forecast",
"metric": "avg(rate(node_cpu_seconds_total{mode=\"user\"}[5m]))",
"datasourceId": 1,
"interval": 300,
"trainingWindow": "90d",
"forecastWindow": "7d",
"algorithm": { "name": "prophet", "config": {} }
}'
# 2. Verify job is producing the predicted-value metric (may take a few minutes).
# <datasourceId> must match the datasourceId used above (find it via
# GET /api/datasources), or run the query from Explore instead.
curl -s -H "Authorization: Bearer <token>" \
'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_forecast_upper{job="cpu-forecast"}' \
| jq '.data.result | length'
# Expect > 0
# 3. Add an alert that fires when actual exceeds the upper bound
# expr: avg(rate(node_cpu_seconds_total{mode="user"}[5m]))
# > ml_forecast_upper{job="cpu-forecast"} * 1.12. Outlier alert — one service deviates from peers
# 1. Create outlier job (DBSCAN — groups peers, flags the odd one)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/outlier \
-H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
-d '{
"name": "service-error-outliers",
"metric": "sum(rate(http_requests_total{status=~\"5..\"}[5m])) by (service)",
"datasourceId": 1,
"interval": 300,
"algorithm": { "name": "dbscan", "sensitivity": 0.5, "config": { "epsilon": 0.5 } }
}'
# 2. Verify the score metric exists (<datasourceId> must match the
# datasourceId used above, or run the query from Explore instead)
curl -s -H "Authorization: Bearer <token>" \
'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_outlier_score{job="service-error-outliers"}' \
| jq '.data.result | length'
# 3. Alert when ml_outlier_score{job="service-error-outliers"} > 0.8 for 5m3. Run a Sift investigation
# 1. Trigger from API (or from Explore / Incident / OnCall)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-sift-app/resources/sift/v1/investigations \
-H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
-d '{ "name":"checkout-spike","start":"2024-02-01T10:00:00Z","end":"2024-02-01T10:30:00Z",
"filters":{"service":"checkout","namespace":"production"} }'
# 2. The response includes an investigation ID — open it in the UI:
# https://<stack>.grafana.net/a/grafana-sift-app/investigations/<id>
# 3. Verify analyses ran — each of the 8 checks shows ✔ or ✖ with linked evidence.See [`references/sift.md`](references/sift.md) for the full 8-analysis table.
4. Wire up the LLM Plugin
# 1. Provision (provisioning/plugins/llm.yaml — see references/llm-and-graph.md)
apiVersion: 1
apps:
- type: grafana-llm-app
jsonData: { openAIUrl: https://api.openai.com, openAIModel: gpt-4o }
secureJsonData: { openAIKey: sk-... }# 2. Restart Grafana, then verify the health endpoint reports the configured provider
curl -s -H "Authorization: Bearer <token>" \
https://<stack>.grafana.net/api/plugins/grafana-llm-app/health | jq
# Expect: {"status":"ok", ...}
# 3. Verify in a panel — open any panel, click the Assistant icon, ask "what does this query do?"See [`references/llm-and-graph.md`](references/llm-and-graph.md) for Assistant capabilities, Knowledge Graph search syntax, and Adaptive Metrics recommendations.
Resources
- [Machine Learning docs](https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/)
- [Sift](https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/sift/)
- [Grafana Assistant](https://grafana.com/docs/grafana-cloud/visualizations/grafana-assistant/)
- [LLM Plugin](https://grafana.com/grafana/plugins/grafana-llm-app/)
Read more
name: ml-ai license: Apache-2.0 description: Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN outlier detection), Sift (8-analysis automated root-cause), Knowledge Graph + RCA Workbench, and the LLM Plugin (OpenAI / Anthropic / Azure / Ollama / vLLM / LiteLLM). Use when you want anomaly alerts without static thresholds, natural-language querying, automated incident investigation, dashboards generated from a sentence, or a managed LLM proxy for plugins — even when the user says "alert when something looks weird", "explain this PromQL", "find the root cause", "make this a dashboard", or "wire Claude into Grafana" without naming any of these products.
Grafana Cloud AI & ML
> **Docs**: https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/
ML alerting + automated RCA + LLM-powered Assistant in one Grafana Cloud stack.
Prerequisites
- Grafana Cloud stack (Pro / Advanced — most features GA, some in preview)
- API token with `plugins:write` for ML / Sift / LLM-plugin endpoints
- For Dynamic Alerting: at least 14 days (ideally 90d) of history for the metric you want to forecast
Common Workflows
1. Forecasting alert with Dynamic Alerting
# 1. Create forecast job (Prophet — learns daily/weekly seasonality)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/forecast \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{
"name": "cpu-forecast",
"metric": "avg(rate(node_cpu_seconds_total{mode=\"user\"}[5m]))",
"datasourceId": 1,
"interval": 300,
"trainingWindow": "90d",
"forecastWindow": "7d",
"algorithm": { "name": "prophet", "config": {} }
}'
# 2. Verify job is producing the predicted-value metric (may take a few minutes).
# <datasourceId> must match the datasourceId used above (find it via
# GET /api/datasources), or run the query from Explore instead.
curl -s -H "Authorization: Bearer <token>" \
'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_forecast_upper{job="cpu-forecast"}' \
| jq '.data.result | length'
# Expect > 0
# 3. Add an alert that fires when actual exceeds the upper bound
# expr: avg(rate(node_cpu_seconds_total{mode="user"}[5m]))
# > ml_forecast_upper{job="cpu-forecast"} * 1.12. Outlier alert — one service deviates from peers
# 1. Create outlier job (DBSCAN — groups peers, flags the odd one)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/outlier \
-H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
-d '{
"name": "service-error-outliers",
"metric": "sum(rate(http_requests_total{status=~\"5..\"}[5m])) by (service)",
"datasourceId": 1,
"interval": 300,
"algorithm": { "name": "dbscan", "sensitivity": 0.5, "config": { "epsilon": 0.5 } }
}'
# 2. Verify the score metric exists (<datasourceId> must match the
# datasourceId used above, or run the query from Explore instead)
curl -s -H "Authorization: Bearer <token>" \
'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_outlier_score{job="service-error-outliers"}' \
| jq '.data.result | length'
# 3. Alert when ml_outlier_score{job="service-error-outliers"} > 0.8 for 5m3. Run a Sift investigation
# 1. Trigger from API (or from Explore / Incident / OnCall)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-sift-app/resources/sift/v1/investigations \
-H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
-d '{ "name":"checkout-spike","start":"2024-02-01T10:00:00Z","end":"2024-02-01T10:30:00Z",
"filters":{"service":"checkout","namespace":"production"} }'
# 2. The response includes an investigation ID — open it in the UI:
# https://<stack>.grafana.net/a/grafana-sift-app/investigations/<id>
# 3. Verify analyses ran — each of the 8 checks shows ✔ or ✖ with linked evidence.See [`references/sift.md`](references/sift.md) for the full 8-analysis table.
4. Wire up the LLM Plugin
# 1. Provision (provisioning/plugins/llm.yaml — see references/llm-and-graph.md)
apiVersion: 1
apps:
- type: grafana-llm-app
jsonData: { openAIUrl: https://api.openai.com, openAIModel: gpt-4o }
secureJsonData: { openAIKey: sk-... }# 2. Restart Grafana, then verify the health endpoint reports the configured provider
curl -s -H "Authorization: Bearer <token>" \
https://<stack>.grafana.net/api/plugins/grafana-llm-app/health | jq
# Expect: {"status":"ok", ...}
# 3. Verify in a panel — open any panel, click the Assistant icon, ask "what does this query do?"See [`references/llm-and-graph.md`](references/llm-and-graph.md) for Assistant capabilities, Knowledge Graph search syntax, and Adaptive Metrics recommendations.
Resources
- [Machine Learning docs](https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/)
- [Sift](https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/sift/)
- [Grafana Assistant](https://grafana.com/docs/grafana-cloud/visualizations/grafana-assistant/)
- [LLM Plugin](https://grafana.com/grafana/plugins/grafana-llm-app/)
Public skills for working with Grafana, Prometheus, Loki, Tempo, Pyroscope, k6, and the broader LGTM observability stack. Compatible with Claude Code, Cursor, Codex, and any tool supporting the Agent Skills open standard.
Repo: grafana/skills
Other skills on grafana-skills.
- /admission-control
Use when the user asks to "write a validator", "add validation", "implement admission control", "write a mutating webhook", "add a mutation handler", "validate incoming resources", "implement admission logic", "add admission webhooks", "write ingress validation", or asks how to
Open skill - /app-sdk-concepts
Use when starting any grafana-app-sdk work — scaffolding a Grafana app, initializing a Grafana App Platform app, picking a deployment mode (standalone operator / grafana/apps / frontend-only), wiring app-specific config, or onboarding to the SDK. Covers `grafana-app-sdk` CLI
Open skill - /cue-kind-definition
Author CUE kind definitions for grafana-app-sdk apps - schemas, versioning, field constraints, named type definitions, custom routes, and codegen configuration. Scaffolds kinds via `grafana-app-sdk project kind add`, writes spec/status schemas with type constraints (regex, enum,
Open skill - /reconciler-logic
Implement reconcilers and watchers for grafana-app-sdk apps — write `TypedReconciler[*MyKind]` reconcile functions, apply generation-based skip patterns, do conflict-safe status updates via `resource.UpdateObject`, configure `BasicReconcileOptions` (namespace, label/field
Open skill - /adaptive-metrics
Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config, unused-metric detection, and Alloy remote_write fallback. Use when investigating a high
Open skill - /admin
Manage Grafana Cloud accounts — organizations, stacks, RBAC roles and assignments, SSO/SAML/OAuth/GitHub auth, service accounts for CI/CD, user invites, team membership, and API-driven provisioning. Creates stacks via the Cloud API, mints service-account tokens, applies role
Open skill

