Skip to content

/monitoring-observability

Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (as_type, score_current_span, should_export_span, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality

shell
$ npx -y skills add yonatangross/orchestkit --skill monitoring-observability --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/monitoring-observability
How auto-invocation works

Context preview

The summary Claude sees to decide when to auto-load this skill.

Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (as_type, score_current_span, should_export_span, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality

SKILL.md

monitoring-observability.SKILL.md
name: monitoring-observability
license: MIT
compatibility: "Claude Code 2.1.220+."
description: Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (as_type, score_current_span, should_export_span, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.
tags: [monitoring, observability, prometheus, grafana, langfuse, tracing, metrics, drift-detection, logging]
context: fork
version: 3.0.0
author: OrchestKit
user-invocable: false
disable-model-invocation: true
complexity: medium
persuasion-type: reference
targets:
  # Python SDK. The JS/TS SDK is a separate line at 5.x (@langfuse/* packages) — see
  # references/langfuse-js-v5.md. The self-hosted platform is a third axis (v3) and is not an
  # SDK version.
  - library: langfuse
    version: ">=4.0.0"
upstream-version-tested: "4.14.2"
metadata:
  category: document-asset-creation
allowed-tools:
  - Read
  - Glob
  - Grep
  - WebFetch
  - WebSearch
path_patterns: ["**/metrics/**", "**/tracing/**", "prometheus.*", "grafana/**"]

Monitoring & Observability

A wrap around Prometheus, Grafana, OpenTelemetry and Langfuse, not a re-teaching of them. This skill carries OrchestKit's delta (version floors, house decisions, scars) and points at the vendor for everything else. Start at `references/ork-delta.md`.

Upstream coverage (do not restate)

These topics are fully covered first-party. Read the source, do not add a local copy.

| Topic | First-party source | |-------|--------------------| | Prometheus metric types, RED method, cardinality, PromQL | <https://prometheus.io/docs/practices/> | | Alertmanager grouping, inhibition, escalation, runbooks | <https://prometheus.io/docs/alerting/latest/configuration/> | | Grafana dashboards, Loki and LogQL, Promtail | <https://grafana.com/docs/> | | OpenTelemetry spans, sampling, context propagation | <https://opentelemetry.io/docs/> | | Langfuse Python SDK (`@observe`, `as_type`, `score_current_span`, `should_export_span`, `LangfuseMedia`) | <https://langfuse.com/docs/sdk/python> | | Langfuse v2 to v4 Python and v3 to v5 JS migration paths | <https://langfuse.com/docs/sdk/python/v4-migration> | | Langfuse self-hosting (ClickHouse, Redis, S3, Helm) | <https://langfuse.com/docs/deployment/self-host> | | Langfuse cost tracking, model pricing, Metrics API v2 | <https://langfuse.com/docs/model-usage-and-cost> | | Langfuse scores, online evaluators, annotation queues, prompt management | <https://langfuse.com/docs/scores/overview> | | Langfuse framework integrations (LangChain, LangGraph, CrewAI, Pydantic AI, Bedrock, LiveKit) | <https://langfuse.com/docs/integrations> | | Agent Graphs, observation types, rendered tool calls | <https://langfuse.com/docs/tracing-features/agent-graphs> | | PSI, KS test, KL and JS divergence, Wasserstein, embedding drift | <https://www.evidentlyai.com/blog/data-drift-detection-large-datasets> | | EWMA control charts | <https://www.itl.nist.gov/div898/handbook/pmc/section3/pmc324.htm> | | structlog, Winston, correlation IDs, log sampling | <https://www.structlog.org/en/stable/> |

Quick Reference

| Category | Rules | Impact | When to Use | |----------|-------|--------|-------------| | [Infrastructure Monitoring](#infrastructure-monitoring) | 1 | CRITICAL | Grafana dashboards, Golden Signals, SLO/SLI | | [LLM Observability](#llm-observability) | 1 | HIGH | Langfuse tracing, observation types, agent graphs | | [Silent Failures](#silent-failures) | 3 | HIGH | Tool skipping, quality degradation, loop/token spike alerting |

**Total: 5 rules across 3 categories.** Drift detection, cost tracking, eval scoring, Prometheus instrumentation and alert-rule authoring moved to the upstream sources listed above.

Quick Start

# Langfuse v4 LLM tracing: semantic as_type plus inline scoring
from langfuse import observe, get_client

@observe(as_type="generation", name="analyze_content")
async def analyze_content(content: str):
    get_client().update_current_trace(
        user_id="user_123", session_id="session_abc",
        tags=["production", "orchestkit"],
    )
    result = await llm.generate(content)
    get_client().score_current_span(name="response_quality", value=0.85)
    return result
# Prometheus RED method, wired the way this repo expects (bounded labels only)
from prometheus_client import Counter, Histogram

http_requests = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
http_duration = Histogram('http_request_duration_seconds', 'Request latency',
    buckets=[0.01, 0.05, 0.1, 0.5, 1, 2, 5])

Infrastructure Monitoring

Dashboard and health-check patterns. Metric instrumentation and alert-rule syntax are upstream.

| Rule | File | Key Pattern | |------|------|-------------| | Grafana Dashboards | `rules/monitoring-grafana.md` | Golden Signals, SLO/SLI, health checks |

> **CC 2.1.161 — OTEL resource attributes as metric labels:** `OTEL_RESOURCE_ATTRIBUTES` values are now attached as labels on metric datapoints, so usage metrics can be sliced by custom dimensions (team, repo, environment). Add label selectors to dashboards for multi-tenant / per-team cost and usage tracking.

LLM Observability

Langfuse-based tracing for LLM applications. Cost tracking, scoring and drift statistics are upstream; what stays here is how this repo wires traces.

| Rule | File | Key Pattern | |------|------|-------------| | Langfuse Traces | `rules/llm-langfuse-traces.md` | @observe decorator, OTEL spans, agent graphs |

Silent Failures

Detection and alerting for silent failures in LLM agents.

| Rule | File | Key Pattern | |------|------|-------------| | Tool Skipping | `rules/silent-tool-skipping.md` | Expected vs actual tool calls, Langfuse traces | | Quality Degradation | `rules/silent-degraded-quality.md` | Heuristics + LLM-as-judge, z-score baselines | | Silent Alerting | `

Read more
Read it on GitHub ↗

Showing the first part of this file.

Ships withorchestkit

The Complete AI Development Toolkit for Claude Code — 114 skills, 37 agents, 212 hooks. Production-ready patterns for full-stack development.

Get the whole plugin, auto-invoked
Stats
212
Stars
0
Views
22
Forks
Active
Maintenance
TypeScript
Language
MIT
License
31m ago
Last commit
7mo ago
Created

Repo: yonatangross/orchestkit

Other skills on orchestkit.