ai-output-validation
Validates, parses, and sanitizes AI-generated outputs before they reach end users or downstream systems. Structured output enforcement, schema validation, and…
Structured logging, distributed tracing, and alerting for AI systems and traditional services. You can't fix what you can't see.
$ npx -y skills add DevelopersGlobal/ai-agent-skills --skill observability --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/observabilityContext preview
The summary Claude sees to decide when to auto-load this skill.
Structured logging, distributed tracing, and alerting for AI systems and traditional services. You can't fix what you can't see.
name: observability description: Structured logging, distributed tracing, and alerting for AI systems and traditional services. You can't fix what you can't see. category: harden applies-to: [claude, gemini, cursor, copilot, any] version: 1.0.0
Observability is the ability to understand the internal state of a system from its external outputs. For AI systems this is especially critical: agents make decisions that are hard to interpret without detailed telemetry.
The three pillars: **Logs** (what happened), **Traces** (how long and where), **Metrics** (aggregate health).
1. All logs must be **structured** (JSON, not free text). Fields: `timestamp`, `level`, `service`, `traceId`, `message`, `context`. 2. Log levels used correctly:
3. **Never log secrets, PII, or auth tokens.** 4. For AI systems, log: prompt inputs (sanitized), model outputs, token counts, latency, model version.
**Verify:** Logs are structured JSON. No secrets in logs. AI interactions logged.
5. Every request gets a unique `traceId` generated at the entry point. 6. `traceId` is propagated through all downstream calls (HTTP headers, message queues, agent calls). 7. Each service/agent creates a **span** for its work, with: start time, end time, parent span ID. 8. Use OpenTelemetry as the standard instrumentation library.
**Verify:** You can trace a single request across all services/agents in a single view.
9. Define and track key metrics:
10. Dashboards: one dashboard per service with RED metrics, one dashboard for AI system health.
**Verify:** RED metrics are tracked for every service. AI-specific metrics tracked for AI systems.
11. Alerts must be **actionable** — every alert should have a runbook. 12. Alert on symptoms (high error rate, high latency), not just causes. 13. AI-specific alerts: token budget exceeded, model error rate spike, retrieval failure rate spike. 14. On-call rotation: someone is responsible for every alert at all times.
**Verify:** Every alert has a runbook. On-call rotation defined.
| Excuse | Rebuttal | |--------|----------| | "We'll add monitoring after launch" | You'll be fighting fires blind. Add it before. | | "Console.log is enough" | In production, console.log is noise. Structured logs with context are signals. | | "The AI model handles it internally" | Model internals are a black box. You must observe the inputs and outputs. |
AI agent skills for production grade applications
Validates, parses, and sanitizes AI-generated outputs before they reach end users or downstream systems. Structured output enforcement, schema validation, and…
Design stable, versioned, self-documenting APIs. Easy to use correctly, hard to use incorrectly. Apply Hyrum's Law from day one.
Automated quality gates from commit to production. Every merge to main is potentially shippable. No manual steps in the deployment path.
Get layered, context-aware explanations of unfamiliar code. Understand what it does, why it was written that way, and how to work with it safely.
Structured code review focusing on correctness, security, and maintainability. Correctness before style. Every reviewer comment must be actionable.
Load minimum necessary context into agent context windows. Prevents token bloat, reduces cost, and improves focus. Only load what the current task needs.