Skip to content
AI & Agents
Skill

/observability

Add agent-first observability — structured logs, health endpoints, failure-state persistence, explicit failure modes — so the next agent can diagnose problems unattended. Use when asked to "add logging", "add observability", "add metrics", "make this observable", or when

BOOST
From plugin
gsd-pi
1.3k37 skills13 agents
Install
$ npx -y skills add open-gsd/gsd-pi --skill observability --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/observability

Context preview

The summary Claude sees to decide when to auto-load this skill.

Add agent-first observability — structured logs, health endpoints, failure-state persistence, explicit failure modes — so the next agent can diagnose problems unattended. Use when asked to "add logging", "add observability", "add metrics", "make this observable", or when

SKILL.md

observability.SKILL.md
name: observability
description: Add agent-first observability — structured logs, health endpoints, failure-state persistence, explicit failure modes — so the next agent can diagnose problems unattended. Use when asked to "add logging", "add observability", "add metrics", "make this observable", or when building/refactoring a subsystem that runs unattended (auto-mode engine, background jobs, servers, watchers).

<objective> Instrument code so that a cold-start agent can understand what happened by reading signals, not by rerunning with extra logging. The deliverable is a set of specific instrumentation additions: structured logs at decision points, health/status surfaces for long-running processes, persisted failure state, and explicit failure modes that don't get swallowed. </objective>

<context> gsd-pi's `VISION.md` lists "agent-first observability" as a principle, and the system prompt calls it out: "A future version of you will land in this codebase with no memory… you add observability because you're the one who'll need it at 3am." gsd-pi already exemplifies this — `activity/*.jsonl`, `journal/*.jsonl`, `metrics.json`, `doctor-history.jsonl` — but new code doesn't get that treatment automatically.

This skill is the thinking process for adding it. Not "add logs everywhere" — add the *right* signals at the *right* decision points.

Invocation points:

  • Building auto-mode-style code (loops, dispatch, guards, retries)
  • Adding a background job, watcher, or scheduled task
  • Writing a server or long-running process
  • Refactoring a subsystem that has been hard to debug
  • Addressing a production bug where "we had no visibility" surfaced

</context>

<core_principle> **LOG DECISIONS, NOT ACTIVITY.** "Entering function X" is noise. "Dispatched unit `slice/S02` after guard check passed because `status=pending`" is signal. Every log line should answer a question a future debugger will ask.

**FAIL LOUDLY AND PERSIST THE REASON.** Silent `try/catch` that returns `undefined` is an anti-pattern. If something fails, the failure state needs to be somewhere a fresh agent can find it — a JSONL, a status file, a health endpoint.

**OBSERVABILITY IS NOT FREE.** Every log allocation, every metric, every health check costs CPU and disk. Add only what you would actually read. </core_principle>

<process>

Step 1: Map the failure modes

Before instrumenting, list what can go wrong:

1. **What inputs could be invalid?** External API responses, user-submitted data, filesystem state, env vars. 2. **What external dependencies could fail?** Network, DB, child processes, filesystem permissions. 3. **What internal invariants could break?** State transitions, lock acquisition, concurrency assumptions. 4. **What silent corruption is possible?** Truncated writes, partial transactions, stale caches.

This map tells you where to instrument. Don't instrument uniformly — instrument at the decision points where these failures would manifest.

Step 2: Structured logs at decision points

For each decision the code makes that could plausibly go wrong later:

  • **What decision?** ("Dispatching unit X", "Retrying with backoff", "Skipping validation because flag set")
  • **Why?** ("status=pending", "previous attempt exit=1", "--dev flag set")
  • **What would a future debugger want to know?** (The values that drove the choice.)

Format:

log.info({
  event: "unit-dispatched",
  unitType: "slice",
  unitId: "S02",
  reason: "pending",
  attempt: 1,
  flowId,
});

Use the project's existing logger if one exists. In gsd-pi, follow the patterns in `src/resources/extensions/gsd/activity-log.ts` and `src/resources/extensions/gsd/journal.ts` — structured JSONL, one event per line, with `ts`, `event`, and domain-specific fields.

Avoid:

  • `console.log("here")` — what does "here" mean in six months?
  • Logging secrets, tokens, or PII — ever.
  • Formatting structured data into a prose string — it can't be grepped or filtered.

Step 3: Persist failure state

When something fails in a way the caller can't immediately handle, write the failure state to disk:

await writeAtomically(
  resolve(".gsd/runtime/last-error.json"),
  JSON.stringify({
    ts: new Date().toISOString(),
    phase: "execute",
    unitId,
    error: { message, stack, code },
    retryCount,
  })
);

A fresh agent reading `.gsd/runtime/` sees what happened last, what was retried, and where the process stopped. Pattern exists already in gsd-pi — reuse the `atomic-write.ts` helpers and the `.gsd/runtime/` and `.gsd/forensics/` directories.

Step 4: Health and status surfaces

For long-running processes:

  • **Health endpoint** (HTTP server) or **status file** (CLI tool). Cheap to call, no side effects. Returns current state: `{status: "healthy" | "degraded" | "down", ...diagnostics}`.
  • **Digest view** — a small representation of recent work. In gsd-pi, this is `STATE.md` and the health widget. In a server, it's `/internal/status` with last 10 request summaries.
  • **Minimal metrics** — counters for the 3–5 things that matter (requests, errors, active jobs). Not everything — just what drives alerts.

Don't build a metrics empire. Build exactly what you'd check at 3am.

Step 5: Explicit failure modes

Replace silent handling with explicit:

// Bad
try {
  return await db.getUser(id);
} catch {
  return null;
}

// Good
try {
  return await db.getUser(id);
} catch (err) {
  log.error({ event: "db-getuser-failed", userId: id, err: serializeError(err) });
  throw new DatabaseError("Failed to load user", { cause: err, userId: id });
}

The caller now knows the failure happened, gets an error type it can branch on, and a log line exists for forensics.

Step 6: Remove the scaffolding

Before shipping, cull the ad-hoc instrumentation you used while debugging. Keep only:

  • Decision-point logs that a future agent would use
  • Persistent failure state
  • Health/status surfaces
  • Explicit failure modes

Drop:

  • Temporary `console.log
Read more
Ships withgsd-pi

GSD Pi is a local-first coding agent for planning, implementing, verifying, and tracking project work from the command line.

Get the whole plugin

Other skills on gsd-pi.