Skip to content
Development
Skill

/annotate-spans

Write effective, consistent annotations on LLM/agent spans and traces, and coach the user on annotation practice. Load this whenever you are about to record structured feedback with the `batch_span_annotate` tool, or when the user asks how to annotate, label, score, or review

From plugin
phoenix
11k8 skills7 agents
Install
$ npx -y skills add arize-ai/phoenix --skill annotate-spans --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/annotate-spans

Context preview

The summary Claude sees to decide when to auto-load this skill.

Write effective, consistent annotations on LLM/agent spans and traces, and coach the user on annotation practice. Load this whenever you are about to record structured feedback with the `batch_span_annotate` tool, or when the user asks how to annotate, label, score, or review

SKILL.md

annotate-spans.SKILL.md
name: annotate-spans
description: >
  Write effective, consistent annotations on LLM/agent spans and traces, and coach the user on annotation practice. Load this whenever you are about to record structured feedback with the `batch_span_annotate` tool, or when the user asks how to annotate, label, score, or review spans/traces, build a failure taxonomy, or set up human/LLM review. Do NOT load for: pure analysis with no intent to save feedback (use debug-trace), latency or cost statistics, or prompt authoring (use playground).
summary: Create consistent span or trace annotations and help design useful feedback taxonomies.

Annotating Spans and Traces

An annotation is durable, structured feedback attached to a span or trace: a `name` (the dimension being judged), an optional `label` and/or `score` (the outcome), and an `explanation` (why). Annotations are not throwaway commentary — they accumulate into a dataset the user filters, aggregates, and iterates against.

A good annotation earns its place by being useful *later*:

  • **Filterable** — `annotations['answer_relevance'].label == 'fail'` returns the spans you meant.
  • **Aggregatable** — counting labels across spans yields a failure rate that tells the user where to focus.
  • **Auditable** — months later, the explanation still justifies the judgment without rerunning anything.
  • **Curatable** — failing spans can be pulled into a dataset to drive evals or fixes.

This skill governs the *judgment* behind annotations. The `batch_span_annotate` tool description governs the *mechanics* (one array, ID requirements, update keying); follow both, and never contradict the tool's naming and identifier rules.

What Makes an Annotation Useful

1. **Grounded in observed behavior, not generic quality vibes.** Annotate what actually went wrong or right in *this* span. "Cited a refund policy that does not exist in the retrieved context" beats a free-floating `hallucination_score: 0.3`. Generic dimensions like `helpfulness` or `coherence` are rarely grounded in the application's real failure modes — prefer names that point at a concrete behavior.

2. **One dimension per annotation.** `name` is the rubric dimension; the outcome lives in `label`/`score`. Use `name: "tool_selection"`, `label: "incorrect"` — not `name: "wrong_tool"`. If you find yourself judging two things at once (e.g., retrieval relevance *and* answer faithfulness), write two annotations.

3. **Target the most specific responsible span.** Annotate the LLM span for model output, the tool span for tool behavior, the retriever span for retrieval quality. Reserve root agent/chain spans for genuinely end-to-end judgments (task success, trajectory). A faithfulness failure pinned to the right LLM span is actionable; the same label on the root span forces the user to hunt.

4. **Judge the first failure, not every downstream symptom.** Errors cascade — bad retrieval produces a bad answer. Annotate the root cause where it occurred. Add a separate annotation downstream only when it reveals an independent problem, not a consequence of the first.

5. **Prefer crisp labels over fuzzy scores.** A binary or small categorical label (`pass`/`fail`, `relevant`/`irrelevant`, `correct`/`partial`/`incorrect`) is easy to apply consistently and easy to aggregate. Use a numeric `score` only when the scale is genuinely meaningful and defined; put the rubric, scale, or threshold in `metadata` so the number is interpretable later.

6. **Explanations are specific observations, not restatements.** Write what you saw, citing the evidence. Good: "Returned chunks about onboarding; user asked about cancellation — no relevant chunk retrieved." Weak: "The retrieval was bad." Always include an explanation for any score, any failure, any unclear label, or any judgment the user might want to revisit.

7. **Be consistent across spans.** The same dimension must use the same `name` and the same label vocabulary everywhere, or filtering and rate computation break. The project's annotation configs *are* that shared vocabulary — annotate into a config's `name` and labels rather than deciding fresh each run (see [Work From the Project's Annotation Configs](#work-from-the-projects-annotation-configs)). Keep names stable across runs (no `_v2`/`_new` suffixes).

8. **Set `annotatorKind` honestly.** `LLM` for your own judgment, `HUMAN` only when recording feedback the user explicitly gave, `CODE` for deterministic checks. Don't record your own opinion as `HUMAN`.

Work From the Project's Annotation Configs

An **annotation config** is the project's codified rubric for one dimension: a `name`, a type (categorical / continuous / freeform), and the allowed outcomes — a categorical config's `values` (each a `label` with an optional `score`), or a continuous config's `lowerBound`/`upperBound`. Configs are the source of truth for annotation vocabulary: they drive the annotation UI, keep names and labels consistent across runs, and let a later visit reuse the grading criteria instead of reinventing it. Annotate *into* configs rather than inventing a name and label set each time.

Before writing any annotation:

1. **Pull the project's configs first.** Query `Project.annotationConfigs` and read the existing names and their label/score schemes. This is the established rubric — prefer it over anything you would invent. (Resolve the project node id as described in the `phoenix-graphql` skill.)

   query ProjectAnnotationConfigs($projectId: ID!) {
     node(id: $projectId) {
       ... on Project {
         annotationConfigs(first: 100) {
           edges { node {
             __typename
             ... on AnnotationConfigBase { name description annotationType }
             ... on CategoricalAnnotationConfig { id optimizationDirection values { label score } }
             ... on ContinuousAnnotationConfig { id optimizationDirection lowerBound upperBound }
             ... on FreeformAnnotationConfig { id }
           } }
Read more
Ships withphoenix

AI Observability & Evaluation

Get the whole plugin

Other skills on phoenix.