annotations
Annotations are named labels/scores attached to spans, traces, or experiment runs by humans, code, or LLM judges.
$ npx -y skills add arize-ai/phoenix --agent claude-codeHow it fires
How this agent gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
Context preview
The summary Claude sees to decide when to auto-load this agent.
Annotations are named labels/scores attached to spans, traces, or experiment runs by humans, code, or LLM judges.
Agent definition
annotations.mdAnnotations
Annotations are named labels/scores attached to spans, traces, or experiment runs by humans, code, or LLM judges.
Fields
`SpanAnnotation` and `TraceAnnotation` share these fields:
- `name`, `label`, `score`, `explanation`
- `annotatorKind` (`AnnotatorKind` enum: human / LLM / code)
- `metadata`, `identifier`, `createdAt`, `updatedAt`
`SpanAnnotation` additionally has `spanId: GlobalID!` and `span`; `TraceAnnotation` has `trace` (but no `traceId` global-id field).
`ExperimentRunAnnotation` has the same scalar shape plus `error: String`, and uses `startTime`/`endTime` instead of `createdAt`/`updatedAt`.
Reading annotations
- Per span: `Span.spanAnnotations { name label score explanation annotatorKind }`.
- Project-wide discovery and rollups: `Project.spanAnnotationNames`, `Project.spanAnnotationSummary`, `Project.traceAnnotationsNames`, `Project.traceAnnotationSummary` — use these to learn which annotation names exist before drilling in.
- In a span `filterCondition`, reference annotations as `annotations['<name>'].label` / `.score` (or the legacy `evals['<name>']`).
Example
Spans that an LLM judge labelled as hallucinated, with the annotation detail:
query Hallucinations($id: ID!) {
node(id: $id) {
... on Project {
spans(first: 20, filterCondition: "annotations['Hallucination'].label == 'hallucinated'") {
edges {
node {
spanId
spanAnnotations { name label score explanation }
}
}
}
}
}
}Read more
Annotations
Annotations are named labels/scores attached to spans, traces, or experiment runs by humans, code, or LLM judges.
Fields
`SpanAnnotation` and `TraceAnnotation` share these fields:
- `name`, `label`, `score`, `explanation`
- `annotatorKind` (`AnnotatorKind` enum: human / LLM / code)
- `metadata`, `identifier`, `createdAt`, `updatedAt`
`SpanAnnotation` additionally has `spanId: GlobalID!` and `span`; `TraceAnnotation` has `trace` (but no `traceId` global-id field).
`ExperimentRunAnnotation` has the same scalar shape plus `error: String`, and uses `startTime`/`endTime` instead of `createdAt`/`updatedAt`.
Reading annotations
- Per span: `Span.spanAnnotations { name label score explanation annotatorKind }`.
- Project-wide discovery and rollups: `Project.spanAnnotationNames`, `Project.spanAnnotationSummary`, `Project.traceAnnotationsNames`, `Project.traceAnnotationSummary` — use these to learn which annotation names exist before drilling in.
- In a span `filterCondition`, reference annotations as `annotations['<name>'].label` / `.score` (or the legacy `evals['<name>']`).
Example
Spans that an LLM judge labelled as hallucinated, with the annotation detail:
query Hallucinations($id: ID!) {
node(id: $id) {
... on Project {
spans(first: 20, filterCondition: "annotations['Hallucination'].label == 'hallucinated'") {
edges {
node {
spanId
spanAnnotations { name label score explanation }
}
}
}
}
}
}Repo: arize-ai/phoenix
Other agents on phoenix.
- VENDOR
- Upstream: https://github.com/DougTrajano/pydantic-ai-skills - Original version: 0.10.1 (SHA `d1a19c8e6dbbee5e6726c8de70aadca2966fb190`) - Vendored: 2026-05-19 - License: MIT — preserved verbatim in [LICENSE](./LICENSE)
Open agent - datasets
There is **no `getDatasetByName`** — fetch via `node(id:) { ... on Dataset { ... } }` or the `datasets(filter: DatasetFilter, sort)` connection.
Open agent - experiments
An experiment is one run of a prompt or pipeline over every example in a dataset.
Open agent - project-spans-traces
- `Project.spans(timeRange, first, after, sort: SpanSort, rootSpansOnly: Boolean, filterCondition: String)` → connection of `Span`. There is **no `traces` connection on `Project`** — use `spans(rootSpansOnly: true)` for root spans, which is usually one per trace though nothing
Open agent - prompts
There is **no `getPromptByName`** — fetch via `node(id:)` or the `prompts(filter: PromptFilter, labelIds)` connection.
Open agent - sessions
A session groups the traces of one multi-turn conversation.
Open agent

