experiments
An experiment is one run of a prompt or pipeline over every example in a dataset.
$ npx -y skills add arize-ai/phoenix --agent claude-codeHow it fires
How this agent gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
Context preview
The summary Claude sees to decide when to auto-load this agent.
An experiment is one run of a prompt or pipeline over every example in a dataset.
Agent definition
experiments.mdExperiment, ExperimentRun
An experiment is one run of a prompt or pipeline over every example in a dataset.
Reaching an experiment
There is **no `getExperimentById`** — reach an `Experiment` via `node(id:)`, `Dataset.experiments`, or `compareExperiments`.
Experiment fields
- `name`, `description`, `sequenceNumber`, `repetitions`, `isEphemeral`
- `dataset`, `datasetVersion`, `project`
- `runs(first, after, sort: ExperimentRunSort)` — **forward-only** (no `last`/`before`)
- `runCount`, `expectedRunCount`
- `errorRate`, `averageRunLatencyMs`, `costSummary`, `costDetailSummaryEntries`
- `annotationSummaries { annotationName meanScore minScore maxScore count errorCount }`
Comparison
For candidate comparison prefer `compareExperiments(baseExperimentId: GlobalID!, compareExperimentIds: [GlobalID!]!, first, after, filterCondition)` over fetching each experiment's runs separately. Related: `experimentRunMetricComparisons(baseExperimentId, compareExperimentIds)` and `validateExperimentRunFilterCondition(condition, experimentIds)`.
Example
query ExperimentMetrics($id: ID!) {
node(id: $id) {
... on Experiment {
name
sequenceNumber
runCount
errorRate
averageRunLatencyMs
annotationSummaries { annotationName meanScore count errorCount }
}
}
}Repo: arize-ai/phoenix
Other agents on phoenix.
- VENDOR
- Upstream: https://github.com/DougTrajano/pydantic-ai-skills - Original version: 0.10.1 (SHA `d1a19c8e6dbbee5e6726c8de70aadca2966fb190`) - Vendored: 2026-05-19 - License: MIT — preserved verbatim in [LICENSE](./LICENSE)
Open agent - annotations
Annotations are named labels/scores attached to spans, traces, or experiment runs by humans, code, or LLM judges.
Open agent - datasets
There is **no `getDatasetByName`** — fetch via `node(id:) { ... on Dataset { ... } }` or the `datasets(filter: DatasetFilter, sort)` connection.
Open agent - project-spans-traces
- `Project.spans(timeRange, first, after, sort: SpanSort, rootSpansOnly: Boolean, filterCondition: String)` → connection of `Span`. There is **no `traces` connection on `Project`** — use `spans(rootSpansOnly: true)` for root spans, which is usually one per trace though nothing
Open agent - prompts
There is **no `getPromptByName`** — fetch via `node(id:)` or the `prompts(filter: PromptFilter, labelIds)` connection.
Open agent - sessions
A session groups the traces of one multi-turn conversation.
Open agent

