Skip to content
Automation
Skill

/phoenix-arize-setup

Arize Phoenix observability platform setup for LLM debugging and evaluation

From plugin
babysitter
1.8k200 skills3 agents21 commands1 MCP
Install
$ npx -y skills add a5c-ai/babysitter --skill phoenix-arize-setup --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/phoenix-arize-setup

Context preview

The summary Claude sees to decide when to auto-load this skill.

Arize Phoenix observability platform setup for LLM debugging and evaluation

SKILL.md

phoenix-arize-setup.SKILL.md
name: phoenix-arize-setup
description: Arize Phoenix observability platform setup for LLM debugging and evaluation
allowed-tools:
  - Read
  - Write
  - Edit
  - Bash
  - Glob
  - Grep
graph:
  domains: [domain:software-engineering]
  specializations: [specialization:ai-agents-conversational]
  skillAreas: [skill-area:agent-debugging-logging, skill-area:eval-driven-development]
  roles: [role:ml-engineer, role:backend-engineer]
  workflows: [workflow:ml-model-lifecycle, workflow:feature-development]

Phoenix Arize Setup Skill

Capabilities

  • Set up Phoenix local server
  • Configure tracing instrumentation
  • Design evaluation experiments
  • Implement embedding visualizations
  • Set up retrieval analysis
  • Create custom evaluations with LLM-as-judge

Target Processes

  • llm-observability-monitoring
  • agent-evaluation-framework

Implementation Details

Core Features

1. **Tracing**: OpenTelemetry-based LLM traces 2. **Evals**: LLM-as-judge evaluations 3. **Embeddings**: Visualization and drift detection 4. **Retrieval**: RAG quality analysis 5. **Datasets**: Experiment management

Instrumentation

  • OpenAI auto-instrumentation
  • LangChain instrumentation
  • LlamaIndex instrumentation
  • Custom span creation

Configuration Options

  • Phoenix server setup
  • Trace sampling
  • Evaluation metrics
  • Embedding models
  • Export settings

Best Practices

  • Comprehensive instrumentation
  • Regular evaluation runs
  • Monitor embedding drift
  • Analyze retrieval quality

Dependencies

  • arize-phoenix
  • openinference-instrumentation-openai
Read more
Ships withbabysitter

Enforce obedience on agentic workforces. Manage extremely complex workflows through deterministic, hallucination-free self-orchestration.

Get the whole plugin

Other skills on babysitter.