Skip to content
Development
Skill

/agent-evaluation

Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement quality.

From plugin
context-engineering-kit
1.3k134 skills23 agents1 command
Install
$ npx -y skills add NeoLabHQ/context-engineering-kit --skill agent-evaluation --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/agent-evaluation

Context preview

The summary Claude sees to decide when to auto-load this skill.

Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement quality.

SKILL.md

agent-evaluation.SKILL.md
name: agent-evaluation
description: Evaluate and improve Claude Code commands, skills, and agents. Use when testing prompt effectiveness, validating context engineering choices, or measuring improvement quality.

Evaluation Methods for Claude Code Agents

Evaluation of agent systems requires different approaches than traditional software or even standard language model applications. Agents make dynamic decisions, are non-deterministic between runs, and often lack single correct answers. Effective evaluation must account for these characteristics while providing actionable feedback. A robust evaluation framework enables continuous improvement, catches regressions, and validates that context engineering choices achieve intended effects.

Core Concepts

Agent evaluation requires outcome-focused approaches that account for non-determinism and multiple valid paths. Multi-dimensional rubrics capture various quality aspects: factual accuracy, completeness, citation accuracy, source quality, and tool efficiency. LLM-as-judge provides scalable evaluation while human evaluation catches edge cases.

The key insight is that agents may find alternative paths to goals—the evaluation should judge whether they achieve right outcomes while following reasonable processes.

**Performance Drivers: The 95% Finding** Research on the BrowseComp evaluation (which tests browsing agents' ability to locate hard-to-find information) found that three factors explain 95% of performance variance:

| Factor | Variance Explained | Implication | |--------|-------------------|-------------| | Token usage | 80% | More tokens = better performance | | Number of tool calls | ~10% | More exploration helps | | Model choice | ~5% | Better models multiply efficiency |

Implications for Claude Code development:

  • **Token budgets matter**: Evaluate with realistic token constraints
  • **Model upgrades beat token increases**: Upgrading models provides larger gains than increasing token budgets
  • **Multi-agent validation**: Validates architectures that distribute work across subagents with separate context windows

Evaluation Challenges

Non-Determinism and Multiple Valid Paths

Agents may take completely different valid paths to reach goals. One agent might search three sources while another searches ten. They might use different tools to find the same answer. Traditional evaluations that check for specific steps fail in this context.

**Solution**: The solution is outcomes, not exact execution paths. Judge whether the agent achieves the right result through a reasonable process.

Context-Dependent Failures

Agent failures often depend on context in subtle ways. An agent might succeed on complex queries but fail on simple ones. It might work well with one tool set but fail with another. Failures may emerge only after extended interaction when context accumulates.

**Solution**: Evaluation must cover a range of complexity levels and test extended interactions, not just isolated queries.

Composite Quality Dimensions

Agent quality is not a single dimension. It includes factual accuracy, completeness, coherence, tool efficiency, and process quality. An agent might score high on accuracy but low in efficiency, or vice versa.

An agent might score high on accuracy but low in efficiency.

**Solution**: Evaluation rubrics must capture multiple dimensions with appropriate weighting for the use case.

Evaluation Rubric Design

Multi-Dimensional Rubric

Effective rubrics cover key dimensions with descriptive levels:

**Instruction Following** (weight: 0.30)

  • Excellent (1.0): All instructions followed precisely
  • Good (0.8): Minor deviations that don't affect outcome
  • Acceptable (0.6): Major instructions followed, minor ones missed
  • Poor (0.3): Significant instructions ignored
  • Failed (0.0): Fundamentally misunderstood the task

**Output Completeness** (weight: 0.25)

  • Excellent: All requested aspects thoroughly covered
  • Good: Most aspects covered with minor gaps
  • Acceptable: Key aspects covered, some gaps
  • Poor: Major aspects missing
  • Failed: Fundamental aspects not addressed

**Tool Efficiency** (weight: 0.20)

  • Excellent: Optimal tool selection and minimal calls
  • Good: Good tool selection with minor inefficiencies
  • Acceptable: Appropriate tools with some redundancy
  • Poor: Wrong tools or excessive calls
  • Failed: Severe tool misuse or extremely excessive calls

**Reasoning Quality** (weight: 0.15)

  • Excellent: Clear, logical reasoning throughout
  • Good: Generally sound reasoning with minor gaps
  • Acceptable: Basic reasoning present
  • Poor: Reasoning unclear or flawed
  • Failed: No apparent reasoning

**Response Coherence** (weight: 0.10)

  • Excellent: Well-structured, easy to follow
  • Good: Generally coherent with minor issues
  • Acceptable: Understandable but could be clearer
  • Poor: Difficult to follow
  • Failed: Incoherent

Scoring Approach

Convert dimension assessments to numeric scores (0.0 to 1.0) with appropriate weighting. Calculate weighted overall scores. Set passing thresholds based on use case requirements (typically 0.7 for general use, 0.85 for critical operations).

Evaluation Methodologies

LLM-as-Judge

Using an LLM to evaluate agent outputs scales well and provides consistent judgments. Design evaluation prompts that capture the dimensions of interest. LLM-based evaluation scales to large test sets and provides consistent judgments. The key is designing effective evaluation prompts that capture the dimensions of interest.

Provide clear task description, agent output, ground truth (if available), evaluation scale with level descriptions, and request structured judgment.

**Evaluation Prompt Template**:

You are evaluating the output of a Claude Code agent.

## Original Task
{task_description}

## Agent Output
{agent_output}

## Ground Truth (if available)
{expected_output}

## Evaluation Criteria
For each criterion, assess the output and provide:
1. Scor
Read more
Ships withcontext-engineering-kit

A hand-crafted collection of advanced context engineering techniques and patterns with minimal token footprint, focused on improving agent result quality and predictability.

Get the whole plugin