Skip to content
Development
Skill

/subagent-testing

Test skills via TDD in fresh subagents. Use when validating behavior or preventing bias.

From plugin
claude-night-market
337200 skills59 agents162 commands1 MCP
Install
$ npx -y skills add athola/claude-night-market --skill subagent-testing --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/subagent-testing

Context preview

The summary Claude sees to decide when to auto-load this skill.

Test skills via TDD in fresh subagents. Use when validating behavior or preventing bias.

SKILL.md

subagent-testing.SKILL.md
name: subagent-testing
description: 'Test skills via TDD in fresh subagents. Use when validating behavior or preventing bias.'
alwaysApply: false
category: testing
tags:
- testing
- validation
- TDD
- subagents
- fresh-instances
token_budget: 30
progressive_loading: true
modules:
- modules/testing-patterns.md
model_hint: standard

Subagent Testing - TDD for Skills

Test skills with fresh subagent instances to prevent priming bias and validate effectiveness.

When NOT To Use

  • Writing the skill under test (use `abstract:skill-authoring`)
  • A static quality audit with no execution (use `abstract:skills-eval`)

Table of Contents

1. [Overview](#overview) 2. [Why Fresh Instances Matter](#why-fresh-instances-matter) 3. [Testing Methodology](#testing-methodology) 4. [Quick Start](#quick-start) 5. [Detailed Testing Guide](#detailed-testing-guide) 6. [Success Criteria](#success-criteria)

Overview

**Fresh instances prevent priming:** Each test uses a new Claude conversation to verify the skill's impact is measured, not conversation history effects.

Why Fresh Instances Matter

The Priming Problem

Running tests in the same conversation creates bias:

  • Prior context influences responses
  • Skill effects get mixed with conversation history
  • Can't isolate skill's true impact

Fresh Instance Benefits

  • **Isolation**: Each test starts clean
  • **Reproducibility**: Consistent baseline state
  • **Measurement**: Clear before/after comparison
  • **Validation**: Proves skill effectiveness, not priming

Testing Methodology

Three-phase TDD-style approach:

Phase 1: Baseline Testing (RED)

Test without skill to establish baseline behavior.

Phase 2: With-Skill Testing (GREEN)

Test with skill loaded to measure improvements.

Phase 3: Rationalization Testing (REFACTOR)

Test skill's anti-rationalization guardrails.

Quick Start

# 1. Create baseline tests (without skill)
# Use 5 diverse scenarios
# Document full responses

# 2. Create with-skill tests (fresh instances)
# Load skill explicitly
# Use identical prompts
# Compare to baseline

# 3. Create rationalization tests
# Test anti-rationalization patterns
# Verify guardrails work

Detailed Testing Guide

For complete testing patterns, examples, and templates:

  • **[Testing Patterns](modules/testing-patterns.md)** - Full TDD methodology
  • **[Test Examples](modules/testing-patterns.md)** - Baseline, with-skill, rationalization tests
  • **[Analysis Templates](modules/testing-patterns.md)** - Scoring and comparison frameworks

Success Criteria

  • **Baseline**: Document 5+ diverse baseline scenarios
  • **Improvement**: ≥50% improvement in skill-related metrics
  • **Consistency**: Results reproducible across fresh instances
  • **Rationalization Defense**: Guardrails prevent ≥80% of rationalization attempts

See Also

  • **skill-authoring**: Creating effective skills
  • **test-skill**: Automated skill testing command

Exit Criteria

  • [ ] Baseline (RED) phase documents at least 5 diverse scenarios run in fresh Claude instances

without the skill active, with full response text recorded.

  • [ ] With-skill (GREEN) phase uses identical prompts in new fresh instances (not continuations

of the baseline conversation) and shows >= 50% improvement on skill-related metrics.

  • [ ] Rationalization (REFACTOR) phase shows skill guardrails blocking >= 80% of rationalization

attempts tested across at least 3 pressure scenarios.

  • [ ] Results are reproducible: the same prompts in a new fresh instance produce consistent

outcomes, confirming the effect is not conversation-history priming.

Read more
Ships withclaude-night-market

A plugin marketplace for Claude Code. Install only the plugins you need to run git workflows, code review, spec-driven development, and autonomous agents from inside your Claude Code session.

Get the whole plugin

Other skills on claude-night-market.