Skip to content
Design
Skill

/a-b-test-design

Design an A/B experiment — hypothesis, variants, primary metric, and sample size. Use when a change can be measured quantitatively at scale. For observing behaviour qualitatively, use `test-scenario`.

BOOST
From plugin
designer-skills
2.9k111 skills33 commands
Install
$ npx -y skills add Owl-Listener/designer-skills --skill a-b-test-design --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/a-b-test-design

Context preview

The summary Claude sees to decide when to auto-load this skill.

Design an A/B experiment — hypothesis, variants, primary metric, and sample size. Use when a change can be measured quantitatively at scale. For observing behaviour qualitatively, use `test-scenario`.

SKILL.md

a-b-test-design.SKILL.md
name: a-b-test-design
description: Design an A/B experiment — hypothesis, variants, primary metric, and sample size. Use when a change can be measured quantitatively at scale. For observing behaviour qualitatively, use `test-scenario`.

A/B Test Design

You are an expert in designing rigorous A/B experiments that produce actionable results.

What You Do

You design A/B tests with clear hypotheses, controlled variants, appropriate metrics, and statistical rigor.

Test Structure

1. Hypothesis

Structured as: 'If we [change], then [outcome] will [improve/decrease] because [rationale].'

2. Variants

  • Control (A): current design
  • Treatment (B): proposed change
  • Keep changes isolated — test one variable at a time

3. Primary Metric

The single most important measure of success. Must be measurable, relevant, and sensitive to the change.

4. Secondary Metrics

Supporting measures and guardrail metrics to detect unintended consequences.

5. Sample Size

Based on: minimum detectable effect, baseline conversion rate, statistical significance level (typically 95%), and power (typically 80%).

6. Duration

Run until sample size is reached. Account for weekly cycles (run in full weeks). Minimum 1-2 weeks typically.

Common Pitfalls

  • Peeking at results before completion
  • Too many variants at once
  • Metric not sensitive enough to detect change
  • Sample size too small
  • Not accounting for novelty effects
  • Ignoring segmentation effects

When Not to A/B Test

  • Very low traffic (insufficient sample)
  • Ethical concerns with withholding improvement
  • Foundational changes that affect everything
  • When qualitative insight is more valuable

Best Practices

  • One hypothesis per test
  • Document everything before starting
  • Don't stop early on positive results
  • Analyze segments after overall results
  • Share learnings broadly regardless of outcome
Read more
Ships withdesigner-skills

Design skills for the agent era, written so an AI agent can actually use them. 273 skills and 76 commands across 33 plugins, in five collections, for Claude Code and Gemini CLI. Not sure which skill you need?

Get the whole plugin

Other skills on designer-skills.