Skip to content

Claude Code Testing skills :)

Flowy lists 556 skills in the Testing category for Claude Code, across 57 plugins. A skill is a folder of instructions an agent loads while you work. Every one here shows what is inside before you install, who wrote it, and whether it is Auto-invoked, meaning a FLOW.md router fires it as you prompt.

Browse all skills →Testing plugins →

add-e2e-selectors
HotSkill

add-e2e-selectors

Add reliable @grafana/e2e-selectors to interactive elements and key containers in the Grafana frontend. Use when adding e2e selectors, data-testid attributes,…

@grafana@grafanaView Skill
frontend-testing-strategy
HotSkill

frontend-testing-strat…

Write unit and E2E tests for Grafana frontend code (React/TypeScript, any package or feature area) to the conventions this repo expects. Use when adding,…

@grafana@grafanaView Skill
panel-testing-strategy
HotSkill

panel-testing-strategy

Write unit and E2E tests for Grafana visualization panels and viz utilities to the conventions this repo expects. Use when adding, backfilling, or reviewing…

@grafana@grafanaView Skill
promptfoo-evals
Skill

promptfoo-evals

Write, refine, run, and QA promptfoo evaluation suites: promptfooconfig.yaml, prompts, providers, vars, tests, assertions, model-graded rubrics, transforms,…

@promptfoo@promptfooView Skill
deepeval
Skill

deepeval

DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow,…

@confident-ai@confident-aiView Skill
deepeval-otel
Skill

deepeval-otel

Export raw OpenTelemetry traces from an AI application to Confident AI's Observatory. TRIGGER when the user wants to send OpenTelemetry or OTLP traces/spans…

@confident-ai@confident-aiView Skill
deepeval-tracing
Skill

deepeval-tracing

Instrument an AI application with DeepEval's native tracing so its behavior is visible in Confident AI. TRIGGER when the user wants to add DeepEval tracing or…

@confident-ai@confident-aiView Skill
dev
Skill

dev

Development workflows for the playwright-cli repository. Use when the user asks about rolling dependencies, releasing, or other repo maintenance tasks.

browser
Skill

browser

Automate browsers with the Vibium CLI. Use to navigate websites, inspect pages, fill forms, extract page data, debug UI behavior, capture screenshots and…

@vibiumdev@vibiumdevView Skill
check
Skill

check

Independently check application acceptance criteria in a live browser or saved recording with the Vibium CLI. Use for a formal verification step in the…

@vibiumdev@vibiumdevView Skill
setup
Skill

setup

Set up or update TDD Guard for the current project. Detects the test framework, installs or updates the matching reporter, and configures or migrates its…

@nizos@nizosView Skill
e2e-test
Skill

e2e-test

Write a new Playwright E2E test for Vortex. Accepts a plain-text description of what to test, or a Linear issue ID to fetch the description automatically. Uses…

@nexus-mods@nexus-modsView Skill
watch-log
Skill

watch-log

Watch or investigate a Vortex log file (rotation- and session-aware). A router over six modes loaded on demand — live tail, session/crash/error investigation,…

@nexus-mods@nexus-modsView Skill
crabbox
Skill

crabbox

Detect and use Crabbox for repository tests and validation on remote runners. Use when crabbox.yaml or .crabbox.yaml exists, the crabbox CLI is available, or…

@openclaw@openclawView Skill
crabbox-quickstart
Skill

crabbox-quickstart

First contact with Crabbox: run your repository's tests inside a disposable Docker or Podman container on your own machine, no account and no cloud spend, then…

@openclaw@openclawView Skill
bom
Auto-invokedSkill

bom

BOM (Bill of Materials) management for electronics projects — the workflow skill that coordinates DigiKey, Mouser, LCSC, element14, JLCPCB, PCBWay, and KiCad…

@aklofas@aklofasView Skill
datasheets
Auto-invokedSkill

datasheets

Extract structured specifications from electronic component datasheet PDFs — pinouts, electrical characteristics, peripherals, topology, and features. Cache…

@aklofas@aklofasView Skill
digikey
Auto-invokedSkill

digikey

Search DigiKey for electronic components and download datasheets — primary source for prototype orders and the preferred API method for fetching datasheets.…

@aklofas@aklofasView Skill
criterium
Skill

criterium

Use this skill when users ask about benchmarking Clojure code, measuring performance, profiling execution time, or using the criterium library. Covers the…

@hugoduncan@hugoduncanView Skill
skill-upper
Skill

skill-upper

Capture and review Agent Skill observations, and create, run, diagnose, or iteratively improve Skill evaluations (evals) with the skill-up CLI / 采集和审核 Agent…

@alibaba@alibabaView Skill
agent-qa-authoring
Skill

agent-qa-authoring

Use when creating, editing, validating, or running agent-qa tests, suites, or hooks. Prefer agent-qa MCP tools, enforce canonical agent-qa IDs, and use the…

@vostride@vostrideView Skill
agent-qa-debug-fix
Skill

agent-qa-debug-fix

Use after an agent-qa run has failed and you need to debug, patch, and verify the issue using MCP evidence, logs, artifacts, and local code changes instead of…

@vostride@vostrideView Skill
agent-qa-result-triage
Skill

agent-qa-result-triage

Use when investigating failed agent-qa runs, inspecting artifacts, classifying failures, or comparing recent runs. Prefer agent-qa MCP run/artifact tools and…

@vostride@vostrideView Skill
00-arena-router
Skill

00-arena-router

Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate…

@agentscope-ai@agentscope-aiView Skill
00-meta-eval
Skill

00-meta-eval

Use when the user wants to build an evaluation system for an LLM/agent application but doesn't know where to start — they have traces, prompts, RAG pipelines,…

@agentscope-ai@agentscope-aiView Skill
old-coder
Skill

old-coder

Evidence-first development — surround the implementation with an executable spec and a gauntlet of constraints (tests, types, coverage, mutation) so…

@amazingang@amazingangView Skill
old-coder-api
Skill

old-coder-api

Design, change, or review an HTTP/JSON API surface — endpoints, request/response shapes, authentication and authorization, pagination, idempotency, rate…

@amazingang@amazingangView Skill
skillgrade-graders
Skill

skillgrade-graders

Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple…

@mgechev@mgechevView Skill
skillgrade-setup
Skill

skillgrade-setup

Sets up and runs skillgrade evaluation pipelines for Agent Skills. Use when initializing eval configurations, running trials, reviewing results, or integrating…

@mgechev@mgechevView Skill
create-adapter
Skill

create-adapter

Scaffold a new Harbor benchmark adapter by running `harbor adapter init` and then guide implementation using the Adapters Agent Guide as the authoritative spec.

@zli12321@zli12321View Skill
publish
Skill

publish

Publish a Harbor task or dataset to the registry. Use when the user wants to upload, publish, or share tasks or datasets/benchmarks on the Harbor registry.

@zli12321@zli12321View Skill
audit
Skill

audit

View audit logs, decision traces, and session history for AI transparency. ACTION_TYPES (19 entries) include PDCA events (phase_transition, gate_passed/failed,…

@popup-studio-ai@popup-studio-aiView Skill
bkend-auth
Skill

bkend-auth

bkend.ai authentication — email/social login, JWT tokens, RBAC, session management. Triggers: bkend auth, bkend login, bkend signup, bkend JWT, bkend RBAC

@popup-studio-ai@popup-studio-aiView Skill
canary-automate
Skill

canary-automate

Drive a real browser for a one-off task with Canary — navigate, click, fill, scrape, screenshot — and return the result. Nothing is recorded. Use when the user…

@0xnyn@0xnynView Skill
canary-review
Skill

canary-review

Open and triage recorded Canary sessions in the local viewer. Use when the user wants to look at, replay, compare, or triage a recorded session — or asks what…

@0xnyn@0xnynView Skill
canary-scripting
Skill

canary-scripting

The Canary sandbox scripting API for browser automation. Use when writing or debugging a Canary script — looking up how to open a page, click, fill, extract…

@0xnyn@0xnynView Skill
e2e-testing
Skill

e2e-testing

AI-powered E2E testing for any app — Flutter, React Native, iOS, Android, Electron, Tauri, KMP, .NET MAUI. Connects via MCP to running apps so the agent can…

kane-cli
Skill

kane-cli

Browser automation + AI test authoring via kane-cli - run browser objectives, generate & refine test scenarios/cases from a description, design…

kane-clitoday243
@lambdatest@lambdatestView Skill
add-seed-skills
Skill

add-seed-skills

Use when adding or editing QA skills in seed-skills/ or getting them onto the live qaskills.sh catalog, e.g. "add N new skills", "create a seed skill for X",…

@pramoddutta@pramodduttaView Skill
api-testing-rest
Skill

api-testing-rest

Comprehensive RESTful API testing patterns covering HTTP methods, status codes, request/response validation, authentication, error handling, and contract…

@pramoddutta@pramodduttaView Skill
claude-code-qa
Skill

claude-code-qa

The complete QA skill for Claude Code — turn Claude into an expert QA engineer that picks the right test type, writes reliable Playwright, Cypress, and pytest…

@pramoddutta@pramodduttaView Skill
wio
Skill

wio

Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing…

wio
Skill

wio

Testing workflow skill for finding high-value test candidates, writing focused tests, generating realistic workloads, reviewing test value, and diagnosing…

yapi
Skill

yapi

Query and sync YApi interface documentation. Use when user mentions "yapi 接口文档", YAPI docs, asks for request/response details, or needs docs sync. Also…

glance-test
Skill

glance-test

Run E2E browser tests on any web application using Glance MCP. Use when the user says "test this page," "check this URL," "run E2E tests," "browser test,"…

@debugbase@debugbaseView Skill
arch-check
Skill

arch-check

Use to check a feature's code against the charter's architecture rules — dependency layering, cycles, forbidden patterns, file naming, file size. Triggers —…

@swingerman@swingermanView Skill
atdd
Skill

atdd

Use to drive feature work through the Acceptance Test Driven Development workflow — Given/When/Then specs before code, a project-specific test pipeline, and…

@swingerman@swingermanView Skill
atdd-mutate
Skill

atdd-mutate

Use to add a third validation layer to the ATDD workflow — after acceptance tests verify WHAT and unit tests verify HOW, mutation testing verifies the tests…

@swingerman@swingermanView Skill
analyze
Auto-invokedSkill

analyze

Analyze a finished coder-eval run and write analysis.md — cluster failures into systemic patterns and recommend fixes. Use when the user wants to know why a…

@uipath@uipathView Skill
check-skill
Auto-invokedSkill

check-skill

Generate and run a coder-eval activation suite for a Claude Code skill. Use when the user asks whether a skill triggers, wants to test skill activation, or…

@uipath@uipathView Skill
ci
Auto-invokedSkill

ci

Generate a GitHub Actions workflow that runs a coder-eval suite as a CI gate or on a schedule, using the published composite action — with the agent runtime,…

@uipath@uipathView Skill
knowledge-base
Skill

knowledge-base

Retrieves and updates project-specific prompt knowledge from comparison evidence and user feedback. Use only for prompt analysis or post-comparison learning…

@shinpr@shinprView Skill
prompt-optimization
Skill

prompt-optimization

Improves LLM-facing context while preserving intent, execution boundaries, and proportional work. Use when creating or reviewing prompts, agent definitions,…

@shinpr@shinprView Skill
recipe-eval-prompt
Skill

recipe-eval-prompt

Compares original and optimized prompts through repeated blind paired execution in git worktrees. Use when evaluating prompt improvement effects or learning…

@shinpr@shinprView Skill
code-review
Skill

code-review

Use when a major project step has been completed and needs review against the plan and coding standards. Also use when someone says 'review this', 'check my…

e2e-playwright
Skill

e2e-playwright

Battle-tested Playwright E2E testing patterns for Next.js/React apps. Use when writing, running, debugging, or fixing Playwright tests. Also triggers on 'e2e',…

460 more in Testing. See them all →