/autonomous-testing
An AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type. Inspired by the edubites autonomous test runner pattern, generalized for Claude Bootstrap + Maggy.
$ npx -y skills add alinaqi/maggy --skill autonomous-testing --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/autonomous-testing
Context preview
The summary Claude sees to decide when to auto-load this skill.
An AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type. Inspired by the edubites autonomous test runner pattern, generalized for Claude Bootstrap + Maggy.
SKILL.md
autonomous-testing.SKILL.mdAutonomous Testing Agent
Overview
An AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type. Inspired by the edubites autonomous test runner pattern, generalized for Claude Bootstrap + Maggy.
Pipeline
Source Scan → Discover Gaps → Generate Tests → Execute → Evaluate → Report → Fix Loop
Phase 1: Discover — What Needs Testing?
Auto-detect project type:
Python → scan for *.py files, extract public functions/classes
TypeScript → scan for *.ts/*.tsx files, extract exports
API → scan FastAPI/Express routes, extract endpoints + methods
Web → scan React/Vue components, extract user flows
Map existing tests:
Python → pytest --collect-only
TypeScript → vitest --list
API → scan tests/ for endpoint coverage
Compute coverage gaps:
- Functions with 0 tests
- API endpoints with 0 tests
- Components with 0 tests
- Branches with <80% coverage
Phase 2: Generate — AI-Written Tests
For each uncovered function/endpoint/component:
1. Read source code → understand inputs, outputs, edge cases
2. Generate test scaffold using ~/bin/deepseek --pro
3. Include: happy path, error cases, edge cases, auth checks
4. Write to appropriate test directory
Model routing for generation:
- Simple functions → ~/bin/deepseek --flash (cheap, fast)
- Complex logic → ~/bin/deepseek --pro (thorough)
- Auth/security tests → ~/bin/deepseek --pro (quality-critical)
Phase 3: Execute — Run Everything
# Python
pytest -x --cov --cov-report=json
# TypeScript
npx vitest run --coverage
# E2E (if Playwright detected)
npx playwright test
# Parse results → structured TestRun { pass/fail, coverage, duration, failures[] }Phase 4: Evaluate — AI-Powered Assessment
For each test failure:
1. Capture: test name, error message, stack trace, source code diff
2. Classify failure:
- TEST_BUG: test is wrong (outdated expectation, bad mock)
- CODE_BUG: code is wrong (regression, edge case)
- ENV_BUG: environment issue (missing dep, config)
3. AI evaluation: ~/bin/deepseek --pro analyzes failure and classifies
For E2E/web tests:
- Capture screenshots at failure points
- ~/bin/gemini --flash evaluates visual state (multimodal)Phase 5: Fix — Autonomous Repair
TEST_BUG → regenerate test with corrected expectation
CODE_BUG → propose fix with ~/bin/deepseek --pro, apply, re-run
ENV_BUG → report to user with fix instructions
Auto-fix loop:
while test_failures > 0 and attempts < 3:
for each failure:
classify → fix → re-run
if fixed: record as "auto-fixed"
if not: escalate to CLAUDE tierPhase 6: Report — Structured Output
{
"project": "my-app",
"timestamp": "2026-05-16T12:00:00Z",
"summary": {
"tests_run": 247,
"passed": 231,
"failed": 12,
"auto_fixed": 8,
"needs_manual": 4,
"coverage": 0.83
},
"gaps_found": 15,
"tests_generated": 15,
"next_actions": [
"4 manual fixes needed in auth module",
"Coverage gap: src/payment.py has 0 tests",
"3 E2E flows untested: signup, checkout, profile-edit"
]
}Integration with Maggy
Maggy Dashboard → Testing tab shows:
- Coverage trend over time
- Auto-generated test count
- Failure classification (TEST_BUG vs CODE_BUG)
- "Generate tests for gaps" one-click button
Heartbeat job: auto-generate tests weekly for new untested code
Auto-review hook: triggers test generation after significant PR merges
Usage
# Discover test gaps
maggy test discover
# Generate tests for all gaps
maggy test generate --all
# Generate tests for specific module
maggy test generate --module auth
# Run full test cycle (discover → generate → execute → fix → report)
maggy test autonomous
# Watch mode — auto-test on file changes
maggy test watch
Configuration
// ~/.claude/testing-config.json
{
"auto_generate": true,
"auto_fix": true,
"max_fix_attempts": 3,
"min_coverage": 0.8,
"generate_model": "deepseek-pro",
"evaluate_model": "gemini-flash",
"fix_model": "deepseek-pro",
"exclude_patterns": ["*/migrations/*", "*/node_modules/*"]
}Read more
Autonomous Testing Agent
Overview
An AI-driven testing agent that auto-discovers, generates, executes, evaluates, and fixes tests for any project type. Inspired by the edubites autonomous test runner pattern, generalized for Claude Bootstrap + Maggy.
Pipeline
Source Scan → Discover Gaps → Generate Tests → Execute → Evaluate → Report → Fix Loop
Phase 1: Discover — What Needs Testing?
Auto-detect project type: Python → scan for *.py files, extract public functions/classes TypeScript → scan for *.ts/*.tsx files, extract exports API → scan FastAPI/Express routes, extract endpoints + methods Web → scan React/Vue components, extract user flows Map existing tests: Python → pytest --collect-only TypeScript → vitest --list API → scan tests/ for endpoint coverage Compute coverage gaps: - Functions with 0 tests - API endpoints with 0 tests - Components with 0 tests - Branches with <80% coverage
Phase 2: Generate — AI-Written Tests
For each uncovered function/endpoint/component: 1. Read source code → understand inputs, outputs, edge cases 2. Generate test scaffold using ~/bin/deepseek --pro 3. Include: happy path, error cases, edge cases, auth checks 4. Write to appropriate test directory Model routing for generation: - Simple functions → ~/bin/deepseek --flash (cheap, fast) - Complex logic → ~/bin/deepseek --pro (thorough) - Auth/security tests → ~/bin/deepseek --pro (quality-critical)
Phase 3: Execute — Run Everything
# Python
pytest -x --cov --cov-report=json
# TypeScript
npx vitest run --coverage
# E2E (if Playwright detected)
npx playwright test
# Parse results → structured TestRun { pass/fail, coverage, duration, failures[] }Phase 4: Evaluate — AI-Powered Assessment
For each test failure:
1. Capture: test name, error message, stack trace, source code diff
2. Classify failure:
- TEST_BUG: test is wrong (outdated expectation, bad mock)
- CODE_BUG: code is wrong (regression, edge case)
- ENV_BUG: environment issue (missing dep, config)
3. AI evaluation: ~/bin/deepseek --pro analyzes failure and classifies
For E2E/web tests:
- Capture screenshots at failure points
- ~/bin/gemini --flash evaluates visual state (multimodal)Phase 5: Fix — Autonomous Repair
TEST_BUG → regenerate test with corrected expectation
CODE_BUG → propose fix with ~/bin/deepseek --pro, apply, re-run
ENV_BUG → report to user with fix instructions
Auto-fix loop:
while test_failures > 0 and attempts < 3:
for each failure:
classify → fix → re-run
if fixed: record as "auto-fixed"
if not: escalate to CLAUDE tierPhase 6: Report — Structured Output
{
"project": "my-app",
"timestamp": "2026-05-16T12:00:00Z",
"summary": {
"tests_run": 247,
"passed": 231,
"failed": 12,
"auto_fixed": 8,
"needs_manual": 4,
"coverage": 0.83
},
"gaps_found": 15,
"tests_generated": 15,
"next_actions": [
"4 manual fixes needed in auth module",
"Coverage gap: src/payment.py has 0 tests",
"3 E2E flows untested: signup, checkout, profile-edit"
]
}Integration with Maggy
Maggy Dashboard → Testing tab shows: - Coverage trend over time - Auto-generated test count - Failure classification (TEST_BUG vs CODE_BUG) - "Generate tests for gaps" one-click button Heartbeat job: auto-generate tests weekly for new untested code Auto-review hook: triggers test generation after significant PR merges
Usage
# Discover test gaps maggy test discover # Generate tests for all gaps maggy test generate --all # Generate tests for specific module maggy test generate --module auth # Run full test cycle (discover → generate → execute → fix → report) maggy test autonomous # Watch mode — auto-test on file changes maggy test watch
Configuration
// ~/.claude/testing-config.json
{
"auto_generate": true,
"auto_fix": true,
"max_fix_attempts": 3,
"min_coverage": 0.8,
"generate_model": "deepseek-pro",
"evaluate_model": "gemini-flash",
"fix_model": "deepseek-pro",
"exclude_patterns": ["*/migrations/*", "*/node_modules/*"]
}Turn Claude Code into a self-reviewing, test-enforced engineering system that remembers context across sessions — then route work across 13 models from a single dashboard.
Repo: alinaqi/maggy
Other skills on maggy.
- /aeo-optimization
AI Engine Optimization - semantic triples, page templates, content clusters for AI citations
Open skill - /agent-teams
Claude Code Agent Teams - default team-based development with strict TDD pipeline enforcement
Open skill - /agentic-development
Build AI agents with Pydantic AI (Python) and Claude SDK (Node.js)
Open skill - /ai-models
Latest AI models reference - Claude, OpenAI, Gemini, Eleven Labs, Replicate
Open skill - /android-java
Android Java development with MVVM, ViewBinding, and Espresso testing
Open skill - /android-kotlin
Android Kotlin development with Coroutines, Jetpack Compose, Hilt, and MockK testing
Open skill

