adopt
Brownfield onboarding — audits existing project artifacts for template format compliance (not just existence), classifies gaps by impact, and produces a…
Detect non-deterministic (flaky) tests by reading CI run logs or test result history. Aggregates pass rates per test, identifies intermittent failures, recommends quarantine or fix, and maintains a flaky test registry. Best run during Polish phase or after multiple CI runs.
$ npx -y skills add Donchitos/Claude-Code-Game-Studios --skill test-flakiness --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/test-flakinessContext preview
The summary Claude sees to decide when to auto-load this skill.
Detect non-deterministic (flaky) tests by reading CI run logs or test result history. Aggregates pass rates per test, identifies intermittent failures, recommends quarantine or fix, and maintains a flaky test registry. Best run during Polish phase or after multiple CI runs.
name: test-flakiness description: "Detect non-deterministic (flaky) tests by reading CI run logs or test result history. Aggregates pass rates per test, identifies intermittent failures, recommends quarantine or fix, and maintains a flaky test registry. Best run during Polish phase or after multiple CI runs." argument-hint: "[ci-log-path | scan | registry]" user-invocable: true allowed-tools: Read, Glob, Grep, Write, Edit, Bash model: sonnet
A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.
**Output:** Updated `tests/regression-suite.md` quarantine section + optional `production/qa/flakiness-report-[date].md`
**When to run:**
---
**Modes:**
standard log output directories
section and provide remediation guidance for already-known flaky tests
`registry`
---
Check for test result artifacts:
ls -t .github/ 2>/dev/null ls -t test-results/ 2>/dev/null
For Godot projects: GdUnit4 outputs XML results compatible with JUnit format. Check `test-results/` for `.xml` files.
For Unity projects: game-ci test runner outputs NUnit XML to `test-results/` by default.
For Unreal projects: automation logs go to `Saved/Logs/`. Grep for `Result: Success` and `Result: Fail` patterns.
If a path argument is provided, read that file directly.
If no logs found: > "No CI log data found. To detect flaky tests, this skill needs test result > history from multiple runs. Options: > 1. Run the test suite at least 3 times and collect the output logs > 2. Check CI pipeline output and save a log to `test-results/` > 3. Run `/test-flakiness registry` to review tests already flagged as flaky > in `tests/regression-suite.md`"
Stop and ask the user which option to pursue.
---
For each CI log or result file found, parse:
**JUnit XML format** (GdUnit4 / Unity):
**Plain text logs**:
Build a table: `test_id → [run1_result, run2_result, run3_result, ...]`
---
A test is **flaky** if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.
Flakiness thresholds:
genuinely rare failure
For each flaky test, classify the likely cause:
| Cause | Symptoms | Fix direction | |-------|----------|---------------| | **Timing / async** | Fails after awaiting signals or timers; pass rate correlates with system load | Add explicit await/synchronisation; avoid time-based delays | | **Order dependency** | Fails when run after specific other tests; passes in isolation | Add proper setup/teardown; ensure test isolation | | **Random seed** | Fails intermittently with no pattern; involves RNG | Pass explicit seed; don't use `randf()` in tests | | **Resource leak** | Fails more often later in a test run | Fix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity) | | **External state** | Fails when a file, scene, or global exists from a prior test | Isolate test from file system; use in-memory mocks | | **Floating point** | Fails on comparisons like `== 0.5` | Use epsilon comparison (`is_equal_approx`, `Assert.AreApproximately`) | | **Scene/prefab load race** | Fails when scenes are not yet ready | Await one frame after instantiation; use `await get_tree().process_frame` |
Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.
---
For each flaky test:
**Quarantine (High flakiness):** > "Quarantine this test immediately. Disable it in CI by adding > `@pytest.mark.skip` / `[Ignore]` / `GdUnitSkip` annotation. Log it in > `tests/regression-suite.md` quarantine section. The test is now opt-in only. > Fix the root cause before removing quarantine."
**Investigate and fix soon (Moderate):** > "This test is intermittently unreliable. Root cause appears to be [cause]. > Suggested fix: [specific fix based on cause classification]. Do not quarantine > yet — fix the test directly."
**Monitor (Low/suspected):** > "This test shows suspected flakiness. Collect more run data before > quarantining. Note it as 'suspected' in the regression suite."
---
## Flakiness Detection Results **Runs analysed**: [N] **Tests tracked**: [N] ### Flaky Tests Found | Test | System | Fail Rate | Likely Cause | Recommendation | |------|--------|-----------|--------------|----------------| | [test_name] | [system] | [N]% |
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
Repo: Donchitos/Claude-Code-Game-Studios
Brownfield onboarding — audits existing project artifacts for template format compliance (not just existence), classifies gaps by impact, and produces a…
Creates an Architecture Decision Record (ADR) documenting a significant technical decision, its context, alternatives considered, and consequences. Every major…
Validates completeness and consistency of the project architecture against all GDDs. Builds a traceability matrix mapping every GDD technical requirement to…
Guided, section-by-section Art Bible authoring. Creates the visual identity specification that gates all asset production. Run after /brainstorm is approved…
Audits game assets for compliance with naming conventions, file size budgets, format standards, and pipeline requirements. Identifies orphaned assets, missing…
Generate per-asset visual specifications and AI generation prompts from GDDs, level docs, or character profiles. Produces structured spec files and updates the…