analyze-selective-test…
Analyzes Xcode selective testing effectiveness for a test run, showing which test targets…
Compares a single test case's behavior across two branches, analyzing pass/fail status, duration, flakiness, and failure details. Useful for investigating test regressions introduced by a feature branch.
$ npx -y skills add tuist/tuist --skill compare-test-case --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/compare-test-caseContext preview
The summary Claude sees to decide when to auto-load this skill.
Compares a single test case's behavior across two branches, analyzing pass/fail status, duration, flakiness, and failure details. Useful for investigating test regressions introduced by a feature branch.
name: compare-test-case description: Compares a single test case's behavior across two branches, analyzing pass/fail status, duration, flakiness, and failure details. Useful for investigating test regressions introduced by a feature branch.
You'll typically receive a test case identifier and two branches. Follow these steps:
1. Run `tuist test case show <id-or-identifier> --json` to get the test case metrics. 2. Run `tuist test case run list <identifier> --json` to see runs across branches. 3. Compare behavior between base and head branches. 4. Inspect failures with `tuist test case run show <run-id> --json`. 5. Summarize findings with root cause analysis.
tuist test case show <test-case-id> --json
tuist test case show Module/Suite/TestCase --json
Discover flaky or failing tests to investigate:
tuist test case list --flaky --json --page-size 10
Key fields from the response:
List test case runs filtered by the test case, and look at the `git_branch` field:
tuist test case run list <identifier> --json --page-size 20
Separate runs by branch. For each branch, compute:
| Metric | Base branch | Head branch | Verdict | |---|---|---|---| | Pass rate | e.g. 100% | e.g. 60% | REGRESSION | | Avg duration | e.g. 0.5s | e.g. 2.1s | REGRESSION | | Flaky runs | 0 | 3 | NEW FLAKINESS | | Last status | success | failure | REGRESSION |
Classify the change:
For each failing run on the head branch:
tuist test case run show <test-case-run-id> --json
Examine:
Based on the comparison:
Produce a summary with:
1. **Test case info**: Name, module, suite, overall reliability. 2. **Base branch behavior**: Pass rate, avg duration, flaky count. 3. **Head branch behavior**: Pass rate, avg duration, flaky count. 4. **Verdict**: What changed and classification. 5. **Root cause**: Hypothesis based on failure analysis. 6. **Recommendations**: Specific file paths, line numbers, and fix suggestions.
Example:
Test Case Comparison: AuthModuleTests/LoginTests/test_login_with_expired_token Overall reliability: 85% (was 100% before head branch) Base (main): Pass rate: 100% (15/15 runs) Avg duration: 0.3s Flaky: No Head (feature/auth-refactor): Pass rate: 60% (3/5 runs) Avg duration: 0.5s Flaky: Yes (2 flaky runs) Verdict: NEWLY FLAKY -- test was stable on main but intermittently fails on feature branch Root cause: The auth refactor introduced an async token refresh that races with the test's synchronous assertion. Failures show "Expected status 401, got nil" at Tests/AuthModuleTests/LoginTests.swift:42, suggesting the response arrives before the token refresh completes. Recommendations: - Add an await/expectation before the assertion at LoginTests.swift:42 - Consider mocking the token refresh to make the test deterministic
Tuist supercharges your build system, whether you build with Xcode, Gradle, or Bazel.
Analyzes Xcode selective testing effectiveness for a test run, showing which test targets…
Compares two Xcode build runs to identify duration regressions, cache changes, and new…
Compares two app bundles to identify size changes, new or removed artifacts, and platform…
Compares two `tuist cache` runs to identify cache hit rate changes and root-cause analysis of…
Compares two `tuist generate` runs to identify cache hit rate changes and root-cause analysis…
Compares two Gradle build runs to identify duration regressions, cache changes, and task…