trace-claude-code
Automatically trace Claude Code conversations to Braintrust for observability. Captures sessions, conversation turns, and tool calls as hierarchical traces.
Comprehensive testing workflow - unit tests ∥ integration tests → E2E tests
$ npx -y skills add parcadei/Continuous-Claude-v3 --skill test --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/testContext preview
The summary Claude sees to decide when to auto-load this skill.
Comprehensive testing workflow - unit tests ∥ integration tests → E2E tests
name: test description: Comprehensive testing workflow - unit tests ∥ integration tests → E2E tests
Run comprehensive test suite with parallel execution.
┌─────────────┐ ┌───────────┐
│ diagnostics │ ──▶ │ arbiter │ ─┐
│ (type check)│ │ (unit) │ │
└─────────────┘ └───────────┘ │
├──▶ ┌─────────┐
┌───────────┐ │ │ atlas │
│ arbiter │ ─┘ │ (e2e) │
│ (integ) │ └─────────┘
└───────────┘
Pre-flight Parallel Sequential
(~1 second) fast tests slow tests| # | Agent | Role | Execution | |---|-------|------|-----------| | 1 | **arbiter** | Unit tests, type checks, linting | Parallel | | 1 | **arbiter** | Integration tests | Parallel | | 2 | **atlas** | E2E/acceptance tests | After 1 passes |
1. **Fast feedback**: Unit tests fail fast 2. **Parallel efficiency**: No dependency between unit and integration 3. **E2E gating**: Only run slow E2E tests if faster tests pass
Before running tests, check for type errors - they often cause test failures:
tldr diagnostics . --project --format text 2>/dev/null | grep "^E " | head -10
**Why diagnostics first?**
**If errors found:** Fix them BEFORE running tests. Type errors usually mean tests will fail anyway.
**If clean:** Proceed to Phase 1.
For large test suites, find only affected tests:
tldr change-impact --session # or for explicit files: tldr change-impact src/changed_file.py
This returns which tests to run based on what changed. Skip this for small projects or when you want full coverage.
# Run both in parallel Task( subagent_type="arbiter", prompt=""" Run unit tests for: [SCOPE] Include: - Unit tests - Type checking - Linting Report: Pass/fail count, failures detail """, run_in_background=true ) Task( subagent_type="arbiter", prompt=""" Run integration tests for: [SCOPE] Include: - Integration tests - API tests - Database tests Report: Pass/fail count, failures detail """, run_in_background=true ) # Wait for both [Check TaskOutput for both]
Task( subagent_type="atlas", prompt=""" Run E2E tests for: [SCOPE] Include: - End-to-end flows - Acceptance tests - UI tests if applicable Report: Pass/fail count, screenshots on failure """ )
User: /test → All unit + integration + E2E tests
User: /test authentication → Only auth-related tests
User: /test --quick → Only unit tests (skip integration and E2E)
User: /test the new payment feature Claude: Starting /test workflow for payment feature... Phase 0: Pre-flight diagnostics... $ tldr diagnostics . --project --format text | grep "^E " (no type errors found) Phase 1: Running parallel tests... ┌─────────────────────────────────────────┐ │ arbiter: Running unit tests... │ │ arbiter: Running integration tests... │ └─────────────────────────────────────────┘ arbiter: ✅ 45/45 unit tests passing arbiter: ✅ 12/12 integration tests passing Phase 2: Running E2E tests... atlas: ✅ 8/8 E2E tests passing Test Summary: ┌─────────────┬─────────┬────────┐ │ Type │ Passed │ Failed │ ├─────────────┼─────────┼────────┤ │ Unit │ 45 │ 0 │ │ Integration │ 12 │ 0 │ │ E2E │ 8 │ 0 │ ├─────────────┼─────────┼────────┤ │ TOTAL │ 65 │ 0 │ └─────────────┴─────────┴────────┘ All tests passing! ✅
User: /test Claude: Starting /test workflow... Phase 0: Pre-flight diagnostics... $ tldr diagnostics . --project --format text | grep "^E " E src/payment.py:45:12: Argument of type 'str' not assignable to 'int' E src/refund.py:23:8: Return type 'None' not assignable to 'float' Found 2 type errors. Fixing before running tests... [Claude fixes the type errors] Re-running diagnostics... clean. Phase 1: Running parallel tests...
If Phase 1 fails:
arbiter: ❌ 43/45 tests passing 2 failures: - test_payment_validation: expected 'invalid' got 'valid' - test_refund_calculation: off by $0.01 Stopping workflow. Fix failures before running E2E tests.
A persistent, learning, multi-agent development environment built on Claude Code Continuous Claude transforms Claude Code into a continuously learning system that maintains context across sessions, orchestrates specialized agents, and eliminates wasting
Repo: parcadei/Continuous-Claude-v3
Automatically trace Claude Code conversations to Braintrust for observability. Captures sessions, conversation turns, and tool calls as hierarchical traces.
Guide for integrating Agentica SDK with Claude Code CLI proxy
Reference guide for Agentica multi-agent infrastructure APIs