agent-activity
Streams what the agent is doing into the room, as rows the desktop client renders in an **events drawer** above the composer (collapsed: avatar, pulsing dots,…
Drive a fixed suite of spoken tests against a voice agent ("subject") from a co-located machine ("prober"), measure response latency / clarity / accuracy, diff against baseline, and report to the owner over Telegram.
$ npx -y skills add sonichi/sutando --skill voice-agent-test-harness --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/voice-agent-test-harnessContext preview
The summary Claude sees to decide when to auto-load this skill.
Drive a fixed suite of spoken tests against a voice agent ("subject") from a co-located machine ("prober"), measure response latency / clarity / accuracy, diff against baseline, and report to the owner over Telegram.
Drive a fixed suite of spoken tests against a voice agent ("subject") from a co-located machine ("prober"), measure response latency / clarity / accuracy, diff against baseline, and report to the owner over Telegram.
**Design:** [docs/voice-agent-test-framework.md](../../docs/voice-agent-test-framework.md)
> **v1 (macOS).** Real audio path: TTS via `gemini-tts` + `afplay`, mic capture via `sox` `rec` (CoreAudio), voice-onset via numpy RMS, STT + judge via Gemini (Sutando-standard, `GEMINI_API_KEY`). Manual trigger; reports to owner only. Each prober-side component is tested; the full closed loop needs the second laptop speaking.
So that a half-SKIPPED suite is never mistaken for "mostly fine," here is exactly what executes through the real acoustic path now versus what is stubbed or excluded. A captured live run is committed at [`examples/run-2026-06-06.json`](examples/run-2026-06-06.json).
| Capability | Status today | |---|---| | Single-answer suite (`test_cases.yaml`, `core-v1`) — speak → capture → onset → Gemini STT → judge → score | ✅ **Wired.** Every row runs end-to-end on real audio; `pass` / `fail` / `partial` / `no_response` are all measured outcomes, not stubs. | | Latency / clarity / accuracy scoring + baseline diff + Telegram roll-up | ✅ **Wired** — computed on real captured turns. | | `timer` action test — real side-effect verify (waits, listens for the alarm) | ✅ **Wired.** | | Multi-turn workflow turns (`workflow_cases.yaml`, e.g. the developer code-change flow) | ⚠️ **Partial.** The spoken handling is captured and judged; remote side effects (branch/test/cleanup) are **not observable from the prober**, so these score wording only. | | Gmail / CRM workflow turns | ⛔ **Excluded** — unfinished test setup; omitted from results, not reported as failures. | | Daily auto-scheduling | ⛔ **Not wired** — manual trigger only. |
1. **Subject:** on laptop 2, start a normal Sutando voice session, mic open, speaker up. 2. **Prober:** on laptop 1 (this one), grant Terminal **Microphone** permission (System Settings → Privacy → Microphone), then:
cd ~/GitHub/sutando/skills/voice-agent-test-harness python3 scripts/run_suite.py --quick # --quick shortens the 2-min timer wait to 30s
3. The prober speaks each prompt; the subject replies; the prober measures, transcribes, judges, and prints the roll-up. Add `--deliver` to send the report to your Telegram.
Useful flags:
python3 scripts/run_suite.py --only arithmetic # one test by id python3 scripts/run_suite.py --dry-run # no audio/model; canned data (CI/sanity) python3 scripts/baseline.py --promote results/voice-test/<date>.json # set regression baseline
Tests with an `effect` block (the `timer`) verify the **real side effect**: after the verbal confirmation, the prober waits the timer duration and listens for the alarm actually firing. Confirmation without an observed effect downgrades to `partial`.
My AI Stand — Realtime by Day, Rewriting Itself by Night. Summon my AI superpower. Voice, vision, screen, meetings, calls when I'm engaged. Learns my patterns, ships its own code when I'm not. Runs across my Macs, interacts with people & their Stands.
Repo: sonichi/sutando
Streams what the agent is doing into the room, as rows the desktop client renders in an **events drawer** above the composer (collapsed: avatar, pulsing dots,…
Local Agent Registry — a standalone, dependency-free service that tracks running Claude Code (and other) agent instances. Agents self-register on startup and…
**Prefer the `ag2-space` MCP tools when they are connected and the room exposes them** — availability is per-room and per-actor, so check…
Deterministic final-answer normalizer — a last-step pass for any task that ends in a *precise* answer (a number, a short string, a comma-list). Applies the…
Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable…
Act back on the owner's Bee wearable — the TOOL half of the Bee integration (channels-vs-tools split). The Bee CHANNEL (ag2-sparrow's `sources/bee.py` watcher)…