agent-activity
Streams what the agent is doing into the room, as rows the desktop client renders in an **events drawer** above the composer (collapsed: avatar, pulsing dots,…
Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
$ npx -y skills add sonichi/sutando --skill audio-transcribe --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/audio-transcribeContext preview
The summary Claude sees to decide when to auto-load this skill.
Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
name: audio-transcribe description: Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
Transcribes audio files (voice notes, clips) to text via Gemini 2.5-flash. Used by the Slack, Discord, and Telegram bridges to surface voice-note content in task bodies so the core agent can act on the words instead of a bare file path.
skills/audio-transcribe/scripts/transcribe.py <audio_file_path>
Reads the file, sends it inline (base64) to the Gemini 2.5-flash generateContent endpoint, and prints the transcript to stdout. Exits 0 on success, 1 on any failure (unsupported format, missing API key, network error, API error).
**Fail-open design.** Every bridge wraps the call in a helper that returns `None` on a non-zero exit. A failed transcription never blocks the task — the `[File attached: /path]` line still goes through so the agent can at least see a file was sent.
`.m4a` (Slack voice clips), `.mp3`, `.ogg`, `.oga`, `.opus`, `.wav`, `.webm`, `.aac`, `.flac`, `.mp4`
1. `GEMINI_API_KEY` or `GOOGLE_API_KEY` in the process environment 2. `<workspace>/.env` (resolved via `resolve_workspace()`) 3. `$CLAUDE_CONFIG_DIR/channels/slack/.env` 4. `$CLAUDE_CONFIG_DIR/channels/discord/.env` 5. `$CLAUDE_CONFIG_DIR/channels/telegram/.env`
Each bridge calls `_transcribe_via_skill(local_path)` after downloading a file. The helper locates the skill script relative to `src/` (or the app bundle), runs it, and returns the transcript string or `None`. The bridge then appends either:
Delete `skills/audio-transcribe/` — all three bridges fall back to `[File attached:]` automatically. Core services are unaffected.
My AI Stand — Realtime by Day, Rewriting Itself by Night. Summon my AI superpower. Voice, vision, screen, meetings, calls when I'm engaged. Learns my patterns, ships its own code when I'm not. Runs across my Macs, interacts with people & their Stands.
Repo: sonichi/sutando
Streams what the agent is doing into the room, as rows the desktop client renders in an **events drawer** above the composer (collapsed: avatar, pulsing dots,…
Local Agent Registry — a standalone, dependency-free service that tracks running Claude Code (and other) agent instances. Agents self-register on startup and…
**Prefer the `ag2-space` MCP tools when they are connected and the room exposes them** — availability is per-room and per-actor, so check…
Deterministic final-answer normalizer — a last-step pass for any task that ends in a *precise* answer (a number, a short string, a comma-list). Applies the…
Act back on the owner's Bee wearable — the TOOL half of the Bee integration (channels-vs-tools split). The Bee CHANNEL (ag2-sparrow's `sources/bee.py` watcher)…
Post a coordination message from this bot to the shared bot2bot channel — @-mentioning a specific peer via --to, auto-mentioning only in single-peer fleets,…