/audio-transcribe
Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
$ npx -y skills add sonichi/sutando --skill audio-transcribe --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/audio-transcribe
Context preview
The summary Claude sees to decide when to auto-load this skill.
Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
SKILL.md
audio-transcribe.SKILL.mdname: audio-transcribe
description: Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
audio-transcribe
Transcribes audio files (voice notes, clips) to text via Gemini 2.5-flash. Used by the Slack, Discord, and Telegram bridges to surface voice-note content in task bodies so the core agent can act on the words instead of a bare file path.
How it works
skills/audio-transcribe/scripts/transcribe.py <audio_file_path>
Reads the file, sends it inline (base64) to the Gemini 2.5-flash generateContent endpoint, and prints the transcript to stdout. Exits 0 on success, 1 on any failure (unsupported format, missing API key, network error, API error).
**Fail-open design.** Every bridge wraps the call in a helper that returns `None` on a non-zero exit. A failed transcription never blocks the task — the `[File attached: /path]` line still goes through so the agent can at least see a file was sent.
Supported formats
`.m4a` (Slack voice clips), `.mp3`, `.ogg`, `.oga`, `.opus`, `.wav`, `.webm`, `.aac`, `.flac`, `.mp4`
Key resolution order
1. `GEMINI_API_KEY` or `GOOGLE_API_KEY` in the process environment 2. `<workspace>/.env` (resolved via `resolve_workspace()`) 3. `$CLAUDE_CONFIG_DIR/channels/slack/.env` 4. `$CLAUDE_CONFIG_DIR/channels/discord/.env` 5. `$CLAUDE_CONFIG_DIR/channels/telegram/.env`
Bridge integration
Each bridge calls `_transcribe_via_skill(local_path)` after downloading a file. The helper locates the skill script relative to `src/` (or the app bundle), runs it, and returns the transcript string or `None`. The bridge then appends either:
- `[Voice transcript: <text>]` — when transcription succeeds
- `[File attached: /path]` — when skill is absent or transcription fails
Removing the skill
Delete `skills/audio-transcribe/` — all three bridges fall back to `[File attached:]` automatically. Core services are unaffected.
Read more
name: audio-transcribe description: Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
audio-transcribe
Transcribes audio files (voice notes, clips) to text via Gemini 2.5-flash. Used by the Slack, Discord, and Telegram bridges to surface voice-note content in task bodies so the core agent can act on the words instead of a bare file path.
How it works
skills/audio-transcribe/scripts/transcribe.py <audio_file_path>
Reads the file, sends it inline (base64) to the Gemini 2.5-flash generateContent endpoint, and prints the transcript to stdout. Exits 0 on success, 1 on any failure (unsupported format, missing API key, network error, API error).
**Fail-open design.** Every bridge wraps the call in a helper that returns `None` on a non-zero exit. A failed transcription never blocks the task — the `[File attached: /path]` line still goes through so the agent can at least see a file was sent.
Supported formats
`.m4a` (Slack voice clips), `.mp3`, `.ogg`, `.oga`, `.opus`, `.wav`, `.webm`, `.aac`, `.flac`, `.mp4`
Key resolution order
1. `GEMINI_API_KEY` or `GOOGLE_API_KEY` in the process environment 2. `<workspace>/.env` (resolved via `resolve_workspace()`) 3. `$CLAUDE_CONFIG_DIR/channels/slack/.env` 4. `$CLAUDE_CONFIG_DIR/channels/discord/.env` 5. `$CLAUDE_CONFIG_DIR/channels/telegram/.env`
Bridge integration
Each bridge calls `_transcribe_via_skill(local_path)` after downloading a file. The helper locates the skill script relative to `src/` (or the app bundle), runs it, and returns the transcript string or `None`. The bridge then appends either:
- `[Voice transcript: <text>]` — when transcription succeeds
- `[File attached: /path]` — when skill is absent or transcription fails
Removing the skill
Delete `skills/audio-transcribe/` — all three bridges fall back to `[File attached:]` automatically. Core services are unaffected.
My AI Stand — Realtime by Day, Rewriting Itself by Night. Summon my AI superpower. Voice, vision, screen, meetings, calls when I'm engaged. Learns my patterns, ships its own code when I'm not. Runs across my Macs, interacts with people & their Stands.
Repo: sonichi/sutando
Other skills on sutando.
- /agent-registry
Local Agent Registry — a standalone, dependency-free service that tracks running Claude Code (and other) agent instances. Agents self-register on startup and heartbeat while alive; the Electron overlay and Sutando dashboard read the live list. Use when you need to know which
Open skill - /agent-room-ops
**One skill, multiple tools.** Everything an agent does in a room beyond its task inbox lives here as a tool, so the parity capabilities are self-evidently *one collection* (not N scattered skills). Each tool is a thin **gateway-only** client verb sharing `_gateway.py`; the
Open skill - /bot2bot-post
Post a coordination message from this bot to the shared bot2bot channel — @-mentioning a specific peer via --to, auto-mentioning only in single-peer fleets, never guessing.
Open skill - /call-diagnostics
Analyze phone call observability data, detect problems, track them across calls, and recommend systematic repairs.
Open skill - /claude-codex
Bash wrapper around the local Codex CLI for non-interactive runs from inside Sutando (bridges, cron, scripts). For interactive code review or task hand-off from this Claude Code session, prefer the official `/codex:*` plugin commands; this skill is the file-bridge-compatible
Open skill - /claude-gemini
Use the local Gemini CLI from Claude Code with the user's existing Gemini authentication or API configuration. Use for large-context repo scans, multimodal analysis, second-opinion planning, or structured Gemini runs in the current workspace.
Open skill

