/gemini-tts
Render text to mp3 via Google Gemini Flash TTS. Free-tier eligible (1500 req/day). Use for video narration, demo voiceovers, audio notes. Parallels openai-tts; default for make-viral-video.
$ npx -y skills add sonichi/sutando --skill gemini-tts --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/gemini-tts
Context preview
The summary Claude sees to decide when to auto-load this skill.
Render text to mp3 via Google Gemini Flash TTS. Free-tier eligible (1500 req/day). Use for video narration, demo voiceovers, audio notes. Parallels openai-tts; default for make-viral-video.
SKILL.md
gemini-tts.SKILL.mdname: gemini-tts
description: "Render text to mp3 via Google Gemini Flash TTS. Free-tier eligible (1500 req/day). Use for video narration, demo voiceovers, audio notes. Parallels openai-tts; default for make-viral-video."
user-invocable: true
Gemini TTS
Synthesize speech via Google's `gemini-2.5-flash-preview-tts` (or `-pro-tts` / `-lite-preview-tts` per env override). Reads `GEMINI_API_KEY` from `.env`.
This is offline synthesis — distinct from voice-agent's bidirectional Gemini Live audio. Same model family, different surface (POST text → get audio bytes back, no streaming).
**Usage**: `/gemini-tts [text]`
ARGUMENTS: $ARGUMENTS
Voices
`Aoede` (default — alto, neutral), `Charon` (baritone, news-anchor), `Kore` (mid, expressive), `Puck` (high, conversational). Per Lucy's 2026-05-09 testing: Aoede is the closest match to OpenAI's `sage`.
Audio tags for expression
Inline bracket tags like `[whispers]`, `[excitedly]`, `[slowly]` are interpreted as stylistic direction, not spoken literally. Empirically verified against `gemini-2.5-flash-preview-tts` (per PR #646 comment): `[whispers] hello` → 1.05s audio; `hello` alone → 1.01s. If the tag were spoken literally as 8 words, the clip would be ~5× longer.
bash "$SKILL_DIR/scripts/synthesize.sh" -- "[whispers] Pull request 691 has landed."
Model selection
Default: `gemini-2.5-flash-preview-tts` (free tier, 1500 req/day, $0 within quota).
Override via `GEMINI_TTS_MODEL` env var:
- `gemini-2.5-pro-tts` — paid, higher fidelity
- `gemini-2.5-flash-lite-preview-tts` — preview, faster
- `gemini-3.1-flash-tts-preview` — preview
Examples
bash "$SKILL_DIR/scripts/synthesize.sh" -- "Hello, this is Sutando."
bash "$SKILL_DIR/scripts/synthesize.sh" --voice Charon --out /tmp/intro.mp3 -- "Hi."
GEMINI_TTS_MODEL=gemini-2.5-pro-tts bash "$SKILL_DIR/scripts/synthesize.sh" -- "High-fidelity narration."
Default output path: `results/gemini-tts-{epoch}.mp3`.
Cost
Free tier: $0 within 1500 req/day quota. For our cadence (a few demos a day), stays free indefinitely. Paid (Flash): $0.50 / 1M input tokens + $10.00 / 1M output tokens.
Compared to OpenAI TTS (`gpt-4o-mini-tts`) at ~$0.02 per 60s: Gemini Flash is free-equivalent for typical demo workloads.
When to fall back to openai-tts
The `make-viral-video` skill auto-falls-back to OpenAI TTS when:
- Gemini API returns 4xx/5xx
- Gemini quota hit (429)
- `GEMINI_API_KEY` missing
- `TTS_PROVIDER=OPENAI` env override set
If Invoked As A Slash Command
If ARGUMENTS is empty, ask the user for the text. Otherwise:
bash "$SKILL_DIR/scripts/synthesize.sh" -- "$ARGUMENTS"
Read more
name: gemini-tts description: "Render text to mp3 via Google Gemini Flash TTS. Free-tier eligible (1500 req/day). Use for video narration, demo voiceovers, audio notes. Parallels openai-tts; default for make-viral-video." user-invocable: true
Gemini TTS
Synthesize speech via Google's `gemini-2.5-flash-preview-tts` (or `-pro-tts` / `-lite-preview-tts` per env override). Reads `GEMINI_API_KEY` from `.env`.
This is offline synthesis — distinct from voice-agent's bidirectional Gemini Live audio. Same model family, different surface (POST text → get audio bytes back, no streaming).
**Usage**: `/gemini-tts [text]`
ARGUMENTS: $ARGUMENTS
Voices
`Aoede` (default — alto, neutral), `Charon` (baritone, news-anchor), `Kore` (mid, expressive), `Puck` (high, conversational). Per Lucy's 2026-05-09 testing: Aoede is the closest match to OpenAI's `sage`.
Audio tags for expression
Inline bracket tags like `[whispers]`, `[excitedly]`, `[slowly]` are interpreted as stylistic direction, not spoken literally. Empirically verified against `gemini-2.5-flash-preview-tts` (per PR #646 comment): `[whispers] hello` → 1.05s audio; `hello` alone → 1.01s. If the tag were spoken literally as 8 words, the clip would be ~5× longer.
bash "$SKILL_DIR/scripts/synthesize.sh" -- "[whispers] Pull request 691 has landed."
Model selection
Default: `gemini-2.5-flash-preview-tts` (free tier, 1500 req/day, $0 within quota).
Override via `GEMINI_TTS_MODEL` env var:
- `gemini-2.5-pro-tts` — paid, higher fidelity
- `gemini-2.5-flash-lite-preview-tts` — preview, faster
- `gemini-3.1-flash-tts-preview` — preview
Examples
bash "$SKILL_DIR/scripts/synthesize.sh" -- "Hello, this is Sutando." bash "$SKILL_DIR/scripts/synthesize.sh" --voice Charon --out /tmp/intro.mp3 -- "Hi." GEMINI_TTS_MODEL=gemini-2.5-pro-tts bash "$SKILL_DIR/scripts/synthesize.sh" -- "High-fidelity narration."
Default output path: `results/gemini-tts-{epoch}.mp3`.
Cost
Free tier: $0 within 1500 req/day quota. For our cadence (a few demos a day), stays free indefinitely. Paid (Flash): $0.50 / 1M input tokens + $10.00 / 1M output tokens.
Compared to OpenAI TTS (`gpt-4o-mini-tts`) at ~$0.02 per 60s: Gemini Flash is free-equivalent for typical demo workloads.
When to fall back to openai-tts
The `make-viral-video` skill auto-falls-back to OpenAI TTS when:
- Gemini API returns 4xx/5xx
- Gemini quota hit (429)
- `GEMINI_API_KEY` missing
- `TTS_PROVIDER=OPENAI` env override set
If Invoked As A Slash Command
If ARGUMENTS is empty, ask the user for the text. Otherwise:
bash "$SKILL_DIR/scripts/synthesize.sh" -- "$ARGUMENTS"
My AI Stand — Realtime by Day, Rewriting Itself by Night. Summon my AI superpower. Voice, vision, screen, meetings, calls when I'm engaged. Learns my patterns, ships its own code when I'm not. Runs across my Macs, interacts with people & their Stands.
Repo: sonichi/sutando
Other skills on sutando.
- /agent-registry
Local Agent Registry — a standalone, dependency-free service that tracks running Claude Code (and other) agent instances. Agents self-register on startup and heartbeat while alive; the Electron overlay and Sutando dashboard read the live list. Use when you need to know which
Open skill - /agent-room-ops
**One skill, multiple tools.** Everything an agent does in a room beyond its task inbox lives here as a tool, so the parity capabilities are self-evidently *one collection* (not N scattered skills). Each tool is a thin **gateway-only** client verb sharing `_gateway.py`; the
Open skill - /audio-transcribe
Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable text in tasks.
Open skill - /bot2bot-post
Post a coordination message from this bot to the shared bot2bot channel — @-mentioning a specific peer via --to, auto-mentioning only in single-peer fleets, never guessing.
Open skill - /call-diagnostics
Analyze phone call observability data, detect problems, track them across calls, and recommend systematic repairs.
Open skill - /claude-codex
Bash wrapper around the local Codex CLI for non-interactive runs from inside Sutando (bridges, cron, scripts). For interactive code review or task hand-off from this Claude Code session, prefer the official `/codex:*` plugin commands; this skill is the file-bridge-compatible
Open skill

