acestep
AI music generation with ACE-Step 1.5 — background music, vocal tracks, covers, stem extraction for video production. Use when generating music, soundtracks,…
Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content in OpenMontage. Optional cloud STT provider — preferred when AZURE_SPEECH_KEY is configured; the local
$ npx -y skills add calesthio/OpenMontage --skill azure-speech-to-text --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/azure-speech-to-textContext preview
The summary Claude sees to decide when to auto-load this skill.
Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content in OpenMontage. Optional cloud STT provider — preferred when AZURE_SPEECH_KEY is configured; the local
name: azure-speech-to-text
description: Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content in OpenMontage. Optional cloud STT provider — preferred when AZURE_SPEECH_KEY is configured; the local faster-whisper `transcriber` is the default offline path.
license: MIT
compatibility: Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).
metadata: {"openclaw": {"requires": {"env": ["AZURE_SPEECH_KEY", "AZURE_SPEECH_REGION"]}, "primaryEnv": "AZURE_SPEECH_KEY"}}Transcribe audio to text with **Azure Fast Transcription** — synchronous, word-level timestamps, speaker diarization, and multi-language identification. In OpenMontage this is exposed through the `azure_stt` tool (`capability=analysis`, `provider=azure`). It is an **optional cloud STT provider** — when `AZURE_SPEECH_KEY` is configured, prefer it for cloud transcription. The local `transcriber` tool (faster-whisper) remains the **default offline path** and the fallback when Azure is unavailable.
> Docs: [Fast Transcription](https://learn.microsoft.com/azure/ai-services/speech-service/fast-transcription-create) · [Speech service overview](https://learn.microsoft.com/azure/ai-services/speech-service/spx-overview)
Azure exposes three STT surfaces. OpenMontage uses **Fast Transcription** because the pipeline transcribes **local audio files**:
| Surface | Input | Latency | Needs | |---------|-------|---------|-------| | **Fast Transcription** (used here) | local file, multipart POST | synchronous, sub-real-time | key + region | | Batch Transcription | audio at a URL (Blob + SAS) | async job + polling | Blob storage plumbing | | Speech SDK (`spx`) | mic / stream / file | streaming | native `azure-cognitiveservices-speech` package |
Fast Transcription needs no Blob storage, no SAS URLs, and no native SDK — just `requests` and the two env vars.
Create a **Speech** resource in the [Azure portal](https://portal.azure.com); copy the key and region from its **Keys and Endpoint** page.
export AZURE_SPEECH_KEY=your_speech_resource_key export AZURE_SPEECH_REGION=eastus # your resource's region # export AZURE_SPEECH_ENDPOINT=https://... # optional: overrides region
`azure_stt` reports `AVAILABLE` once `AZURE_SPEECH_KEY` plus either `AZURE_SPEECH_REGION` or `AZURE_SPEECH_ENDPOINT` are set.
Prefer `azure_stt` over `transcriber` unless the run must be offline. Its output matches the `transcriber` schema exactly, so it is a drop-in for `subtitle_gen` and any stage that consumes a transcript.
from tools.tool_registry import registry
registry.discover()
stt = registry._tools["azure_stt"]
result = stt.execute({
"input_path": "projects/my-video/assets/audio/narration.mp3",
# "language": "en", # ISO 639-1 or BCP-47 ("en-US"); omit for auto-ID
# "diarize": True, # speaker labels, no HuggingFace token needed
# "max_speakers": 4,
"output_dir": "projects/my-video/artifacts",
})
if result.success:
segs = result.data["segments"] # [{id,start,end,text,words:[...]}]
words = result.data["word_timestamps"] # flat [{word,start,end,probability}]If `azure_stt` is unavailable (no key) or errors, fall back to `transcriber` (local whisper) — its `execute` signature and output are identical.
when you know the language; it is faster and more accurate than auto-ID.
identification across this shortlist. Narrow it to the languages you actually expect; a huge list slows detection and invites misclassification.
podcasts). Set `max_speakers` to the real upper bound.
The raw Azure response (`phrases[]` with `offsetMilliseconds` / `words[]`) is converted to seconds and the OpenMontage transcript schema:
{
"segments": [
{"id": 0, "start": 0.0, "end": 2.4, "text": "Hello world",
"speaker": 1,
"words": [{"word": "Hello", "start": 0.0, "end": 0.5, "probability": 0.98}]}
],
"word_timestamps": [{"word": "Hello", "start": 0.0, "end": 0.5, "probability": 0.98}],
"language": "en-US",
"duration_seconds": 2.4,
"provider": "azure"
}Note: Fast Transcription has no *per-word* confidence, so each word carries the **phrase** confidence in `probability`.
jobs, use Azure Batch Transcription instead.
you only need speech — smaller upload, same result.
the first and last cues against the source audio.
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
Repo: calesthio/OpenMontage
AI music generation with ACE-Step 1.5 — background music, vocal tracks, covers, stem extraction for video production. Use when generating music, soundtracks,…
Build voice AI agents with ElevenLabs. Use when creating voice assistants, customer service bots, interactive voice characters, or any real-time voice…
Generate AI videos from text prompts using multiple provider gateways. Use when: (1) Generating videos from text descriptions, (2) Creating AI-generated video…
Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API. Use when: (1) Choosing a specific…
Generate neural narration audio using Azure AI Speech (REST text-to-speech). Use when synthesizing voiceovers or narration in OpenMontage. Optional cloud TTS…
Render Mermaid diagrams as SVG and PNG using the Beautiful Mermaid library. Use when the user asks to render a Mermaid diagram.