anachb
Austrian public transport (VOR AnachB) for all of Austria. Query real-time departures, search…
Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.
$ npx -y skills add mitsuhiko/agent-stuff --skill audio-transcription --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/audio-transcriptionContext preview
The summary Claude sees to decide when to auto-load this skill.
Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.
name: audio-transcription description: "Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio."
Use this skill whenever the user asks to transcribe an audio/video file, a Voice Memos export, dictation, lecture, meeting recording, or "bad audio".
1. **Preserve temporary inputs immediately.** Voice Memo share-sheet paths under `~/Library/Containers/com.apple.VoiceMemos/Data/tmp/.com.apple.uikit.itemprovider...` can disappear. Before probing or experimenting, copy the file to stable `/private/tmp/audio-transcription-inputs/`. 2. **Use cached local models, not cloud APIs.** Prefer MLX Whisper via `uvx --from mlx-whisper mlx_whisper`; Hugging Face models must be cached in `~/.cache/huggingface/hub/`. 3. **Force language when known.** For Armin's own dictations this is usually English with an Austrian/German accent, even when the filename is German. Do **not** infer language from filename alone. 4. **For bad audio, run a hallucination-resistant pass.** Use `--condition-on-previous-text False`, `--word-timestamps True`, and `--hallucination-silence-threshold 2`. 5. **Deliver a cleaned best-effort transcript.** Compare model output with timestamps/JSON, remove obvious Whisper loops, and mark uncertain spans as `[unclear]` rather than inventing words.
Run from this skill directory:
cd /Users/mitsuhiko/Development/agent-stuff/skills/audio-transcription ./transcribe-audio.py "/path/to/audio.m4a" --language en --quality balanced
The script:
Useful variants:
# Quick draft, fastest cached model ./transcribe-audio.py audio.m4a --language en --quality fast # Bad/important audio, slower full model ./transcribe-audio.py audio.m4a --language en --quality best \ --prompt "Armin Ronacher dictating about AI, data centers, Vienna, Donauinsel, shareholder value." # Auto language detection when language is genuinely unknown ./transcribe-audio.py audio.m4a --language auto --quality balanced
Default model IDs:
Pre-cache / refresh both models:
cd /Users/mitsuhiko/Development/agent-stuff/skills/audio-transcription ./precache-models.py
Verify cache manually:
find ~/.cache/huggingface/hub -maxdepth 1 -type d -name 'models--mlx-community--whisper-large-v3*' -print
If a model is already cached, `mlx_whisper` should say `Fetching 4 files: 100%` almost instantly.
If the helper script is not suitable, use this command directly:
mkdir -p /private/tmp/audio-transcriptions/manual uvx --from mlx-whisper mlx_whisper "/stable/copy/of/audio.m4a" \ --model mlx-community/whisper-large-v3-turbo \ --language en \ --condition-on-previous-text False \ --word-timestamps True \ --hallucination-silence-threshold 2 \ --output-format all \ --output-dir /private/tmp/audio-transcriptions/manual \ --output-name transcript \ --verbose False
For especially rough audio, replace the model with `mlx-community/whisper-large-v3-mlx`.
Inspect the generated `.txt` first, then the `.srt`/`.json` around suspicious areas.
Red flags that require rerun or cleanup:
When finalizing, lightly punctuate and paragraph the transcript, but do not over-edit uncertain content.
Armin's personal Pi Coding Agent package: reusable skills, extensions, prompt commands, themes, and a few supporting utilities that I use across projects. The package is published to npm as mitsupi.
Repo: mitsuhiko/agent-stuff
Austrian public transport (VOR AnachB) for all of Austria. Query real-time departures, search…
Search, read, and extract attachments from Apple Mail's local storage. Query emails by…
Reverse engineer binaries using Ghidra's headless analyzer. Decompile executables, extract…
Interact with GitHub using the `gh` CLI. Use `gh issue`, `gh pr`, `gh run`, and `gh api` for…
Access Google Workspace APIs (Drive, Docs, Calendar, Gmail, Sheets, Slides, Chat, People) via…