Skip to content
Development
Skill

/speech-recognition

Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or

From plugin
swift-ios-skills
98186 skills1 MCP
Install
$ npx -y skills add dpearson2699/swift-ios-skills --skill speech-recognition --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/speech-recognition

Context preview

The summary Claude sees to decide when to auto-load this skill.

Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or

SKILL.md

speech-recognition.SKILL.md
name: speech-recognition
description: "Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or adopting SpeechAnalyzer, SpeechTranscriber, DictationTranscriber, AssetInventory, and async result streams on iOS 26+."

Speech Recognition

Transcribe live and pre-recorded audio to text using Apple's Speech framework. Covers `SpeechAnalyzer` / `SpeechTranscriber` (iOS 26+) and `SFSpeechRecognizer` (iOS 10+) fallback guidance.

**Scope boundary:** Use this skill for speech-to-text recognition, speech authorization, microphone capture plumbing, and result handling. Hand off text analysis, language identification after transcription, sentiment, embeddings, and translation to `natural-language`; hand off audio playback UI to `avkit`; hand off summarization or generation over transcripts to `apple-on-device-ai`.

Contents

  • [SpeechAnalyzer Strategy (iOS 26+)](#speechanalyzer-strategy-ios-26)
  • [SFSpeechRecognizer Setup](#sfspeechrecognizer-setup)
  • [Authorization](#authorization)
  • [Live Microphone Transcription](#live-microphone-transcription)
  • [Pre-Recorded Audio File Recognition](#pre-recorded-audio-file-recognition)
  • [On-Device vs Server Recognition](#on-device-vs-server-recognition)
  • [Handling Results](#handling-results)
  • [Common Mistakes](#common-mistakes)
  • [Review Checklist](#review-checklist)
  • [References](#references)

SpeechAnalyzer Strategy (iOS 26+)

Use `SpeechAnalyzer` for modern iOS 26+ speech analysis, especially long-form recordings, live transcription, time-indexed transcripts, and fully on-device flows. Keep `SFSpeechRecognizer` for iOS 10+ deployment targets, server-backed locale coverage, or existing callback/delegate implementations.

Read [SpeechAnalyzer patterns](references/speechanalyzer-patterns.md) when implementing an iOS 26+ transcription pipeline, model asset handling, volatile results, or file/buffer examples.

SpeechAnalyzer setup checklist

1. Choose the module:

  • `SpeechTranscriber` for the newer general-purpose on-device model.
  • `DictationTranscriber` when `SpeechTranscriber` is unavailable for the

current device or locale and dictation-compatible support is acceptable.

  • `SpeechDetector` only in conjunction with a transcriber when voice

activity detection is worth the accuracy/power tradeoff. 2. Check support before creating the session:

  • `SpeechTranscriber.isAvailable`
  • `SpeechTranscriber.supportedLocale(equivalentTo:)`
  • `SpeechTranscriber.installedLocales` / `supportedLocales` when showing

language choices. 3. Pick a documented preset:

  • `.transcription` for basic accurate transcription.
  • `.progressiveTranscription` for live UI updates.
  • `.timeIndexedProgressiveTranscription` when playback highlighting needs

`audioTimeRange`. 4. Install required assets with `AssetInventory.assetInstallationRequest`. 5. Convert live audio buffers to `SpeechAnalyzer.bestAvailableAudioFormat(compatibleWith:)` before yielding `AnalyzerInput`. 6. Consume module results from their `AsyncSequence` in a separate task. 7. Finish explicitly with `finalizeAndFinish(through:)`, `finalizeAndFinishThroughEndOfInput()`, or `cancelAndFinishNow()`.

Do not use an `offlineTranscription` preset; Apple does not document one. Finishing an `AsyncStream` input sequence does not finish the analyzer session.

SFSpeechRecognizer Setup

Creating a recognizer with locale

import Speech

// Default locale (user's current language)
let recognizer = SFSpeechRecognizer()

// Specific locale
let recognizer = SFSpeechRecognizer(locale: Locale(identifier: "en-US"))

// Check if recognition is available for this locale
guard let recognizer, recognizer.isAvailable else {
    print("Speech recognition not available")
    return
}

Monitoring availability changes

final class SpeechManager: NSObject, SFSpeechRecognizerDelegate {
    private let recognizer = SFSpeechRecognizer()!

    override init() {
        super.init()
        recognizer.delegate = self
    }

    func speechRecognizer(
        _ speechRecognizer: SFSpeechRecognizer,
        availabilityDidChange available: Bool
    ) {
        // Update UI — disable record button when unavailable
    }
}

Authorization

Request **both** speech recognition and microphone permissions before starting live transcription. Add these keys to `Info.plist`:

  • `NSSpeechRecognitionUsageDescription`
  • `NSMicrophoneUsageDescription`
import Speech
import AVFoundation

func requestPermissions() async -> Bool {
    let speechStatus = await withCheckedContinuation { continuation in
        SFSpeechRecognizer.requestAuthorization { status in
            continuation.resume(returning: status)
        }
    }
    guard speechStatus == .authorized else { return false }

    let micStatus: Bool
    if #available(iOS 17, *) {
        micStatus = await AVAudioApplication.requestRecordPermission()
    } else {
        micStatus = await withCheckedContinuation { continuation in
            AVAudioSession.sharedInstance().requestRecordPermission { granted in
                continuation.resume(returning: granted)
            }
        }
    }
    return micStatus
}

Live Microphone Transcription

The standard pattern: `AVAudioEngine` captures microphone audio → buffers are appended to `SFSpeechAudioBufferRecognitionRequest` → results stream in.

import Speech
import AVFoundation

final class LiveTranscriber {
    private let recognizer = SFSpeechRecognizer(locale: Locale(identifier: "en-US"))!
    private let audioEngine = AVAudioEngine()
    private var recognitionRequest: SFSpeechAudioBufferRecognitionRequest?
    private var recognitionTask: SFSpeechRecognitionTask?

    func startTranscribing() throws {
        /
Read more
Ships withswift-ios-skills

86 agent skills optimized for iOS 26+ development with Swift 6.3 and modern Apple frameworks.

Get the whole plugin
Stats
981
Stars
50
Forks
Active
Maintenance
Python
Language
9d ago
Last commit
5mo ago
Created

Repo: dpearson2699/swift-ios-skills

Other skills on swift-ios-skills.