Parakeet TDT 0.6B v3
NVIDIA’s speech-recognition model, 4-bit quantized to about 607 MB. It provides punctuation and capitalization across 25 languages.
Under the hood
From shortcut to finished text or natural audio, the speech work happens on your Mac. Here is what runs locally, what uses the network, and why macOS asks for access.
Dictation, cleanup, and Read Aloud run locally on Apple Silicon.
Used for setup downloads, licensing, updates, and optional diagnostics.
Microphone for dictation and Accessibility for shortcuts and insertion.
Vocaldi downloads its speech models once, then dictation, cleanup, and Read Aloud run locally on Apple Silicon. Core tools keep working offline.
Both models run through Apple’s MLX framework. Network access is used for the initial model download, licensing, updates, and optional diagnostics, not speech processing.
NVIDIA’s speech-recognition model, 4-bit quantized to about 607 MB. It provides punctuation and capitalization across 25 languages.
An open-weight text-to-speech model with 82 million parameters. It renders all 54 voices locally, without a cloud voice API.
macOS asks before either permission is granted. You can review or revoke both in System Settings.
Used only while you dictate. Vocaldi records the active session, transcribes it locally, then discards the temporary recording.
Lets Vocaldi listen for global shortcuts and insert your finished transcript into the field you already focused.