Speak has moved into Listen. Dictation is now part of Listen, which also records and transcribes your meetings. Same shortcut, same local speech model, nothing uploaded. This page stays up and Speak keeps working, but it is no longer updated.

Speak

Turns your voice into text, wherever you are typing.

Entirely on your Mac. No account, nothing uploaded.

For Apple silicon Macs (M1 or later).

fn to start, speak, press again to stop

What it does

Press a shortcut, say what you mean, press it again. Your words appear in whatever app you are already in: an email, Slack, a search bar or a code comment.

The menu bar shows a red dot while Speak is listening. Your words stay on the clipboard too, so ⌘V works wherever you need it.

Tidied up, if you want it

Both off until you turn them on.

um so i think we should uh basically try the the parakeet model and see if it you know works better

So I think we should basically try the Parakeet model and see if it works better.

AI polish

Punctuation, filler words, false starts and paragraphs, using the model built into macOS. Needs macOS 26. If it fails or takes too long you get the plain transcript, so a dictation is never lost.

Your dictionary

Teach it the names it keeps getting wrong, and set exact replacements that run on every transcript. The replacements need no model, so they work on any supported macOS. Import brings across a dictionary you already built elsewhere.

Your voice never leaves your Mac

Not as a policy. As an architecture.

Speak downloads a speech model once, then runs it on your own hardware using Apple’s MLX framework. Transcription never touches the network. There is no account because there is no server that could hold your recordings, even in principle.

Tidying up is local too. The optional polish pass uses the model built into macOS, running on your Mac like everything else here, so a polished transcript has still never left the machine.

The trade is worth stating plainly: when a local model gets a word wrong, it is wrong. There is no cloud fallback to catch it. A dictionary of your own names and jargon is the fix, and it is built in.

What you need

How fast

About 35 ms to turn a six-second sentence into text, measured warm on an M4 Max. In practice the words are on screen before you have finished letting go of the key. Most people speak at around 150 words a minute and type at around 40.

Its companion

For whole conversations, there is Listen. It records your meetings on your Mac, writes them down and works out who said what.

Other ways to install

Prefer Homebrew?

brew install --cask mugoosse/tap/speak

Apple Silicon, macOS 14 or later. The optional AI polish needs macOS 26.

The speech model downloads on first run only if you choose Parakeet. If you already use Listen with the same model, it is already on your disk, and Speak will find it rather than fetching it twice.