Dictation that stays on your Mac

ParrotFlow turns speech into text with small models that run on your Mac. Built-in steps need no account and no cloud key. Your audio stays on the machine.

What happens when you dictate

ParrotFlow does not stream. It records the whole clip, then runs these steps in order, all on your Mac:

  1. asr. Parakeet turns the clip into words.
  2. vad. Silero VAD checks the clip for speech.
  3. sentence_repair. A small language model takes out the full stops a pause put in the middle of a sentence.
  4. vocabulary. Names you taught ParrotFlow are matched by spelling and by sound, then checked against the sentence.
  5. pipeline. Your own steps run, in the order you list them.

Then the text is pasted into the app you are in.

Which models run on your Mac

ParrotFlow uses very small models. They repair the raw transcript and use your vocabulary in context. No large language model rewrites what you said.

Model What it does Size
Parakeet TDT 0.6B v3 Speech recognition, run by FluidAudio on the Neural Engine about 470 MB
Silero VAD Speech detection in the clip
Qwen3 0.6B Base, 4-bit Decides whether a full stop is real or a pause 320 MB
mmBERT-small Reads the place a word sits in, so a name is not written where the sentence wants a verb 269 MB
A sound model Turns a word into sounds, so a name is found even when the recogniser spells it a new way 81 MB

spaCy is also on the list of small models ParrotFlow uses. macOS provides two more pieces: its spell checker and its language tagger, which the vocabulary step asks about words and names.

The app ships no model weights. Each model is fetched once, then runs from your disk. Expect about 3 GB of disk and 1 GB of memory in all.

espeak-ng is an optional second way to turn words into sounds. It is a separate program, so you install it yourself:

brew install espeak-ng

What stays on your Mac

  • Your audio. A recording is only kept on disk if you turn on logging.audio, and then it stays in ~/.config/parrotflow/recordings.
  • Your transcripts. The record of each dictation, trace.jsonl, is a file beside your recordings.
  • How your voice says names. The voice/ folder beside config.yaml holds what this Mac has heard you say. It stays on the machine.
  • Your screen. When ParrotFlow reads the window you dictate into, it reads it on your Mac, at the moment you press the key.

The built-in steps need no account, no API key and no cloud service.

The network is used to download each model, once. ParrotFlow also asks GitHub’s release API once a day whether there is a new version. That call sends no account and nothing about you or what you dictate. Set updates.after_days to -1 and it never asks.

The one exception: prompt steps

A prompt step asks a language model to rewrite your text, for example to fix grammar or format an email. ParrotFlow has no such step until you add one. You also choose the model it runs on.

models:
  gemma:               # on your Mac, through Ollama
    api: ollama
    model: gemma4:e4b-mlx
    default: true      # what a transform runs on when it names no model
  gpt:                 # remote, for the harder jobs
    api: openai
    model: gpt-5.6-luna

A model run through Ollama stays on your Mac. A remote model receives the text of each step that names it. If a prompt also includes the screen text, that text goes to the remote model too.

By default a remote model’s key goes in the macOS keychain, not in config.yaml. The first time the app loads a config that names one, it asks for the key.

Spoken commands, the ones that start with “hey parrot”, also run on the models you list.

If a model call fails for any reason, the transcript comes back exactly as it arrived. No key, a timeout or no network costs you the rewrite, never the sentence.

Turning the built-in passes off

The two passes that repair the transcript are on by default. Each has a switch in config.yaml. With sentence_repair off, nothing is downloaded for it.

transcription:
  sentence_repair: {enabled: false}
  vocabulary: {enabled: false}

Read more