What happens when you dictate
ParrotFlow does not stream. It records the whole clip, then runs these steps in order, all on your Mac:
asr. Parakeet turns the clip into words.vad. Silero VAD checks the clip for speech.sentence_repair. A small language model takes out the full stops a pause put in the middle of a sentence.vocabulary. Names you taught ParrotFlow are matched by spelling and by sound, then checked against the sentence.pipeline. Your own steps run, in the order you list them.
Then the text is pasted into the app you are in.
Which models run on your Mac
ParrotFlow uses very small models. They repair the raw transcript and use your vocabulary in context. No large language model rewrites what you said.
| Model | What it does | Size |
|---|---|---|
| Parakeet TDT 0.6B v3 | Speech recognition, run by FluidAudio on the Neural Engine | about 470 MB |
| Silero VAD | Speech detection in the clip | |
| Qwen3 0.6B Base, 4-bit | Decides whether a full stop is real or a pause | 320 MB |
| mmBERT-small | Reads the place a word sits in, so a name is not written where the sentence wants a verb | 269 MB |
| A sound model | Turns a word into sounds, so a name is found even when the recogniser spells it a new way | 81 MB |
spaCy is also on the list of small models ParrotFlow uses. macOS provides two more pieces: its spell checker and its language tagger, which the vocabulary step asks about words and names.
The app ships no model weights. Each model is fetched once, then runs from your disk. Expect about 3 GB of disk and 1 GB of memory in all.
espeak-ng is an optional second way to turn words into sounds. It is a separate program, so you install it yourself:
brew install espeak-ng
What stays on your Mac
- Your audio. A recording is only kept on disk if you turn on
logging.audio, and then it stays in~/.config/parrotflow/recordings. - Your transcripts. The record of each dictation,
trace.jsonl, is a file beside your recordings. - How your voice says names. The
voice/folder besideconfig.yamlholds what this Mac has heard you say. It stays on the machine. - Your screen. When ParrotFlow reads the window you dictate into, it reads it on your Mac, at the moment you press the key.
The built-in steps need no account, no API key and no cloud service.
The network is used to download each model, once. ParrotFlow also asks GitHub’s release API once a day whether there is a new version. That call sends no account and nothing about you or what you dictate. Set updates.after_days to -1 and it never asks.
The one exception: prompt steps
A prompt step asks a language model to rewrite your text, for example to fix grammar or format an email. ParrotFlow has no such step until you add one. You also choose the model it runs on.
models:
gemma: # on your Mac, through Ollama
api: ollama
model: gemma4:e4b-mlx
default: true # what a transform runs on when it names no model
gpt: # remote, for the harder jobs
api: openai
model: gpt-5.6-luna
A model run through Ollama stays on your Mac. A remote model receives the text of each step that names it. If a prompt also includes the screen text, that text goes to the remote model too.
By default a remote model’s key goes in the macOS keychain, not in config.yaml. The first time the app loads a config that names one, it asks for the key.
Spoken commands, the ones that start with “hey parrot”, also run on the models you list.
If a model call fails for any reason, the transcript comes back exactly as it arrived. No key, a timeout or no network costs you the rewrite, never the sentence.
Turning the built-in passes off
The two passes that repair the transcript are on by default. Each has a switch in config.yaml. With sentence_repair off, nothing is downloaded for it.
transcription:
sentence_repair: {enabled: false}
vocabulary: {enabled: false}
Read more
- Spelling from what is on your screen
- Rules, scripts and prompts for your dictation
- How it works, in the repository
- Transcription: the speech model, its limits, and the passes that cover them