Format dictation with a local model

Run prompt steps on a local model with Ollama. Install it, pull the model, add it to config.yaml, and fix grammar or lay out an email at a key press.

What you will get

You say You press You get
their going to merge it tomorrow its almost done G They’re going to merge it tomorrow; it’s almost done.
hi Marie thanks for the notes yesterday I fixed the two typos and pushed the new version can you take one more look before Friday thanks Sam E the email below
Hi Marie,

Thanks for the notes yesterday.

I fixed the two typos and pushed the new version.

Can you take one more look before Friday?

Thanks,
Sam

Both are real output from gemma4:e4b-mlx through Ollama. A model can word things differently on another run.

What needs a model

Plain dictation needs no model. Speech is turned into text on your Mac, and the built-in steps need nothing running. A model is used for three things:

  • prompt: steps, such as the grammar chip.
  • Commands that start with “hey parrot”.
  • Spelling a name out loud.

A prompt step is the only kind of step that rewrites your words with a model, and only when you add one.

Install Ollama

Ollama runs language models on your Mac. Install it with Homebrew and start it:

brew install ollama && brew services start ollama

Or install the app from ollama.com/download. ParrotFlow’s setup guide asks for Ollama 0.22.0 or later. Older versions cannot run the gemma4 e-series models.

Check that it answers:

curl -s --max-time 3 http://localhost:11434/api/version

Pull the model

ollama pull gemma4:e4b-mlx

This is the model the ParrotFlow README, the installer and a new config name. Check that it is there:

ollama list

The download is several gigabytes. On the Mac this page was tested on, ollama list shows 8.8 GB for it.

Add it to config.yaml

Open ~/.config/parrotflow/config.yaml. From the menu bar, use Settings, then Edit Config. A config written by a recent install already has this block. If yours has no models: key, add it:

models:
  gemma:
    api: ollama
    model: gemma4:e4b-mlx
    keep_loaded: true
    default: true
  • gemma is a name you pick. Transforms use it to name the model.
  • api: ollama is the protocol. With no endpoint:, the app calls Ollama at http://localhost:11434.
  • model is the name as ollama list writes it. If you pulled another model, put its name here.
  • default: true makes this the model for any prompt that names none. With one model, it is the default anyway.
  • keep_loaded: true keeps the model in memory while the app runs. It is on unless you set it to false.

Ollama unloads an idle model after five minutes. The docs measure the next call at 6.7 s with a reload and 1.5 s with the model loaded. They also measure 9.6 GB of memory for gemma4:e4b while it is loaded. On a 16 GB Mac, set keep_loaded: false. On 32 GB, keep it on.

Save the file. The app reloads it at once.

Add a formatting prompt

A new config ships a grammar prompt with a chip on the pill and the key G. It is ready once the model is there.

This one lays out a dictated email. Add it under transforms::

transforms:
  - name: email
    description: lay out a dictated email
    display: Writing the email
    offer: true
    key: e
    prompt: |
      Lay out this dictated email. Put the greeting on its own line.
      Break the body into short paragraphs. Fix punctuation.
      Keep every word the speaker chose. Add nothing.
      Return only the email.

It names no model, so it runs on the default one. To pick one, add model: gemma.

Run it on demand

A prompt step costs a second or more, so run it when you want it.

  • A chip on the pill. offer: true puts the chip there after each dictation. key: e is its letter. Press E, or click the chip. The keys need the Input Monitoring permission. See Permissions.
  • A selection. Select text in any app, tap the hotkey, and press the letter.
  • Out loud. Tap, then hold the hotkey and say the transform’s name, such as “email”. say: on the transform adds other words that reach it.

To run it on every dictation in some apps only, add a pipeline step with a condition:

transcription:
  pipeline:
    # ...your other steps
    - transform: grammar
      app: /slack|outlook/

Then every dictation into Slack or Outlook waits for the model. Leave out app: and every dictation everywhere waits.

Check that it worked

Ask the app what it loaded:

/Applications/ParrotFlow.app/Contents/MacOS/ParrotFlow --check-config

If you installed with Homebrew, parrotflow runs the same binary.

Look for the model and the chips:

  · models            1 reachable
      gemma  ollama  gemma4:e4b-mlx  http://localhost:11434  reasoning off
  · offer             3 transform(s) on the pill, after Vocabulary
      grammar — G
      slack_mentions — S
      email — E

Then run a prompt on a sentence, with no microphone. The empty "" is the spoken instruction, which a chip does not have.

/Applications/ParrotFlow.app/Contents/MacOS/ParrotFlow --prompt grammar "" "their going to merge it tomorrow its almost done"
prompt:      grammar  (confirm on)
instruction: ""
in:          their going to merge it tomorrow its almost done
out:         They're going to merge it tomorrow; it's almost done.
             gemma4:e4b-mlx in 3.85s

The time depends on whether the model is already loaded. The same call took 0.34 s on a second run.

Last, dictate a sentence and press G.

What stays on your Mac

With api: ollama and no endpoint:, the text of a prompt step goes to Ollama on localhost. It does not leave your Mac.

A remote model changes that. Add one with api: openai or api: anthropic and it gets the text of every step that runs on it. --check-config says so for each one:

  · models: "gpt" sends text off this Mac — https://api.openai.com/v1, gpt-5.6-luna, key from …

Which text goes out depends on what runs on it:

  • A transform with model: gpt sends the text it rewrites.
  • If the remote model has default: true, every prompt that names no model uses it.
  • “hey parrot” commands use the default model unless router: under commands: names another one. So with a remote default, every command you say goes out too. The docs say to keep the router local.

Speech recognition and the built-in steps never call these models.

When the model is not running

If Ollama is stopped, the model is missing, or the call times out, the text stays exactly as it arrived. A model call gives up after 20 seconds unless you set timeout_seconds on the model.

Limits

  • A model can change words you meant to keep. Read the result.
  • The shipped grammar prompt scores 16 of 17 on its own test set with gemma4:e4b. The email prompt on this page has no test set. Score your own prompts with --eval before you put one in the pipeline.
  • The chips stay on the pill for a few seconds. After that, select the text and tap the hotkey.