OPTIONAL LOCAL CLEANUP
Improve transcription,
on your Mac.
An optional step that removes fillers and repetitions and improves English grammar after you dictate.
What VoxKey 0.6.0 does and the terms of the models it uses. Off by default. English only. Edits can change words, so review important text.
Turn it on when it helps.
Leave it off if…
You dictate in Spanish or other languages, you want your exact words kept, or you would rather not download up to 2.03 GB of extra models. Plain dictation works on its own.
Turn it on if…
You dictate in English and want fillers, repetitions, and small grammar slips fixed before the text lands, without sending it to a server.
How it works
Open Settings → Dictation and turn on Improve transcription. The models download once and then run inside the app. When the setting is on, three stages turn speech into the text delivered to your editor:
- Whisper transcribes. The speech model you selected converts your recording to text.
- S1-mini cleans up. Superwhisper’s editing model removes fillers and repetitions. It runs through llama.cpp with Metal support, linked directly into the app. Prompt lookup speeds up decoding by reusing matching text from the input.
- GECToR checks grammar. A Core ML grammar model applies restricted edits. Guards protect technical literals and limit word substitutions in this stage.
I integrated and validated these models; I did not create S1-mini or GECToR. Failed, cancelled, unsupported, or oversized improvement requests keep the original recognition result, and turning the setting off restores plain transcription.
Requirements and terms
Improvement requires English as the selected dictation language and up to 2.03 GB of optional model downloads. It pauses for Spanish, other languages, and Automatic; speech recognition keeps working in those languages without it.
The GECToR grammar model is for noncommercial use, under terms separate from VoxKey’s MIT-licensed app. S1-mini has its own terms. Review both in the behavior and validation summary before enabling the setting.
Models can still change names, meaning, or a correction you spoke on purpose. Review important text before sending it.
Try it with your own voice
Dictate the same passage twice into a local document, once with improvement off and once with it on. Use the same microphone, room, and speech model, and note your Mac chip, macOS version, and VoxKey version.
“So, um, the meeting is on Tuesday at three. Please, please send the updated document to María before lunch.”
“La reunión es el martes a las tres. Revisá el documento y avisame si podés enviarlo antes del almuerzo.”
These are suggested prompts, not recorded results.
- Check the words. Compare both versions for missing words, names, accents, and any change in meaning. Score punctuation separately.
- Check the wait. Time from your last spoken word to usable text, including releasing the trigger. Improvement adds its editing time on top of recognition.
- Check your apps. Try the editors you use each day with disposable text. If delivery cannot be confirmed, Safety Net keeps the last result for recovery; it is not a saved history.
See how it fits your writing.
Free, open source, and local. Apple Silicon · macOS 26+
This website measures visits, clicks, and page loading times with PostHog using an anonymous browser identifier. No session recordings. Privacy details.