AI & Technology

Dictation That Never Leaves Your Mac

If you dictate client notes, medical details, or unreleased work, the first question is where the audio goes. Here is what on-device transcription actually means, and what it does not cover.

IJ

Isaac Juracich

September 16, 2026 · 6 min read

Share

You are between appointments, talking through a note that has a client's name in it, a date, and a detail neither of you wants sitting on somebody else's server. Talking it through beats typing it out one-handed. The convenience is obvious. The question is where the audio goes.

For Voice, our Mac dictation tool, the answer is that it stays on your Mac. Whisper turns your speech into text on the machine itself. That one fact changes what you can reasonably dictate, so it is worth understanding precisely what it covers and what it does not.

What "on-device" actually means

On-device transcription means the microphone audio is captured by the app, handed to a speech model running on your own processor, and turned into text without ever being sent anywhere. There is no upload, no queue on a vendor's server, no retention window to read about in a policy page, and no second company in the chain.

Cloud dictation works the other way. Audio, or a live stream of it, goes to a server, gets recognized there, and text comes back. That is not automatically wrong. Plenty of good tools work this way. But it changes the question you are answering. Instead of "do I trust my own Mac," you are now asking whether you trust a company, its subprocessors, its retention defaults, its handling of a future breach, and its policy as it exists three product cycles from now. Those are all answerable questions. They are just a lot more work than the local answer.

The model download is the tell

Local speech recognition needs a model on disk, and models are large. In Voice, Whisper Large v3 Turbo is the default at 574 MB and handles any language. Small and Base are lighter English models for machines or moments where you want less weight. You download the model once, and after that transcription works offline.

That last part is the honest test, and you can run it yourself on any dictation tool that claims to be local. Turn off Wi-Fi. Dictate a paragraph. If you get a transcript, the audio was never going anywhere. If you get an error or an endless spinner, something was leaving the machine. The test takes a moment and tells you more than a marketing page will.

What still leaves if you turn on cleanup

Voice has an optional second step where Claude tidies the transcript: filler words out, punctuation in. When that step is on, the text transcript is sent to Anthropic using your own API key. The audio is never uploaded. Only the words.

Be precise about what that means, because it is easy to feel protected by the wrong part. The audio staying local protects your voice, your background noise, and anyone else audible in the room. It does not sanitize what you said. If you spoke a client's name, the name is in the transcript, and the transcript is what gets sent. On-device transcription is not a redaction step.

So the useful way to think about it is per category of work, not as one global setting you flip once:

  • Sensitive dictation: turn cleanup off and take the raw words. You get a transcript that never left the Mac at any stage, and you clean it up yourself in your editor.
  • Ordinary work: email, internal notes, a first draft of a proposal. Turn cleanup on and let the mechanics get fixed before you read it.
  • Anything you are unsure about: treat it as sensitive. The cost of being wrong is asymmetric, and the raw transcript is still perfectly usable.

Using your own API key matters here too, and not just for billing. The account relationship is directly between you and Anthropic. You read those terms yourself and hold that key yourself, rather than inheriting whatever arrangement a middle layer negotiated and passing your text through it.

Questions worth asking any dictation tool

None of this is specific to our app. If you are evaluating something else, these five questions cover most of the ground:

  1. Where does the audio go? If the answer is not a flat "nowhere," you are on the cloud path and should evaluate it as such.
  2. Where does the text go? This is a separate question from the audio, and vendors sometimes answer only the first one.
  3. Whose account pays for the AI step? Your own key means your own terms. A bundled key means someone else's.
  4. Does it work with the network off? The test described above.
  5. Can you see the text before it moves? A tool that shows you the result before anything is copied lets you catch a transcript you did not intend to produce.

The clipboard belongs in the same conversation

Privacy usually gets discussed as a network question, but on a desktop the more common leak is local and boring. Text ends up in the wrong window. Voice never types into your apps and nothing reaches the clipboard until you have looked at the result and chosen Copy. The panel shows the cleaned version and the original side by side, with Discard and Redo sitting next to the copy button.

We built that as an ergonomics decision, because reviewing beats correcting. It turns out to be a containment decision too. A tool that injects text straight into whatever window has focus is one mistimed app switch away from putting a sentence about a client into a group chat.

What on-device does not buy you

Three limits worth stating plainly, because a tool that oversells its privacy story is worse than one that has none.

It does not secure your Mac. If the disk is not encrypted, the screen does not lock, or someone else has physical access, local storage is not a benefit. On-device processing raises the value of basic machine hygiene rather than replacing it.

It does not make the transcript accurate. A local model can mishear a surname or a drug name exactly as a cloud model can. Reviewing before you paste is a correctness step, not just a privacy one.

It does not settle compliance. If you work under privilege rules, a records retention policy, or a healthcare privacy framework, "the audio never leaves the device" is one input to that analysis, not the conclusion. The people who sign off on your policies should hear how the tool works and decide, and they will usually want to know about the optional cleanup step specifically.

The takeaway

On-device transcription is not a feature you should pay a premium for. It is the default that a dictation tool should have to argue its way out of, because the alternative asks you to extend trust you gain nothing from extending. The useful judgment is not "local good, cloud bad." It is knowing exactly which stage of the pipeline your words are sitting in, and being able to turn off the one stage that moves them.

Voice is a Mac app we distribute directly, on Apple Silicon with macOS 15 or later. There is no public installer yet, so if you want to try it, get in touch and we will sort out access.

Filed Under

PrivacyMac AppsAI ToolsDictation
Share
IJ

Written by

Isaac Juracich

Full-stack engineer building production software for businesses that need it done right. Based in La Crosse, WI.

More about Isaac

Ready to Build?

Hire a web developer who ships

If this post resonated, we'd love to hear what you're working on. Tell us your project and we'll reply within 24 hours with a fixed scope and price.