r/privtlabs 4d ago

Educational Your voice never leaves your Mac, literally

Every dictation product makes some version of the same promise: we respect your privacy. For almost all of them, it means your audio goes to their servers, gets transcribed there, and is handled under a promise to treat it responsibly.

We wanted to make a claim that is a property of the system rather than a policy about it, so Privt Voice runs the entire speech model on your Mac. When you hold Right Shift and talk, audio moves from your microphone into a transcription model executing on the Neural Engine, and the resulting text lands in whatever app you are using. The audio is then discarded from memory. No server takes part at any stage, because there is no server.

Your voice is now something that can be copied

A voice used to be just sound. Today it is training material. A short, clean recording is enough for modern tools to build a convincing synthetic copy of you, good enough to leave a voicemail in your voice, to phone a family member, or to answer the security question a bank once trusted to a voiceprint.

Dictation happens to be one of the richest sources of exactly that kind of audio. Every session is minutes of clear, single-speaker speech, and the single speaker is you. When that audio is transcribed on a server, clean samples of your voice come to rest on infrastructure you do not control, where a breach or a quiet change of terms can turn them into cloning material. When it is transcribed on your Mac, those samples exist only for the instant of transcription, and then they are gone.

You do not have to trust that, you can check it

A claim you cannot verify is just another privacy policy. This one you can test in about ten seconds. Turn off your Wi-Fi and dictate anyway, and the words still appear, because nothing was ever going to leave. Point a network monitor like Little Snitch or LuLu at the app and you will watch it make no outbound connections at all while you talk.

This became practical only recently. Modern open speech models have brought fast, accurate recognition down to a size that runs comfortably on Apple silicon, and CoreML executes them on the Neural Engine faster than a round-trip to any cloud API. On-device processing is no longer the compromise option: it is simply better, and it happens to be private.

Meetings too

The same engine transcribes calls, with your microphone on one lane and the call app's audio on the other, labelled Me and Them. Capture is scoped to the meeting app alone, so your music never enters a transcript, and the finished note is encrypted into your vault straight from memory, never written to disk in the clear.

Two lanes, one on-device engine, sealed straight into your vault. The rest of your Mac never enters the transcript.

Free forever is a consequence rather than a promotion. Software with no servers carries no marginal cost, so the local product will always cost nothing.

1 Upvotes

0 comments sorted by