Local vs cloud dictation: what leaves your laptop?

Khoa Truong Nguyen AnhSources checked Updated

“Local dictation” can describe only the speech-to-text step. If another model rewrites the transcript in the cloud, your audio may stay on the laptop while the words leave it. A useful privacy check follows each stage instead of relying on a single badge or setting.

This guide uses vendor documentation checked on September 11, 2026, with Superwhisper and Wispr Flow as examples. It explains documented behavior and a configuration-check procedure, not a packet-capture audit or a certification of either product.

Ask where audio and text are processed

Superwhisper's sensitive-data guide describes two independently configurable stages: a voice model transcribes audio, then an optional language model processes the text. This creates several distinct configurations:

Processing choices, based on the selected mode; storage and destination checks are separate
Voice stageText stageWhat to check
LocalNoneTranscription runs on-device; still check history and the destination app.
LocalLocalBoth model stages run on-device if supported and downloaded.
LocalCloudAudio can remain local while the transcript is sent for cleanup.
CloudLocal or cloudAudio leaves the device for transcription, regardless of the next stage.

The guide says local language models are currently available on macOS, while Windows users can omit AI post-processing to keep the voice stage local. Do not carry a desktop capability over to mobile without checking that platform. The Superwhisper review covers model choice, pricing and the current free-tier documentation gap.

A training opt-out does not make transcription local

Wispr's training and cloud-sync guide explicitly says transcription happens in the cloud. It describes a renamed control on Mac and iOS: Improve the model for everyone.

According to that page, ON allows audio, transcripts and edits to be used for model improvement; OFF opts out. An earlier Privacy Mode ON preference now appears as Improve the model for everyone OFF. The displayed direction matters, so read the setting's explanation rather than assuming that “on” always means more privacy.

The same page says Android does not yet have a Data & Privacy section and does not give the same instructions for Windows. Other Wispr pages still use the old Privacy Mode label. Check your platform and app version instead of treating one set of controls as universal.

An opt-out is about data use. It does not mean speech is processed on-device, that network access is unnecessary, or that every local and synced copy has been deleted. For whether this trade-off fits your daily use, see the Wispr Flow review.

Check nearby text as well as the microphone

Some dictation features use more than the words you speak. Wispr's Context Awareness documentation describes nearby text, app information and IDE context, and provides a separate control for Context Awareness.

That page still uses older Privacy Mode terminology and lists platform-specific limitations. It also distinguishes standard password fields from custom or web fields. Do not treat automatic field exclusions as proof that every visible secret is excluded from context.

For a configuration review, record whether context reading is enabled and what the current documentation says is sent. Use an empty local document and non-sensitive sample sentences for experiments. If context is not needed for your task, evaluate with it disabled through the controls available on your platform.

Storage, retention and destination are separate questions

After processing, ask where the recording and transcript remain. A local history file, a cloud copy and a clipboard entry are different copies with different lifetimes. Turning off training does not answer the retention question, and deleting local history does not automatically delete a remote copy.

Superwhisper says transcript history is stored locally and describes FileSync for configuration without transcript or audio content. Its current billing guide elsewhere says modes and vocabulary are stored per device and cloud sync is planned. Because those descriptions are not fully aligned, verify whether your version actually exposes FileSync and what it covers; do not infer that every setting syncs.

Finally, inspect the destination. Text inserted into a local unsynced editor has a different next step from text submitted to an online coding assistant, chat app or cloud document. Local recognition does not extend its processing boundary to whatever app receives the transcript.

Interpret zero-retention claims within their scope

Superwhisper's security guide states that its integrated cloud providers operate under API agreements with zero-data-retention terms and no model training. These are vendor commitments described for those integrations, not independent observations made for this article.

The guide also distinguishes bring-your-own-key use, where your provider account and agreement apply, from the integrated service. Neither arrangement automatically supplies the same terms when you dictate into a separate consumer website or another business's application.

For a concrete decision, record the provider, account route, data involved and applicable retention statement. If the documentation or account setting is unclear, mark it unresolved rather than replacing it with a general “private” label.

Run a bounded offline check

Download the dictation data-flow worksheet. It separates what the vendor says, what you configured and what you observed.

  1. Record app, platform, voice model, text model and whether optional context or sync is enabled.
  2. Complete required model downloads and setup while online. Use a local text editor without document sync for the test.
  3. Disable the network connections you are testing, including Ethernet or a hotspot rather than only the Wi-Fi toggle.
  4. Dictate a new harmless sentence. Record whether a transcript appears, whether cleanup completes and any error.
  5. Reconnect and note any queued action you can observe. Record the trial's limits; do not claim you measured unobserved traffic.

If local transcription works but cleanup stalls, the selected cleanup stage may depend on a service. An offline failure can also be caused by missing models or incomplete setup, so it does not by itself prove audio was sent anywhere. Success shows the tested task can complete without the disconnected connections, not that the application never communicates when online.

Choose from the requirement you actually have

If audio must stay on-device, check the voice model first. If the text must also stay there, check cleanup, context, sync and the destination. If cloud processing is acceptable but training is not, inspect the relevant account control and agreement instead of searching only for an offline badge.

Use the Flow versus Superwhisper comparison for the product trade-offs, then keep a record of the exact configuration you chose. Recheck it when you change modes, devices or account settings.

After recording your processing requirements, use the developer voice-tool decision guide to narrow candidates by the job you need them to do.

Some links may earn a commission if you subscribe, at no extra cost to you. Commercial relationships do not determine the recommendation; we state the constraints and alternatives.