Some dictation apps process your voice entirely on your device. Others upload it to a server first. Some retain that audio and use it to train AI models; others discard it immediately or never collect it. This comparison reads each vendor's own privacy page — not marketing copy — to show which is which.
Checked against each vendor's own page on 2026-07-19: Superwhisper and MacWhisper process audio entirely on-device, with no cloud option. Hovor and Apple Dictation default to on-device processing for part of their flow, with cloud processing for the rest. Google's Gboard defaults to on-device for basic voice typing. Wispr Flow states transcription "always happens in the cloud." Aqua Voice, Otter.ai, and Dragon Professional's pages did not document an on-device mode.
Every dictation app needs a speech-recognition model to turn audio into text. That model either runs locally, on the same chip capturing your voice, or on a remote server that receives your audio over the network. This single fact determines whether a company's infrastructure ever touches your voice, and whether that audio can be retained or used to train future models.
Encryption in transit and a well-written privacy policy reduce risk in the cloud model, but they do not eliminate the fact that your audio reached someone else's hardware — which is exactly the fact this comparison is built around.
The table below reflects only what each vendor states on its own privacy policy or product page, fetched 2026-07-19. Where a vendor's page did not address a question, the cell reads "not stated" rather than a guess. Source links point to the exact page checked.
| App | Where audio is processed | Cloud provider named | Audio retained? | Used for AI training? | On-device mode |
|---|---|---|---|---|---|
| Hovor (cloud mode, default) | OpenAI; some workflows route to Deepgram or ElevenLabs instead | Yes — OpenAI (or Deepgram/ElevenLabs) | No — deleted immediately after transcription | Not stated for the cloud path | Yes — optional (free tier, capped; unlimited with Local Unlock or Pro) |
| Hovor (on-device mode) | Your device (Whisper or Parakeet) | On-device when the model is loaded; server fallback while it is loading | On-device when the model is loaded; server fallback while it is loading | N/A | Yes — this is the mode |
| Superwhisper | Your device | N/A — no cloud processing | "Data is not retained on Superwhisper servers" | "Not being used for training AI models" | Yes — entire app is on-device |
| MacWhisper | Your device | N/A — "no data leaves your device" | Not stated (implied local-only) | Not stated | Yes — entire app is on-device |
| Apple Dictation | Device or Apple servers — your device's Keyboard Settings discloses which applies | Not named — "Apple servers" | Not stored, unless you opt in to "Improve Siri and Dictation" | Only if opted in (retained up to 2 years) | Yes, device-dependent — disclosed in Keyboard Settings |
| Google Gboard voice typing | Device, for standard voice typing; Google servers for "Fix it" / detailed edits | Not named | "No text or audio data is stored on Google servers" (detailed-edits feature) | Not stated | Yes — default for basic voice typing |
| Wispr Flow | Cloud — "Transcription always happens in the cloud" | Not named — "secure cloud servers" | Conditional — none if Privacy Mode is on; otherwise processed and may be used per below | Conditional — excluded under Privacy Mode; otherwise "may be used to improve Flow's features and AI models" | Not stated |
| Aqua Voice | Cloud (implied by server-side transcript storage; not explicitly named) | Not named — "hosting services" | Conditional — stored on servers if Privacy Mode is off; not collected if on | Not stated (states data may be used "to improve the product") | Not stated |
| Otter.ai | Cloud | Yes — Amazon Web Services (for compute/storage generally) | "For as long as necessary" — no fixed duration stated | Yes — explicitly, on "de-identified audio recordings and transcriptions" | Not stated |
| Dragon Professional (Nuance / Microsoft) | Not stated on the pages we checked | Not stated | Not stated | Not stated | Not stated on vendor's own page (see note below) |
Note on Dragon: Nuance's Dragon Professional product page describes features and business use cases but does not state, in the pages checked, whether speech recognition runs locally or via the cloud, nor does it address audio retention or training use. We are not asserting Dragon does or does not process audio locally — only that its own pages did not document it as of this check, so the cells above read "not stated" rather than repeating third-party claims.
Hovor's default free and Pro tiers use a cloud workflow: your audio streams to Hovor's server, which forwards it to OpenAI (gpt-4o-mini-transcribe) for transcription, and in some workflow configurations to Deepgram or ElevenLabs instead. The resulting text may be sent to OpenAI again to clean up grammar and punctuation. Per Hovor's privacy policy, audio is never stored on Hovor's server — it is deleted immediately once transcription completes.
On-device speech recognition — Whisper or Parakeet, selected automatically based on device capability (Parakeet needs roughly 4 GB of iOS device RAM; Whisper needs roughly 6 GB) — already runs on the free tier under its own weekly quota; Local Unlock ($49.99 one-time, or 1,999 UAH, family-shareable) or a Pro subscription removes that cap. When the on-device model is loaded, transcription runs entirely on your device and the audio is not uploaded. If the model is still loading — after an app relaunch, or while it is still downloading — that recording falls back to server transcription. Text clean-up in this mode runs via Apple's on-device Foundation models on iOS 26 and later, or is skipped if unavailable; it never uses the local EuroLLM/MLX formatting code that exists in Hovor's source tree but does not function on any platform today. BYOK — included with Pro, or available as a separate one-time BYOK Unlock ($24.99, or 999 UAH) — lets you route text-cleanup or tone formatting to your own OpenAI, Anthropic, or custom-endpoint API key, stored in Keychain (service app.hovor.byok) and sent straight from your device to the provider you chose — never proxied through Hovor's server.
"Audio retention" answers one question: after your speech becomes text, does a copy of the recording still exist on a server? Of the eight competitor pages checked, only three gave a direct, unconditional answer: Superwhisper and MacWhisper (no retention — audio never reaches a server), and Otter.ai (retained "for as long as necessary," with training use stated explicitly).
The remaining five made retention conditional on a setting (Wispr Flow, Aqua Voice), described it as opt-in (Apple), scoped it to one feature (Google Gboard's "Fix it"), or did not address it on the pages we could locate (Dragon). A missing answer is itself a data point: it means the retention question has to be answered by testing the product or reading a longer legal document, not by checking the page most users would actually read.
On-device means the speech-recognition model runs on your phone or computer's own chip. When that model is loaded and active, the audio it processes stays on the device: there is no network request carrying your voice to intercept, log, or retain, because none is made. This guarantee holds only once the on-device model is actually loaded — some apps, including Hovor, can fall back to server transcription if the local model is still loading when a recording starts. Cloud dictation instead uploads your audio to a remote server on every request, where a company you don't control can, in principle, access it before returning a transcript. With cloud dictation, the vendor's own privacy policy is the only thing standing between your voice and their servers — which is why this article reads those policies directly instead of trusting marketing copy.
Based on each vendor's own page, checked 2026-07-19: Superwhisper and MacWhisper both state all processing happens locally, with no audio leaving the device. Hovor offers on-device processing via Whisper or Parakeet as an optional mode, available free-tier under a weekly quota or unlimited with Local Unlock ($49.99/1,999 UAH one-time) or Pro, alongside a default cloud mode. Apple's own page states your device discloses in Keyboard Settings whether Dictation is processed on-device or sent to Apple servers, and Google's Gboard defaults to on-device for standard voice typing, with documented exceptions ("Fix it" and detailed edits) that do involve a server. Wispr Flow, Aqua Voice, Otter.ai, and Dragon Professional's own pages did not state on-device processing as a documented mode — see the comparison table for exact wording and links.
By default, yes: Hovor's cloud workflow sends your audio to OpenAI for transcription (and in some workflow configurations to Deepgram or ElevenLabs instead), then sends the resulting text to OpenAI again for grammar and punctuation clean-up. Per Hovor's privacy policy, audio is never stored on Hovor's server — it is deleted immediately after transcription completes. On-device models (Whisper or Parakeet) already work on the free tier under its weekly quota; Local Unlock ($49.99/1,999 UAH one-time) or a Pro subscription removes that cap. When the on-device model is loaded, audio and text stay on your device and formatting runs via Apple's on-device Foundation models on iOS 26+ or is skipped; if the local model is still loading when a recording starts, that recording falls back to server transcription. Hovor does not hold any GDPR, HIPAA, or SOC 2 certification — this describes technical data flow, not a compliance attestation.
Otter.ai's privacy policy explicitly states it trains "proprietary AI technology on de-identified audio recordings and on transcriptions." Wispr Flow's policy states that with its optional Privacy Mode turned off, "your data may be used to improve Flow's features and AI models" — turning Privacy Mode on removes this use entirely, per Wispr Flow's own wording. Superwhisper's and MacWhisper's own pages state audio is not used for training, consistent with their fully on-device architecture. Aqua Voice's page states retained transcript data (when its Privacy Mode is off) is used "to improve the product," without using the specific word "training." Apple states audio is only used to improve Siri and Dictation if a user separately opts in to "Improve Siri and Dictation," and is otherwise not stored. Dragon's and Google's own pages we checked did not address AI training use of dictation audio specifically.
Hovor gives you both: a free cloud tier (2,000 words/week) and an optional on-device mode (Whisper or Parakeet) that already works free-tier under its own weekly quota, unlimited with Local Unlock or Pro — audio staying on your device once the local model is loaded.
Get Hovor