Audio processing
Live dictation and file transcription run on your device using local engines such as Whisper, Parakeet, or Apple speech technology. Your audio is not sent to oto servers.
oto is built around a simple rule: speech should become text on your hardware, under your control. The app does not need an account, does not upload your audio, and does not include product telemetry.
Live dictation and file transcription run on your device using local engines such as Whisper, Parakeet, or Apple speech technology. Your audio is not sent to oto servers.
Settings, hotkeys, history, and local files stay on the device where you use oto. If you delete them locally, they are gone from oto.
There is no oto account system. We do not ask you to sign in, and we do not have a user profile tied to your voice or transcripts.
The app is intentionally boring from a data perspective. We cannot review your recordings, read your transcripts, or build a profile of how you work.
If you enable optional cloud AI cleanup, only transcript text is sent to the provider you configured. Audio is never sent for that feature. API keys are stored in macOS Keychain where supported.
Off by default. If you turn it on, oto reads the names visible on your screen when you start a recording so people and products are spelled correctly. It happens entirely on your Mac: a single frame is read and discarded immediately. No screen image is stored, and nothing is sent anywhere.
Long-form meeting recording can capture your Mac's system audio so the other side of a call stays separate from your microphone. That audio is captured and transcribed locally, on your device, exactly like your microphone input.
The bundled oto command line tool and the AI-assistant (MCP) integration talk to the app over a private channel inside a local, sandboxed app container. They add ways to reach oto from your own tools; they do not change where transcription happens. Your audio still never leaves your Mac.
If you contact support, we receive the email address and message you send. We use that only to reply and resolve the issue.
Required for live dictation. The audio is processed locally for transcription.
Used for global hotkeys and placing text where your cursor already is.
Used to detect single-key hotkeys, so you can start dictation without a multi-key shortcut.
Used only when you choose an audio or video file to transcribe.
Optional. Used only if you enable Screen Context (to read visible names on your Mac, on device) and, on older macOS versions, to capture system audio for meetings. No screen content is recorded or stored.
Optional. Used only during long-form meeting recording, to keep the other side of a call separate from your microphone. Captured and transcribed on your device.
Last updated: July 8, 2026. Questions about this policy can be sent to oto.transcribe@gmail.com.
Download oto from the Mac App Store, or try the full app free through TestFlight.