Vox is an on-device AI dictation tool that transforms speech into text directly on your Mac or Windows machine. It belongs to the productivity category of voice-to-text applications, designed for anyone who types frequently—from writers and developers to busy professionals. Its core value lies in absolute privacy and speed: by processing everything locally without cloud uploads, Vox eliminates latency, account requirements, and data exposure while delivering clean, formatted text ready to paste.
The primary pain point Vox addresses is the inefficiency and privacy risk of traditional dictation services. Cloud-based dictation tools require an internet connection, often demand account creation, and send audio to remote servers where it may be stored, analyzed, or leaked. This creates friction for users who want quick transcription without sacrificing control over their data. Vox solves this by running entirely on-device, ensuring audio never leaves the machine, no telemetry is collected, and the app works even in airplane mode—turning a privacy concern into a productivity advantage.
Vox leverages local AI models for both transcription and cleanup. For transcription, it uses OpenAI's Whisper or NVIDIA's Parakeet model, which run on the Apple Neural Engine or Metal GPU on Mac, or any DirectX 12 GPU on Windows with CPU fallback. For cleanup—removing filler words, correcting self-corrections, and formatting text—it uses Apple Intelligence on macOS 26+ or a local Gemma 4 model. This dual-model architecture ensures high accuracy while keeping everything on-device, meaning no internet is required after the initial one-time model download. Users can verify the absence of network calls with tools like Little Snitch or GlassWire.
Vox includes context-aware voice modes that adjust the cleanup engine based on the target application. The General mode provides balanced cleanup for any app, dropping fillers and enumerating lists. The Email mode produces formal, fully punctuated text suitable for mail clients like Gmail or Outlook, without inventing salutations. Chat mode tailors output to casual platforms like Slack or Discord, preserving fragments and contractions. Code Comment mode uses present-tense third-person and preserves identifiers verbatim, ideal for Xcode or VS Code. Notes mode generates full sentences with bullets on enumeration. Additionally, users can create custom modes with their own system prompts and auto-trigger rules.
Vox's architecture ensures complete privacy during dictation. No audio, transcripts, or telemetry are ever sent to a server; the only network calls are the one-time model download and an optional update check. This makes Vox a verifiably privacy-first tool, independently checkable with network monitors. After the initial download, Vox works entirely offline, including on a plane. The app requires no account to use—no sign-up, no login. For commercial use, a license requires only a billing email at Stripe. This combination of privacy and offline functionality sets Vox apart from every cloud-dependent dictation alternative.
admin
Vox's workflow is designed for minimal friction: three keys, no setup. Press and hold the default hotkey (⌘⌥. on Mac, Ctrl+Alt+. on Windows) to start listening—a small pill appears near the cursor. Speak naturally, including filler words and self-corrections; Vox's cleanup engine handles them. Release the hotkey, and the cleaned-up text is automatically copied to the clipboard. Then paste it wherever needed with ⌘V or Ctrl+V. Notably, Vox avoids silent keystroke synthesis, which is fragile and requires accessibility permissions. Instead, the user controls the final paste, ensuring compatibility and reliability across applications.
Users can employ Vox in various scenarios. For example, a knowledge worker drafting emails can switch to Email mode, speak naturally, and paste a polished message in seconds—saving up to 40 minutes daily if they type over 3,000 words. A developer writing code comments uses Code Comment mode to dictate verbatim variable names without extra formatting. A project manager on Slack can quickly respond in Chat mode, maintaining a casual tone. A researcher taking notes in Obsidian benefits from Notes mode's automatic bullets and paragraph breaks. The estimated time savings of 3x faster speaking vs typing translates to significant productivity gains, as illustrated by the built-in calculator based on Stanford research.
Vox targets Mac and Windows users with specific hardware requirements: Apple Silicon (M1 or newer) on macOS 14+, and 64-bit Windows 10/11 with DirectX 12 GPU support. Ideal users include freelancers, writers, developers, and knowledge workers who value privacy and speed. Vox is free for personal use under a perpetual Personal Use license. For companies with more than one employee using Vox, commercial licenses start at $12/seat/month. This fair-source model balances accessibility for individuals with sustainable revenue for businesses. In summary, Vox delivers on-device AI dictation that is fast, private, and reliable—turning speech into text without compromises.
Vox is ideal for knowledge workers, writers, developers, and professionals who type extensively throughout the day. Specifically, it suits freelancers seeking privacy-focused dictation without cloud dependency, office workers wanting to reduce typing strain, and developers needing efficient code comment dictation. Companies with multiple employees can adopt Vox for compliant, on-device transcription via commercial licensing. The tool also appeals to privacy-conscious users who work offline frequently, such as travelers or those using restricted networks. Hardware requirements target Mac users with Apple Silicon (M1+) on macOS 14+, and Windows users with 64-bit systems and DirectX 12 GPUs.