Dictation that runs on the PC you already have.
No graphics card? It works, fully. Have an NVIDIA card? You also get the most accurate engine we ship and the local AI studio.
Press a hotkey and speak. NVIDIA's Parakeet v3 transcribes on your CPU, so the no-GPU path is the whole product rather than a trial mode: 11.9% word error across our five published clips, and 2.0% on the fast natural speech clip, where Whisper large-v3 on an RTX 3080 Ti scores 3.1%. Add an NVIDIA card and you get large-v3 itself, still the most accurate overall at 7.8%, plus a local model that rewrites and translates. Your audio and transcripts never leave the machine.
Real dictation on a real machine, no cuts. Everything you see runs on the CPU; the network is doing nothing.
What we store, plainly
No hedging, and nothing here needs taking on trust - run any network monitor while you dictate.
The only outbound calls Voxmelt makes are model downloads, a licence check, and update checks. Full privacy policy
100% local processing - zero cloud, zero cap.
The engines are already warm - Parakeet on your CPU, your AI model on the GPU. While they sit loaded, you get four creative surfaces at no extra cost:
Always-on-top floating widget. Dictate over any app without switching windows - one tap to record, polished text lands at your cursor.
Paste any text, pick a template and tone, the local LLM reshapes it. Emails, tickets, commit messages - no mic needed, same GPU.
Local text-to-speech, runs on your CPU. 4 personas x 8 tones = 32 voices. Export WAV files. No per-character billing, no cloud TTS.
Voice commands control the entire pipeline - start, stop, copy, paste, switch presets - without touching the keyboard.
Two heavy AI models. One GPU. Fully orchestrated.
The all-GPU mode, for people who want it: flip dictation from the default CPU engine to Whisper on CUDA and Voxmelt runs the speech engine and the LLM cleanup engine on the same consumer GPU - swapping them in and out of VRAM as VRAM permits. Tap a model to load it, or flip the dial to swap mode and watch only one ride the card at a time.
Speak. Stop. It's already done.
No "go" button, no second-guessing. The instant your mic cools, the cleanup fires and polished text streams in token by token. Here's the whole cycle, faked for the browser - the real thing runs on your GPU, offline.
For people who keep their data on their own machine.
If you keep a hard rule that sensitive audio never leaves the building, this was built for you. Any Windows PC runs it - an NVIDIA GPU just adds the local AI studio.
Developers
Talk through a commit message or a spec; clean text lands where your cursor is. Works fully air-gapped.
Creators & streamers
Caption and transcribe footage offline on your CPU while the RTX card stays free for editing and games. No upload, no wait.
Writers
Draft at the speed of speech. A local LLM tidies grammar and filler so the first pass reads like a third.
Clinicians
Dictate notes between patients without a single byte of PHI touching the cloud. HIPAA stays simple when nothing leaves the room.
Legal & finance
Privileged calls and filings transcribed on-box. No third-party processor, no data-residency paperwork.
Air-gapped teams
Defence, research, and secure labs run Voxmelt with the network cable unplugged. By design, not by promise.
Talk like you talk. Send what comes out.
Pick whoever sounds most like you. The messy version on the left is the point - fillers, restarts and thinking out loud are what it expects.
The future they announced. The hardware you already own.
Jensen Huang says “the PC is being reinvented.” Satya Nadella wants “unmetered intelligence to every home and every desk.” That future ships this fall on RTX Spark Windows PCs with up to 128 GB of unified memory - and Voxmelt will run even better on it. But you don't have to wait for new silicon: the RTX card in your rig runs the whole local pipeline today.
Voxmelt is independent and not affiliated with, sponsored by, or endorsed by NVIDIA or Microsoft. “NVIDIA,” “RTX Spark,” and “Windows” are trademarks of their respective owners, used here descriptively only.