Beta
Both stages run on your machine.
mumblemuch is voice dictation and voice control for Windows. Two stages do the work: the speech stage, which turns your recording into a transcript, and the cleanup stage, which refines that transcript or decides what you asked for. Both of them can run on your own machine.
Speech and cleanup, both local
The speech stage transcribes what you said.
The cleanup stage is a language model. In Dictate it refines the transcript with a prompt selected by the application that has focus, then pastes the result into that application. In Intent it classifies one utterance into one action.
Set both stages to run here and they use open speech and language models that run on your machine. The path from the microphone to the text is then one machine long.
mumblemuch runs locally by default. No dictated words leave this device unless a cloud provider is configured for that run. Read what stays on this device, and for how long.
Windows draws the same line in its own speech features: Microsoft's page on speech and privacy separates device-based recognition, which processes your voice locally on your device, from online recognition, which uses Microsoft cloud services.
- https://support.microsoft.com/en-us/windows/privacy/speech-voice-activation-inking-typing-and-privacy - read 2026-09-21.
The chain you choose
Speech-to-text and the language model each run their own ordered list of providers. mumblemuch works down the list: when an entry cannot run, the next one takes the work.
Each list starts from a template. Local keeps the work on this machine. Custom points at an endpoint you run. mumblemuch cloud points at ours. Set every stage to the local template and the whole chain stays here.
Each stage is set on its own. Speech can stay here while the language model points somewhere else. Both can stay here.
The lists are yours to order in the workshop. See the provider chain in full.
Your GPU, or your CPU
It runs on your GPU or CPU. mumblemuch looks for CUDA first, then Vulkan, then falls back to the CPU. The detection is automatic.
At startup, and again each time the configuration reloads, mumblemuch checks whether this machine can run each local speech or language entry you have configured. An entry this machine cannot run is marked unusable, and the chain moves on to the next one.
What the setup does
The local models are put in place with the application. The endpoint wizard asks where each stage should run, and you can skip its pages: a skipped page leaves that stage on the local template.
Dictation with the built-in local model is English.
Local is where you start. Moving a stage off this machine is a change you make on purpose.
What the chain does not change
What mumblemuch may read and touch is yours to set, local chain or not. Six trust levels sit on a slider, each one nested inside the next, and the default is glance: it reads the clipboard, the selection and the window title, once per request. Those are the trust levels you grant.
The self-learning process reads the run history kept on this device. What it drafts from that history - a prompt, a lexicon entry - is a proposal you approve.
The steps are the same whichever provider runs them: transcribe, clean up, paste. Read what dictation does with the words.
Two examples
Two examples, not a feature list.
Dictating with every stage local. Every stage is set to run here. Press the dictate shortcut, speak a paragraph into your editor, press it again, and the words arrive there. The transcription and the cleanup both ran here.
Asking about what you copied. Copy an error message, press the intent shortcut, and ask what it means. mumblemuch reads the clipboard for that request, and the answer is pasted where your cursor is. With every stage set to run here, the answer was worked out here too.
mumblemuch is in closed beta
mumblemuch is not publicly available yet. Add your email below to join the closed beta, and we will write to you before launch.