Beta
Controlling Windows by voice
There are two ways to make a computer do what you say. You can learn a fixed set of commands and speak them exactly. Or you can say the thing in your own words and let one action be worked out from that. Windows ships the first kind. mumblemuch is the second kind. Here is what each one asks of you.
Two ways to control a computer by voice
A command grammar is a published list. Every phrase on it maps to one action, and the phrases are written down so you can look them up. That is predictable: one phrase, one action. The cost is that you carry the list, and anything the list does not cover has to be added to it first.
An utterance-first model has no list to learn. A shortcut starts the recording, you say what you want in your own words, and a second press stops it. The words are transcribed, then classified, and one action runs. When nothing is clear, it abstains instead of guessing. The cost is the mirror image: there is no published list to look up in advance.
Neither model replaces the other. A command grammar is precise about the screen - pointers, clicks, overlays and grids are all documented commands. An utterance-first model is precise about the request.
What Windows already gives you
Windows has two separate speech features, and they are not the same thing. Every sentence below is Microsoft's own documentation, with the page it comes from and the date it was read.
Voice access
Voice access is available in Windows 11, version 22H2 and later, and is turned on at Settings > Accessibility > Speech. [1]
It offers both commanding and dictation. [2]
Microsoft publishes a command list, in families with examples, covering opening, closing and switching applications, a number overlay, and a grid you can address to a point on the screen across displays. [3] [4] [5]
It uses on-device speech recognition and works without an internet connection, after a one-time download of the language files at first run. [1] [6] [7]
Its voice shortcuts chain a maximum of eight actions per shortcut, and are documented for English variants only. [8]
Fluid dictation is currently available in English and only on Copilot+ PC devices. [7]
A privacy statement specific to voice access is not documented in the frequently asked questions. [2]
Voice typing
Windows voice typing is the other feature. It offers dictation. [2]
It uses online speech recognition powered by Azure Speech services, and requires an internet connection. [9]
Where Microsoft's own pages disagree
On language coverage the two pages do not match. The setup page lists fifteen locales; the FAQ lists four. Both were read on 2026-09-21. [1] [2]
Sources
- Set up voice access - https://support.microsoft.com/en-us/accessibility/windows/voice-access/set-up-voice-access - read 2026-09-21.
- Voice access frequently asked questions - https://support.microsoft.com/en-us/accessibility/windows/voice-access/voice-access-frequently-asked-questions-faqs - read 2026-09-21.
- Voice access command list - https://support.microsoft.com/en-us/accessibility/windows/voice-access/voice-access-command-list - read 2026-09-21.
- Use voice to work with windows and apps - https://support.microsoft.com/en-us/accessibility/windows/voice-access/use-voice-to-work-with-windows-and-apps - read 2026-09-21.
- Use voice access on a multi-display setup - https://support.microsoft.com/en-us/accessibility/windows/voice-access/use-voice-access-on-a-multi-display-setup - read 2026-09-21.
- Use voice access to control your PC and author text with your voice - https://support.microsoft.com/en-us/accessibility/windows/voice-access/use-voice-access-to-control-your-pc-author-text-with-your-voice - read 2026-09-21.
- Get started with voice access - https://support.microsoft.com/en-us/accessibility/windows/voice-access/get-started-with-voice-access - read 2026-09-21.
- Use voice to create voice access shortcuts - https://support.microsoft.com/en-us/accessibility/windows/voice-access/use-voice-to-create-voice-access-shortcuts - read 2026-09-21.
- Use voice typing to talk instead of type on your PC - https://support.microsoft.com/en-us/accessibility/windows/use-voice-typing-to-talk-instead-of-type-on-your-pc - read 2026-09-21.
Wording taken from Microsoft's pages is normalised to plain ASCII. Each page above was read on the date shown.
Where mumblemuch sits
mumblemuch is voice dictation and voice control software for Windows. It does not replace either feature. It is the other model.
Each mode has its own shortcut, and a shortcut can be a keyboard chord, a mouse button, or a gamepad combination. One press starts the recording, one press stops it.
Dictate types what you said. The transcript is refined by a cleanup prompt chosen by whichever application has focus, then pasted into that application. More on dictation with cleanup per application.
Intent works out what you meant. One utterance is classified, and one action runs: the language model answers the request, and when you say so it answers from the text you copied. The answer is pasted where your cursor is. When nothing is clear, it abstains. More on what a spoken request can do here.
Rehearsal. Show it once. Let it learn. Ask it to do it again.
What mumblemuch may do on your machine is set before you speak. Capability is granted on a Trust slider with six nested levels, and the default is glance. Raising it is your decision. Read the permission model.
There is a fully local path: open speech and language models that run on your machine, on your GPU or CPU. See the steps from press to result.
Four examples
Examples, not a feature list.
Press the Dictate shortcut, speak a paragraph into an email, and press it again. The transcript is cleaned by the prompt for that application and pasted there.
Copy a paragraph, press the Intent shortcut, and ask what it commits you to. mumblemuch reads what you copied for that request, and the answer is pasted where your cursor is.
Copy a sentence, press the Intent shortcut, and say which language you want it in. The translation is pasted where your cursor is.
Press the Intent shortcut and think out loud instead of asking for something. With no clear intent it abstains, and no action runs.
What mumblemuch does not do
mumblemuch has no spoken pointer control. There is no number overlay, no mouse grid, and no command for clicking a control by name. If that is what you need, voice access documents those commands. [3] [4] [5]
Starting a recording is a physical press: a key, a mouse button, or a gamepad button. Speaking is not the only action involved.
Status
mumblemuch is in closed beta. Leave your email below to hear when it opens up.