Beta
Speak, and something happens.
Press a shortcut, say what you need, and mumblemuch works out what you meant. It does what people search for as an AI voice assistant for Windows: one spoken request becomes one action, or an abstain when nothing is clear.
One spoken request, one action.
Each mode has its own shortcut, and a shortcut can be a keyboard chord, a mouse button, or a gamepad combination. One press starts the recording, one press stops it.
When the recording stops, mumblemuch classifies it and runs exactly one continuation. There is no command grammar to memorise: the utterance is classified, not matched. You say the thing; the classification decides what happens next. Here is how a recording becomes text or an action.
Dictate types what you said. Intent works out what you meant. When you want the words themselves rather than the outcome, use dictation instead of an instruction. Windows ships speech features of its own, and they work the other way round: see the Windows speech features you already have.
What a request can do.
You speak one request and mumblemuch classifies it. The language model answers it, and when you say so it answers from the text you copied: say "what I copied" or "the clipboard" and the copied text goes in with the request. The answer is pasted where your cursor is.
When nothing resolves, mumblemuch abstains and nothing happens. It produces nothing rather than a guess.
The design routes spoken requests further - launching applications and scripts, web searches, learned automations - and those routes are not in the current beta build.
Dictate and Intent are two of all three modes.
Where the answer goes.
The answer is pasted into the editor you are working in, at the cursor. Which continuation ran decides what you get, and it is not a setting you pick each time.
What this looks like.
These are examples of what a request can sound like. They are not a list of commands: the utterance is classified, not matched.
- You copy a paragraph, press the shortcut, and say which language you want it in. The translated text is pasted where your cursor is.
- You copy an error message and ask what it means. The explanation is pasted where you are working.
- You copy a long note and ask for the short version. The summary lands at the cursor.
- You say something that does not resolve into a request. mumblemuch does nothing.
No wake word.
Nothing is listening. The microphone opens when you press the shortcut and closes when you press it again. There is no name to say first and no chat window to keep open.
It is not an agent running on its own. One recording produces one classified continuation, and when the classification is not clear it abstains rather than guess. It does one thing per press, and you started the press.
What it is allowed to touch.
You set the ceiling before you speak. A Trust slider grants one of six nested trust levels, and the default is glance.
At glance, mumblemuch may read the clipboard, the current selection and the title of the window you are in, once per request - which is exactly what answering a question about a paragraph you copied needs. Every level above glance carries what glance carries and more, and the level you grant is the level it keeps until you change it.
Speech and language each run an ordered chain of providers and fail over down it. The chain can be the local template, your own endpoint, or mumblemuch cloud. mumblemuch runs locally by default, and nothing leaves this device unless a cloud provider is configured for that run. Here is what stays on this device.
mumblemuch is in closed beta. Leave your email in the form below to join.