Multilingual dictation is trickier than single language because most engines lock onto one language per session. This guide covers how to switch quickly and which tools handle mixed input best.
Pick your primary language per session
Before you press record, choose the language for that session. Most tools require you to reload after switching, so batching by language is more efficient than jumping around.
Switch languages inside Chrome and Edge
Zahvox lets you pick the language from the dropdown before recording.
Google Docs voice typing has a language selector next to the microphone.
Windows Voice Typing follows your Windows display language, so switch it in Settings.
Handle mixed language text
Dictate each language separately when possible. If you must mix, dictate the majority language and hand-type the minority language words. Most engines misrecognize embedded foreign words as similar-sounding native words.
Best practices for non-English dictation
Speak slower than you would in English if the model has less training data.
Add punctuation words explicitly, since auto punctuation is weaker in many languages.
Use a good microphone. Non-English models are more sensitive to noise.
Turn this into a repeatable weekly rhythm
Most people try dictation once, get a messy paragraph, and quietly go back to the keyboard. The trick is to schedule it. Block two short windows a week where multilingual voice typing is the only way you are allowed to draft. Speak fast, do not correct while talking, and let the transcript be ugly on purpose. Then spend five minutes cleaning it up. After three or four weeks the ugly first pass gets noticeably tighter, and your typed drafts start to sound more like your spoken voice too — which is usually the voice readers actually want.
Where Zahvox fits after the raw capture
Zahvox is the second pass, not the first. The browser, the OS, or the app you dictated into gave you words on a page. Zahvox at /tool is where you paste those words, tighten the wording, fix the shape of the sentences, and copy the result back into wherever it needs to live. That two-tool loop — raw capture anywhere, careful edit in Zahvox — is the whole workflow, and it stays the same whether the source was a laptop, a phone, or a meeting transcript.