The Web Speech API sends audio to third-party servers. VPNs and corporate firewalls sometimes block those endpoints or introduce latency that breaks the connection. This guide helps you narrow down which layer is at fault.
Test with the VPN off
First rule out the VPN. Disable it briefly and try to dictate. If it works, the VPN is blocking or slowing the recognition endpoint. Try a different VPN server closer to your location.
Check the firewall
Corporate firewalls sometimes block traffic to speech recognition services. Ask your IT team to allowlist the Google Speech API endpoint for Chrome or the Microsoft speech endpoint for Edge, depending on which browser you use.
Split tunneling helps
If your VPN client supports split tunneling, exclude your browser from the VPN. Dictation traffic then goes over the direct connection and is fast enough to work reliably.
Confirm from a clean network
Try dictating from a personal hotspot. If it works there and not on the office network, the issue is definitely network layer, not the tool.
Turn this into a repeatable weekly rhythm
Most people try dictation once, get a messy paragraph, and quietly go back to the keyboard. The trick is to schedule it. Block two short windows a week where voice-to-text with a vpn or firewall is the only way you are allowed to draft. Speak fast, do not correct while talking, and let the transcript be ugly on purpose. Then spend five minutes cleaning it up. After three or four weeks the ugly first pass gets noticeably tighter, and your typed drafts start to sound more like your spoken voice too — which is usually the voice readers actually want.
Where Zahvox fits after the raw capture
Zahvox is the second pass, not the first. The browser, the OS, or the app you dictated into gave you words on a page. Zahvox at /tool is where you paste those words, tighten the wording, fix the shape of the sentences, and copy the result back into wherever it needs to live. That two-tool loop — raw capture anywhere, careful edit in Zahvox — is the whole workflow, and it stays the same whether the source was a laptop, a phone, or a meeting transcript.
Looking for something else? Browse the full
guides library or open the
free voice tool and practice as you read.