Wordcast
Back to blog

The best free text-to-speech voices on Windows, explained

Windows ships two very different sets of voices: old robotic SAPI5 ones and much better neural Natural voices. Here is how to find them and turn them on.

2026-05-27 · 6 min read

5 min listen

Windows has shipped with text-to-speech for decades, and that history is exactly why the voices confuse people. Out of the box you get a small handful of old voices with names like David, Zira, and Mark. They run entirely offline, they respond instantly, and they sound like a train station announcement from 2005. These are the SAPI5 voices, the ones the classic Speech settings and most old desktop apps use. If you have ever pressed a Read Aloud button on Windows and winced, this is probably what you heard.

The good voices are hiding one menu deeper. In the last few years Microsoft added a set of neural voices it calls Natural voices, with names like Microsoft Aria and Microsoft Guy. These are a different technology entirely. Instead of stitching together recorded fragments, they generate speech from a neural model, and the difference is not subtle. Aria pauses for commas, rises at questions, and does not sound like it is reading each word off a separate card. They are the same quality tier as the neural voices you hear in Microsoft Edge, and they are free.

The catch is that most of these Natural voices are not fully offline. Many of them stream from Microsoft's servers, which is how they can sound so good without shipping a large model to your machine. That means a Natural voice needs an internet connection, and the text you read gets sent to Microsoft to be turned into audio. Windows does also offer a smaller number of on-device Natural voices for accessibility, but the widest and best-sounding selection is the online set. This is the core tradeoff on Windows: the offline SAPI5 voices are private and instant but robotic, and the online Natural voices are excellent but leave your machine.

To add the Natural voices, open Windows Settings and look under Time and Language, then Speech. Recent versions of Windows 11 have a Manage voices section there where you can add voices, and each one downloads or registers a Natural voice you can then use. If you cannot find it in that exact spot, the other reliable path is through Narrator's settings, which lets you add its natural voices, and through the Accessibility settings, which surface the same neural voices. Menu names shift between Windows builds, so if the wording does not match, search Settings for the word voices and follow whichever result lets you add or manage them.

There is a separate lever that matters just as much: language packs. The Web Speech API in a browser can only offer voices that are actually installed on your system, so if you want a French, German, or Japanese voice, you usually have to add that language first. Go to Settings, then Time and Language, then Language and region, and add the language you want. When Windows installs it, it pulls in that language's speech voices too, and they show up in the browser voice list afterward. If a language you expected is missing from a TTS tool, a missing language pack is the most common reason.

Microsoft Edge is the easiest place to hear what these voices can do. Edge's Read Aloud feature uses the Natural voices directly, so if you open any article in Edge and pick Read Aloud, the voice menu there is a good preview of the best quality Windows can produce. It is worth doing this once as a reference point, because it tells you what a good result sounds like before you go hunting for the same voice in a browser TTS tool.

Now the part that trips people up. The Web Speech API exposes whatever voices your browser decides to surface, and different browsers make different choices. Edge tends to expose the full set of Microsoft Online Natural voices, so you will see entries labeled something like Microsoft Aria Online Natural. Chrome on the same machine often does not show those exact online voices, because Chrome surfaces its own voice engine and a different slice of the system voices. So it is completely normal to open a browser TTS tool in Chrome and Edge side by side on one Windows PC and get two different voice lists. Neither is broken. The browser, not the website, controls which engines appear.

That has a practical consequence when you use any browser-based TTS tool, Wordcast included. The tool can only offer the voices your browser hands it, so the fastest way to get the best Windows voices is to open the tool in Edge and look for a voice with Natural or Online in its name. Pick one of those and you get the neural quality. If you only see David, Zira, and Mark, you are looking at the old SAPI5 set, which usually means either you are in a browser that does not expose the Natural voices or the Natural voices have not been added in Settings yet. Add them, reload the page, and the better options should appear.

A short decision guide. If you care most about privacy or you are often offline, stay on the SAPI5 voices and accept the robotic sound, because they never send your text anywhere. If you care most about how it sounds and you are online anyway, use a Microsoft Natural voice through Edge and enjoy quality that used to cost money. If you want a specific language, install that language pack first or none of the above will show the voice you want. On Windows the voice you get is decided by three things stacked together: which voices you have installed, which language packs you have added, and which browser you opened the tool in. Once you know those three levers, the confusing voice list stops being a mystery and becomes something you can steer.