Text to Speech
Turn text into natural-sounding speech for free — choose from 28 American and British English AI voices, adjust the speed, and download the result as an MP3 or WAV file. Speech is generated on your own device by the open-source Kokoro voice model, so your text is never uploaded, there's no sign-up and no character quota beyond 20,000 per run.
100% private — everything runs in your browser. Conversion happens on your device; nothing is uploaded.
The first time, the voice model (about 90 MB) downloads once and is then cached by your browser.
Open-source credits: voice model Kokoro-82M by hexgrad (Apache-2.0); kokoro-js (Apache-2.0) and Transformers.js (Apache-2.0); pronunciation by phonemizer.js, which includes eSpeak NG — licensed under the GNU GPL v3, loaded unmodified, with full source code available at those links; MP3 encoding by lamejs (LGPL).
Speech is generated on your device by an open-source AI voice model — your text is never uploaded. English only (American and British voices). Very long texts are read sentence by sentence; on a phone, shorter texts work best.
How to convert text to speech
- Type or paste your text.
- Choose a voice and speed.
- Click Generate speech (the voice model downloads once the first time).
- Play it, then download the MP3 or WAV.
What you can use it for
- Create voice-overs for videos, reels and presentations
- Listen to articles, notes and drafts while you work
- Make audio versions of blog posts and newsletters
- Prototype narration for e-learning and product demos
More free tools
Designing with text?
Style, caption and decorate text with our free, private text and image tools.
See all free tools →Frequently asked questions
Is this text to speech really free?
Yes. There's no sign-up, watermark or paid tier — the voice model runs on your own device, so it costs us nothing per use.
Can I use the audio commercially?
The Kokoro voice model is released under the Apache-2.0 licence, which allows commercial use. As always, make sure you have the rights to the text you're reading.
Why does the first run take a while?
The first time, your browser downloads the voice model (about 90 MB). After that it's cached, and generation starts straight away.
Which languages are supported?
English — 20 American and 8 British voices. Other languages aren't supported by this voice model yet.
Is my text uploaded?
No. The text is turned into speech entirely in your browser; only the voice model itself is downloaded.