Audio to Text

Convert audio to text for free with OpenAI's Whisper speech recognition, running privately in your browser. Transcribe MP3, WAV, M4A and even video files in about 100 languages, add timestamps, translate to English, and download TXT, SRT or VTT. No upload, no account and no minutes limit.

100% private — transcription runs on your device. Conversion happens on your device; nothing is uploaded.

100% private — runs in your browser

Transcription runs on your device with OpenAI’s open-source Whisper model — your audio is never uploaded. The model downloads once (40–250 MB) and is then cached by your browser. A fast computer with a modern browser (Chrome or Edge, ideally with WebGPU) transcribes an hour of audio in a few minutes; older devices are slower. Always check names and numbers.

How to convert audio to text

  1. Drop your audio or video file in the box.
  2. Choose the spoken language (or detect automatically) and accuracy.
  3. Click Transcribe — the model downloads once, then runs on your device.
  4. Edit the text if needed, then copy it or download TXT, SRT or VTT.

What you can use it for

  • Transcribe interviews, meetings and lectures
  • Turn voice memos and podcasts into text
  • Write up research recordings without sending them to a cloud service
  • Make notes and quotes from webinars and videos

More free tools

Designing with text?

Style, caption and decorate text with our free, private text and image tools.

See all free tools →

Frequently asked questions

How do I convert an audio file to text?

Drop the file in, pick the language, and click Transcribe. When it finishes, the text appears in an editable box — copy it or download it.

Which languages are supported?

Whisper understands about 100 languages. The menu lists the most common ones; “Detect automatically” works for the rest. You can also translate speech in any supported language straight into English text.

Is my audio uploaded?

No. OpenAI's open-source Whisper model runs inside your browser. The model is downloaded once (40–250 MB) and cached; your audio never leaves your device.

How long does it take?

On a recent laptop with Chrome or Edge, an hour of audio takes a few minutes with the Balanced model (faster with WebGPU). Older phones and computers are slower — try the Fast model.

How accurate is it?

Clear speech with little background noise is transcribed very well. Accents, crosstalk, music and technical terms reduce accuracy — always check names and numbers before publishing.

Can I get timestamps?

Yes — tick Include timestamps for a time-coded transcript, or choose SRT/VTT for subtitle files.

Is there a length limit?

No fixed limit. Long files are processed in two-minute parts with a progress bar; very long recordings (several hours) need plenty of memory.