Turn audio or video — spoken in any of 100+ languages — into English text with timestamps, powered by Whisper running right here on your device. Your recording never leaves your computer.
MP3, WAV, M4A, OGG, FLAC, or the audio track of a video file.
Add an audio or video file under 60 minutes. It's read on your device — never uploaded.
One button — no settings. It detects the spoken language on its own and turns it into English text.
Read it with timestamps, copy the text, or download .txt, .srt, or .vtt subtitles.
No. The speech-recognition model (OpenAI's Whisper) runs entirely in your browser, so your audio never leaves your device. The only things downloaded are the model and its runtime — once, from a CDN — after which they're cached, so repeat runs skip the download.
Because everything runs on your own device, a very long file could make your browser crawl or freeze the tab. If your file is longer, trim it first with the Audio Cutter and bring the clip back here — it's the same suite, one click away.
It uses Whisper, the same model family behind many transcription services. On browsers with WebGPU (Chrome, Edge, Safari 18+) it runs several times faster than real time; on others it falls back to CPU, which is slower.
Around 100 for the spoken audio — it detects the language automatically. The transcript always comes out in English, translating other languages for you. Nothing to set.
Common audio (MP3, WAV, M4A, OGG, FLAC) and video files — it transcribes the audio track of a video too. The first time you run it, it downloads the model (~300 MB on a GPU, less on other devices); after that it's cached.
Yes. Alongside plain text (.txt), you can export .srt and .vtt subtitle files with timestamps, ready to drop onto a video.