Transcribe video and audio to text in 99 languages
Whisper large-v3 with automatic language detection and word-level timestamps. You get paragraph text, properly split SRT and VTT subtitles and, if you want, an English translation. Up to 3 hours and 1 GB.

How it works
Upload the video or audio
MP4, MOV, WebM, MKV, MP3, WAV, M4A… up to 1 GB. Pick the language or leave it on auto.
Whisper transcribes with timestamps
The large-v3 model with voice detection: it skips silence and music and timestamps every word.
Download text and subtitles
TXT in paragraphs, SRT and VTT ready for YouTube or your editor, JSON with the words. Deleted in 1 hour.
Frequently asked questions
Which languages are supported?
All 99 Whisper languages: Spanish, English, French, German, Portuguese, Italian, Japanese, Chinese, Arabic, Russian, Hindi… Detected automatically; you can pin it if the audio mixes languages.
How accurate is it?
We run the large model (large-v3), the same one behind the reference paid tools, with a voice filter so nothing is invented during silence. With clean audio it is close to a human transcript.
How are the subtitles formatted?
The way the trade does it: at most 2 lines of 42 characters, up to 7 seconds per cue, cuts on pauses and punctuation. Drop them straight into YouTube, Premiere or DaVinci.
Can I translate it?
Yes: tick “add English translation” and you also get TXT, SRT and VTT in English (Whisper's own translation).
How long does it take?
About 1 minute per 20-30 minutes of audio when warm, plus 1-2 minutes if the engine starts cold. The bar shows an estimate.
Do you keep my videos?
No. The video is processed and deleted: 1 hour after download, 24 h at most. No copies.



















