Transcribe audio or video
Drop a file to get text, subtitles or synced lyrics. Recognition runs in your browser, the file is never uploaded.
The first run downloads the speech model (about 250–600 MB), then the browser keeps it cached.
How it works
Animation: a video file is dropped in, its sound is transcribed word by word with timestamps, then saved as SRT.
- 1
Add a file
Any audio or video: a song you made, a voice memo, a reel.
- 2
Pick the language
Choose the language of the speech or vocals. It makes recognition noticeably more accurate.
- 3
Get text or subtitles
Copy the text with timings or save SRT subtitles or AE Subs for After Effects.
Frequently Asked Questions
Is my file uploaded anywhere?
No. The file is decoded and recognized inside your browser. Only the model weights are downloaded from Hugging Face.
How accurate is it with music?
Sung vocals over loud instruments are harder than speech. Pick the right language for better results.
Why is the first run slow?
The browser downloads the model once and caches it. With WebGPU recognition is several times faster than on CPU.
Which formats can I export?
SRT subtitles and AE Subs: a file with line, word and letter timings for our After Effects plugin. The copy button gives you text with line timestamps.