Skip to content
Content Creator Tools

Transcribe audio or video

Drop a file to get text, subtitles or synced lyrics. Recognition runs in your browser, the file is never uploaded.

The first run downloads the speech model (about 250–600 MB), then the browser keeps it cached.

How it works

Animation: a video file is dropped in, its sound is transcribed word by word with timestamps, then saved as SRT.

  1. 1

    Add a file

    Any audio or video: a song you made, a voice memo, a reel.

  2. 2

    Pick the language

    Choose the language of the speech or vocals. It makes recognition noticeably more accurate.

  3. 3

    Get text or subtitles

    Copy the text with timings or save SRT subtitles or AE Subs for After Effects.

Frequently Asked Questions

Is my file uploaded anywhere?

No. The file is decoded and recognized inside your browser. Only the model weights are downloaded from Hugging Face.

How accurate is it with music?

Sung vocals over loud instruments are harder than speech. Pick the right language for better results.

Why is the first run slow?

The browser downloads the model once and caches it. With WebGPU recognition is several times faster than on CPU.

Which formats can I export?

SRT subtitles and AE Subs: a file with line, word and letter timings for our After Effects plugin. The copy button gives you text with line timestamps.

Feedback or an idea

What is it