Automatic subtitles

Turn speech in a video or audio file into subtitles and a transcript. Speech recognition runs locally with Whisper, nothing is uploaded.

Drop a video or audio file hereMP4 · MOV · WebM · MP3 · WAV

🔒 Your file and the transcript stay on this device.

How it works

How it works

  1. 1

    Add what goes in

    Drop your MP4, MOV, M4V, WebM and 8 more formats here, or pick it from your device.

  2. 2

    Start it

    Set the spoken language, translation to English and the subtitle text, then start. Everything runs in this browser — no upload, no queue, no account.

  3. 3

    Take the result

    The SRT · VTT · TXT · WebM is ready the moment the run finishes. Save it straight to your device.

What this tool does

What goes in
Reads MP4, MOV, M4V, WebM and 8 more formats. One file at a time.
What comes out
You get SRT · VTT · TXT · WebM.
What happens to quality
Your file is read, not changed. The original stays exactly as it was.
Where it stops
The first run downloads a small model to your device. After that it works offline. A long or high-resolution video is limited by your device's memory. Phones give up earlier than laptops — half an hour of 4K can be too much. Which video codecs work depends on your browser. Chrome and Edge read the most, including H.265 on many machines; Safari and Firefox refuse some. A file that will not open here usually opens in Chrome.
Automatic subtitles
Runs OpenAI’s Whisper tiny model in your browser: 99 languages with automatic detection, optional translation into English and editable timings. The model is downloaded once (about 41 MB) and then works offline; burnt-in subtitles are rendered as a new WebM video.

Supported formats

MP4MOVWebMMKVMP3M4AWAV → SRTVTTTXTWebM

Good to know

Frequently asked questions

How accurate are the subtitles?

Whisper tiny is the smallest Whisper model. Clear speech in common languages works well; music, several people talking at once or heavy accents cause mistakes. Every line can be corrected before you download.

Why is there a download before the first run?

Speech recognition needs a model of about 41 MB. It comes from usetool.app itself, is checked against a fixed checksum and then stays in your browser, so later runs work without internet.

What files can I use here?

Reads MP4, MOV, M4V, WebM and 8 more formats. One file at a time.

What do I get back?

You get SRT · VTT · TXT · WebM.

Does the file lose quality?

Your file is read, not changed. The original stays exactly as it was.

When does it not work?

The first run downloads a small model to your device. After that it works offline. A long or high-resolution video is limited by your device's memory. Phones give up earlier than laptops — half an hour of 4K can be too much. Which video codecs work depends on your browser. Chrome and Edge read the most, including H.265 on many machines; Safari and Firefox refuse some. A file that will not open here usually opens in Chrome.

Does the video lose quality?

Only when it has to be re-encoded. Cutting, joining and removing the sound track copy the picture data unchanged, so those are lossless. Compressing and adding a watermark do re-encode, which is why they take longer.

How long may a video be?

There is no fixed limit. What decides is the memory of your device: a phone manages a few minutes comfortably, a computer handles far more. If a long video fails, splitting it first and working on the parts usually helps.

Private by design

Your files stay on your device

Your file and the transcript stay on this device.

Other tools

Ready-made

Workflows for this file type