Word-highlight subtitles
Burn word-by-word highlighted subtitles into a short video.
Word-highlight subtitles
Drop your files here
or choose files from your device
How to use Word-highlight subtitles
Word-highlight subtitles transcribes a short clip with word-level timestamps and burns those words onto the video as they are spoken. Captions use bold white text with a black outline at the bottom of the frame, and you can set the font size from 12 to 64. The output is an MP4 of up to 20 seconds at 4 frames per second.
- Choose a source file. If the tool accepts multiple files, arrange them in the order you want.
- Set language code (blank = detect), maximum seconds, subtitle font size.
- Run word-highlight subtitles, review the output, then use the available copy or download controls.
What this tool supports
Processes locally at most 20 seconds at 4 frames per second and up to 640 pixels wide. Frame-by-frame editing can flicker. Selection coordinates refer to processed frames. Subtitles are sampled at frame times.
- Language code (blank = detect)
- Set this value for your task.
- Maximum seconds
- Range: 1 to 20
- Subtitle font size
- Range: 12 to 64
Limits and processing
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.
Example source
A useful tool makes everyday work easier. Review the results before sharing.
Frequently asked questions
How are word-by-word captions different from normal burned subtitles?
This tool asks the speech model for word-level timestamps, so each caption cue is a single word shown during its spoken interval. Standard burned subtitles use longer phrase-level cues instead.
Why are some words missing from my word-level captions?
Captions are drawn on frames sampled at 4 frames per second. A word spoken between two sampled frames, lasting under a quarter second, may never be visible. Review the output for skipped words.
Can I download the word timings as an SRT file?
The current implementation returns only the MP4 with captions burned in. For an SRT or WebVTT file, use the automatic subtitle generator, which outputs phrase-level cues as text.
Where is my input processed?
Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.
What are the input limits?
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.