Burn subtitles into video
Transcribe and burn subtitle overlays into a video.
Burn subtitles into video
Drop your files here
or choose files from your device
How to use Burn subtitles into video
Burn subtitles into video transcribes the speech in a short clip and draws the captions directly onto the frames, returning an MP4 with the original audio. Captions appear as bold white text with a black outline, centered near the bottom and wrapped to at most four lines. You can set the font size from 12 to 64 and the language.
- Choose a source file. If the tool accepts multiple files, arrange them in the order you want.
- Set language code (blank = detect), maximum seconds, subtitle font size.
- Run burn subtitles into video, review the output, then use the available copy or download controls.
What this tool supports
Processes locally at most 20 seconds at 4 frames per second and up to 640 pixels wide. Frame-by-frame editing can flicker. Selection coordinates refer to processed frames. Subtitles are sampled at frame times.
- Language code (blank = detect)
- Set this value for your task.
- Maximum seconds
- Range: 1 to 20
- Subtitle font size
- Range: 12 to 64
Limits and processing
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.
Example source
A useful tool makes everyday work easier. Review the results before sharing.
Frequently asked questions
Can I burn my own SRT file into a video?
No. This tool creates captions by transcribing the clip with a local Whisper model and then draws them on the frames. It does not accept an existing subtitle file as input.
How long can a video be for burned-in subtitles?
Up to 20 seconds, set by Maximum seconds with a default of 10. Frames are processed at 4 frames per second and up to 640 pixels wide, so the output is a short, low-frame-rate MP4.
Can I change the subtitle size and style?
You can set the font size from 12 to 64 pixels; the default is 24. The style is fixed: bold sans-serif white text with a black outline, centered at the bottom of the frame.
Where is my input processed?
Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.
What are the input limits?
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.