Subtitle translator
Translate subtitle text while preserving its timestamps.
Subtitle translator
How to use Subtitle translator
The subtitle translator translates SRT or WebVTT cue text between English and Spanish while keeping every cue's start and end time. Translation runs in the browser with Marian models downloaded on first use. Set source and target language codes to en and es in either direction. The result downloads as translated.srt, with a limit of 200 cues and 30,000 characters.
- Paste the source into the input editor, or load the built-in example.
- Set source language code, target language code.
- Run subtitle translator, review the output, then use the available copy or download controls.
What this tool supports
English ↔ Spanish translation uses browser Marian models, downloaded on first run. Document translation exports extracted text in Markdown and does not rebuild page layout.
- Source language code
- Set this value for your task.
- Target language code
- Set this value for your task.
Limits and processing
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.
Example source
1 00:00:01,000 --> 00:00:03,000 Hello, world.
Frequently asked questions
Which languages can the subtitle translator handle?
Only English to Spanish and Spanish to English. Enter en and es as the source and target codes in either order. Any other combination returns an error explaining the supported pair.
Does translating subtitles change the timestamps?
No. Each cue's start and end times are copied unchanged and only the text is translated. Translated lines may be longer than the originals, so the tool suggests reviewing wording and line length.
Can I translate a VTT file and keep it as VTT?
WebVTT input is accepted, but the translated output is always written as SRT. Convert it back with an SRT to WebVTT converter if your player needs VTT.
Where is my input processed?
Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.
What are the input limits?
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.