Automatic subtitle generator
Transcribe speech to timestamped SRT or WebVTT subtitles.
Automatic subtitle generator
Drop your files here
or choose files from your device
د کارولو طریقه Automatic subtitle generator
The automatic subtitle generator transcribes speech from an audio or video file and returns timestamped captions as SRT or WebVTT. It runs the Whisper tiny model in the browser after a one-time model download. You can leave the language blank for automatic detection or enter a code, and set how many seconds to transcribe from 1 to 120.
- Choose a source file. If the tool accepts multiple files, arrange them in the order you want.
- Set language code (blank = detect), maximum seconds, subtitle format.
- Run automatic subtitle generator, review the output, then use the available copy or download controls.
د دې وسیلې وړتیاوې
Downloads a browser inference model on first run, then processes source locally in a cancellable Web Worker. Background removal uses BEN2; transcription uses Whisper tiny. Review recognition errors and image edges.
- Language code (blank = detect)
- Set this value for your task.
- Maximum seconds
- Range: 1 to 120
- Subtitle format
- SRT · WebVTT
محدودیتونه او پروسس
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.
د بېلګې سرچینه
A useful tool makes everyday work easier. Review the results before sharing.
عامې پوښتنې
How long a video can I auto-generate subtitles for?
The tool transcribes from the beginning of the file up to the Maximum seconds value, which ranges from 1 to 120 with 30 as the default. Longer recordings need to be split and processed in parts.
Should I set a language code for automatic subtitles?
Leaving it blank lets Whisper detect the language. If detection picks the wrong language, enter a code such as en or fr. When no speech is recognized, the tool asks for a clearer recording.
Can I edit the timing of auto-generated subtitles?
The output is plain SRT or WebVTT text you can copy or download. Review words and timings against the audio, since the tool's notice warns that recognition errors happen. A subtitle timing editor can shift cues afterward.
Where is my input processed?
Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.
What are the input limits?
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.