# Automatic subtitle generator

> Transcribe speech to timestamped SRT or WebVTT subtitles.

[Open tool](https://www.toolcabana.com/nd/auto-subtitles) · [Audio, video and files](https://www.toolcabana.com/nd/category/audio-video-and-files)

Tool ID: auto-subtitles. Requested language: nd. Description language: en. Complete guide translation: no; untranslated sections use English.

## Overview (en)

The automatic subtitle generator transcribes speech from an audio or video file and returns timestamped captions as SRT or WebVTT. It runs the Whisper tiny model in the browser after a one-time model download. You can leave the language blank for automatic detection or enter a code, and set how many seconds to transcribe from 1 to 120.

## Supported tasks (en)

Downloads a browser inference model on first run, then processes source locally in a cancellable Web Worker. Background removal uses BEN2; transcription uses Whisper tiny. Review recognition errors and image edges.

## Steps (en)

1. Choose a source file. If the tool accepts multiple files, arrange them in the order you want.
2. Set language code (blank = detect), maximum seconds, subtitle format.
3. Run automatic subtitle generator, review the output, then use the available copy or download controls.

## Settings

- Language code (blank = detect) (en; key: language; type: text)
- Maximum seconds (en; key: duration; type: number); minimum: 1; maximum: 120
- Subtitle format (en; key: format; type: select): SRT [en], WebVTT [en]

## Limitations (en)

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.

## Example input

```text
A useful tool makes everyday work easier. Review the results before sharing.
```

## Privacy and connections (en)

Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.

## Questions (en)

### How long a video can I auto-generate subtitles for?

The tool transcribes from the beginning of the file up to the Maximum seconds value, which ranges from 1 to 120 with 30 as the default. Longer recordings need to be split and processed in parts.

### Should I set a language code for automatic subtitles?

Leaving it blank lets Whisper detect the language. If detection picks the wrong language, enter a code such as en or fr. When no speech is recognized, the tool asks for a clearer recording.

### Can I edit the timing of auto-generated subtitles?

The output is plain SRT or WebVTT text you can copy or download. Review words and timings against the audio, since the tool's notice warns that recognition errors happen. A subtitle timing editor can shift cues afterward.

### Where is my input processed?

Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.

### What are the input limits?

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.

## Related tools

- [Burn subtitles into video](https://www.toolcabana.com/nd/burn-subtitles)
- [Word-highlight subtitles](https://www.toolcabana.com/nd/word-subtitles)
- [Subtitle translator](https://www.toolcabana.com/nd/subtitle-translate)
- [Speaker-labelled transcript](https://www.toolcabana.com/nd/speaker-transcript)
