# Offline neural text-to-speech

> Synthesize English speech with a local Supertonic voice.

[Open tool](https://www.toolcabana.com/cu/local-neural-tts) · [Audio, video and files](https://www.toolcabana.com/cu/category/audio-video-and-files)

Tool ID: local-neural-tts. Requested language: cu. Description language: en. Complete guide translation: no; untranslated sections use English.

## Overview (en)

Offline neural text-to-speech turns up to 1,000 characters of English text into a downloadable speech.wav file. It runs a Supertonic neural voice model in the browser, downloading the model and voice vector on first use. You can adjust speaking speed from 0.8 to 1.2 in steps of 0.05, with 1 as the default.

## Supported tasks (en)

Generates an English WAV download locally with Supertonic. Downloads the model and voice vector on first run. Source text stays on-device.

## Steps (en)

1. Paste the source into the input editor, or load the built-in example.
2. Set speaking speed.
3. Run offline neural text-to-speech, review the output, then use the available copy or download controls.

## Settings

- Speaking speed (en; key: speed; type: number); minimum: 0.8; maximum: 1.2

## Limitations (en)

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.

## Example input

```text
A useful tool makes everyday work easier. Review the results before sharing.
```

## Privacy and connections (en)

Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.

## Questions (en)

### Can I download the generated speech as an audio file?

Yes. The voice is rendered locally and returned as speech.wav, a 16-bit mono WAV file. This differs from browser read-aloud, which only plays audio and cannot be saved.

### Does offline neural text-to-speech support other languages?

No. The voice is English only, and input must be between 1 and 1,000 characters. For other languages, a browser speech tool that uses installed system voices may be an option.

### Does offline text-to-speech work without internet?

The first run needs a connection to download the model and voice vector. After that, models may remain in the browser cache, and the text itself is processed on your device rather than sent to a speech service.

### Where is my input processed?

Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.

### What are the input limits?

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.

## Related tools

- [Automatic subtitle generator](https://www.toolcabana.com/cu/auto-subtitles)
- [Burn subtitles into video](https://www.toolcabana.com/cu/burn-subtitles)
- [Word-highlight subtitles](https://www.toolcabana.com/cu/word-subtitles)
- [Subtitle translator](https://www.toolcabana.com/cu/subtitle-translate)
