ToolCabana

Offline neural text-to-speech

Synthesize English speech with a local Supertonic voice.

Favorites are saved in this browser. Find them in My favorites.

Offline neural text-to-speech

Your input stays in this browser unless stated otherwise
Clear your input, files and result, and restore the default settings.
Load sample text to explore what this tool can do. This replaces your current input.
0 charactersClear the source text.

د کارولو طریقه Offline neural text-to-speech

Offline neural text-to-speech turns up to 1,000 characters of English text into a downloadable speech.wav file. It runs a Supertonic neural voice model in the browser, downloading the model and voice vector on first use. You can adjust speaking speed from 0.8 to 1.2 in steps of 0.05, with 1 as the default.

  1. Paste the source into the input editor, or load the built-in example.
  2. Set speaking speed.
  3. Run offline neural text-to-speech, review the output, then use the available copy or download controls.

د دې وسیلې وړتیاوې

Generates an English WAV download locally with Supertonic. Downloads the model and voice vector on first run. Source text stays on-device.

Speaking speed
Range: 0.8 to 1.2

محدودیتونه او پروسس

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.

د بېلګې سرچینه
A useful tool makes everyday work easier. Review the results before sharing.

عامې پوښتنې

Can I download the generated speech as an audio file?

Yes. The voice is rendered locally and returned as speech.wav, a 16-bit mono WAV file. This differs from browser read-aloud, which only plays audio and cannot be saved.

Does offline neural text-to-speech support other languages?

No. The voice is English only, and input must be between 1 and 1,000 characters. For other languages, a browser speech tool that uses installed system voices may be an option.

Does offline text-to-speech work without internet?

The first run needs a connection to download the model and voice vector. After that, models may remain in the browser cache, and the text itself is processed on your device rather than sent to a speech service.

Where is my input processed?

Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.

What are the input limits?

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.