ToolCabana
Language preview: no interface translation is available for this language yet. Showing English. Use English

Speech to text

Transcribe speech with a supported browser service.

Favorites are saved in this browser. Find them in My favorites.

Speech to text

Microphone audio is processed in this browser after a model download
Clear your input, files and result, and restore the default settings.

Your voice, in text

Choose a language and recording limit. Your browser will ask for microphone permission.

How to use Speech to text

Speech to text records from your microphone and returns a written transcript you can copy or download as transcript.txt. The default Local browser mode records up to the chosen limit and transcribes with a Whisper model in the page; the Browser speech service mode uses the browser's recognition API instead. You set a BCP 47 language and a recording limit of 5 to 120 seconds.

  1. Enter the requested values in the settings panel.
  2. Adjust the basic settings, then open Advanced mode for export and fine-tuning options. Custom settings stay active when the panel is closed.
  3. Run speech to text, review the output, then use the available copy or download controls.

What this tool supports

Requires a supported browser speech-recognition API and microphone permission. The recording limit starts after recognition begins. Final recognized text can be copied or downloaded; empty and permission-denied sessions produce an error. Your browser may send audio to its speech provider.

Processing
Local browser · Browser speech service
Recognition language (BCP 47)
Set this value for your task.
Recording limit (seconds)
Range: 5 to 120

Limits and processing

Choose a recording limit of 5–120 seconds. Recognition must start within 15 seconds, and transcript text is bounded to 30,000 characters. Browser support, microphone permission and a functioning speech provider are required.

Example source
A handy tool for every little task.

Frequently asked questions

What is the difference between Local browser and Browser speech service?

Local browser records a clip, then transcribes it with a Whisper model downloaded to the page, so audio stays in the browser. Browser speech service uses the browser's speech recognition API, which may send audio to its provider and is not available in every browser.

Can I upload an audio file to transcribe?

No. This tool works from live microphone input in both modes. Recording starts after you grant microphone permission and stops at the recording limit or when you cancel. If permission is denied, the tool shows an error.

How long can a speech to text recording be?

Set the recording limit between 5 and 120 seconds; the default is 30. In local mode the clip must be at most about two minutes and 20 MB. With the browser service, the transcript is capped at 30,000 characters.

Where is my input processed?

Your browser may use a remote speech provider. Its own permission and privacy settings apply.

What are the input limits?

Choose a recording limit of 5–120 seconds. Recognition must start within 15 seconds, and transcript text is bounded to 30,000 characters. Browser support, microphone permission and a functioning speech provider are required.