ToolCabana

Meeting notes from audio

Export a timestamped transcript and extract action candidates.

Favorites are saved in this browser. Find them in My favorites.

Meeting notes from audio

Your input stays in this browser unless stated otherwise
Clear your input, files and result, and restore the default settings.

Drop your files here

or choose files from your device

Conas úsáid a bhaint as Meeting notes from audio

Meeting notes from audio transcribes a recording in the browser and produces a Markdown file with timestamped transcript sections plus a list of action candidates. Action candidates are transcript segments containing phrases such as will, need to, should, action, please or must. You also get a plain transcript.txt, and can process from 1 to 120 seconds.

  1. Choose a source file. If the tool accepts multiple files, arrange them in the order you want.
  2. Set language code (blank = detect), maximum seconds.
  3. Run meeting notes from audio, review the output, then use the available copy or download controls.

Cad a thacaíonn an uirlis seo leis

Local Whisper transcription produces timestamped transcript sections. Meeting action candidates use phrase matching; these are extractive drafts, not a generated summary. Review transcription errors.

Language code (blank = detect)
Set this value for your task.
Maximum seconds
Range: 1 to 120

Teorainneacha agus próiseáil

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.

Foinse shamplach
A useful tool makes everyday work easier. Review the results before sharing.

Ceisteanna coitianta

Does meeting notes from audio write a summary?

No. It does not generate new text. The notes contain the transcript split into timestamped sections and a list of segments copied word for word because they contain action-like phrases. Review them as a draft.

How are action items found in the transcript?

A segment becomes an action candidate if it contains will, need to, should, action, please or must as whole words. If none match, the notes say no action phrases were detected. Real tasks phrased differently may be missed.

How long a meeting recording can I use?

The tool transcribes from the start of the file up to Maximum seconds, which ranges from 1 to 120 with a default of 30. For longer meetings, split the recording into two-minute parts.

Where is my input processed?

Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.

What are the input limits?

Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.