Meeting notes from audio
Export a timestamped transcript and extract action candidates.
Meeting notes from audio
Drop your files here
or choose files from your device
Como usar Meeting notes from audio
Meeting notes from audio transcribes a recording in the browser and produces a Markdown file with timestamped transcript sections plus a list of action candidates. Action candidates are transcript segments containing phrases such as will, need to, should, action, please or must. You also get a plain transcript.txt, and can process from 1 to 120 seconds.
- Choose a source file. If the tool accepts multiple files, arrange them in the order you want.
- Set language code (blank = detect), maximum seconds.
- Run meeting notes from audio, review the output, then use the available copy or download controls.
Que admite esta ferramenta
Local Whisper transcription produces timestamped transcript sections. Meeting action candidates use phrase matching; these are extractive drafts, not a generated summary. Review transcription errors.
- Language code (blank = detect)
- Set this value for your task.
- Maximum seconds
- Range: 1 to 120
Límites e procesamento
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.
Fonte de exemplo
A useful tool makes everyday work easier. Review the results before sharing.
Preguntas frecuentes
Does meeting notes from audio write a summary?
No. It does not generate new text. The notes contain the transcript split into timestamped sections and a list of segments copied word for word because they contain action-like phrases. Review them as a draft.
How are action items found in the transcript?
A segment becomes an action candidate if it contains will, need to, should, action, please or must as whole words. If none match, the notes say no action phrases were detected. Real tasks phrased differently may be missed.
How long a meeting recording can I use?
The tool transcribes from the start of the file up to Maximum seconds, which ranges from 1 to 120 with a default of 30. For longer meetings, split the recording into two-minute parts.
Where is my input processed?
Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.
What are the input limits?
Image inputs: up to 32 megapixels, batches up to 10 images. Audio/video uploads: 50 MB; transcription: 120 seconds. Edited videos: 20 seconds, 4 fps and 640 pixels wide. Translation: English/Spanish, 30,000 characters; subtitles: 200 cues. Neural speech: 1,000 English characters. Model downloads and memory requirements vary by device.