Speaker-labelled transcript
Label transcript segments using a configured diarization model.
Speaker-labelled transcript
Drop your files here
or choose files from your device
របៀបប្រើ Speaker-labelled transcript
The speaker-labelled transcript tool transcribes up to 10 seconds of audio and assigns each word segment to an anonymous speaker slot. It matches Whisper word timestamps to Pyannote speaker segmentation, both run in the browser, and supports up to three slots plus overlap labels such as Speakers 1 + 2. Results download as speakers.json with segments and speaker windows.
- Choose a source file. If the tool accepts multiple files, arrange them in the order you want.
- Set language code (blank = detect), maximum seconds.
- Run speaker-labelled transcript, review the output, then use the available copy or download controls.
អ្វីដែលឧបករណ៍នេះគាំទ្រ
Browser Whisper word timestamps are matched to Pyannote segmentation within a 10-second excerpt, with up to three anonymous speaker slots. Labels do not identify people or persist across runs. Review overlaps and short turns.
- Language code (blank = detect)
- Set this value for your task.
- Maximum seconds
- Range: 1 to 10
ដែនកំណត់ និងការដំណើរការ
Up to 10 seconds, three anonymous speaker slots. Labels apply only within the selected excerpt and are not speaker identities. First inference downloads transcription and segmentation models.
ប្រភពឧទាហរណ៍
A useful tool makes everyday work easier. Review the results before sharing.
សំណួរដែលសួរញឹកញាប់
Can this tool identify who is speaking by name?
No. Labels like Speaker 1 are anonymous slots within the analyzed excerpt. They do not identify people and are not comparable between separate runs, as the tool's output notice states.
Why is the speaker transcript limited to 10 seconds?
Speaker segmentation is run on a single short excerpt with up to three speaker slots. Maximum seconds can be set from 1 to 10, starting at the beginning of the file. Longer recordings need to be cut into short clips.
What does Unknown mean in the speaker labels?
A segment is labelled Unknown when its timing does not overlap any speaker window found by segmentation. Short turns and overlapping speech may also be misassigned, so review each label.
Where is my input processed?
Source files and text stay in this browser. The first run downloads public models from Hugging Face and local runtime assets from the site. Models may remain in the browser cache. No source upload to Vercel, Supabase or a conversion service is required.
What are the input limits?
Up to 10 seconds, three anonymous speaker slots. Labels apply only within the selected excerpt and are not speaker identities. First inference downloads transcription and segmentation models.