Multimodal AI
Speaker diarization
Determining which stretches of an audio recording belong to different speakers.
Example
A meeting written record of speech labels alternating speech as Speaker 1 and Speaker 2.
Why people use it
It helps organize a recording into stretches spoken by different people.
What you'll hear
“Which parts belong to each speaker?”
What this means for you
Review speaker changes before attributing statements to named participants.
Can you control it?
Sometimes
Sometimes. Your choices depend on the tool and your access. The settings available to an everyday user may differ from those available to the people running it.
Common questions
- Does diarization tell you the speakers' names?
- Not by itself. It separates voices without necessarily identifying the people.
- Can overlapping speech cause trouble?
- Yes. When people speak together, the system may merge voices or assign words incorrectly.
- Can it split one person into two speakers?
- Yes. A changed microphone, tone or recording quality can make the same voice appear different.