Speaker diarization is the step that works out who spoke when in a recording. It finds where speech is, turns each stretch into a numeric description of the voice, and groups similar voices under labels such as Speaker 1 and Speaker 2. It does not know names, and it struggles with overlapping talk, very short turns and similar voices.
Speaker diarization answers one question about a recording: who spoke when. A 2021 review of the field defines it as “a task to label audio or video recordings with classes that correspond to speaker identity, or in short, a task to identify ‘who spoke when’”. It is the difference between a wall of text and a transcript you can read as a conversation.
It is also the part of an AI transcript most likely to be wrong in a way that matters. A misheard word is a typo. A sentence given to the wrong person can change who agreed to what.
What does diarization produce?
A timeline. Each entry has a start time, an end time and a label:
| Start | End | Label |
|---|---|---|
| 0:00 | 0:07 | Speaker 1 |
| 0:07 | 0:11 | Speaker 2 |
| 0:11 | 0:19 | Speaker 1 |
The labels are anonymous. The system knows that the voice at 0:00 and the voice at 0:11 are the same, and that the voice at 0:07 is different. It does not know anyone’s name.
Speech recognition produces a second timeline: which word was said when. A speaker-labeled transcript comes from laying one over the other, so each word takes the label of whoever was speaking at that moment.
How does speaker diarization work?
Most systems do a version of these steps. The review above describes the classic pipeline, and open-source toolkits such as pyannote.audio package them as “neural building blocks”: speech activity detection, speaker change detection, overlapped speech detection and speaker embedding.
1. Find the speech
First the system separates speech from everything else: silence, a door, a keyboard, music. This is called voice activity detection. Errors here show up later as words with no speaker or noise treated as talk.
2. Segmentation: cut the speech into single-speaker pieces
Next it looks for the points where one voice stops and another starts, and cuts the audio there. The goal is pieces that each contain one person. Modern models do this at fine time resolution. The segmentation model described by Bredin and Laurent works on five-second chunks and makes a decision every 16 milliseconds, and it also flags regions where two people talk at once.
3. Embeddings: describe each voice as numbers
Each piece goes through a neural network that outputs a list of numbers, called a speaker embedding. Think of it as coordinates for a voice. Pieces spoken by the same person land close together, and pieces from different people land further apart. The embedding captures how a voice sounds, not what was said.
A voiceprint is the same idea: an embedding kept as a reference for one person.
4. Clustering: group the pieces
Finally the system groups embeddings that sit close together. Each group becomes a speaker. Usually it does not know in advance how many people are in the room, so it has to decide the number of groups too. Too few and two people merge into one. Too many and one person splits into two.
Some newer systems train a single model to do all of this at once, known as end-to-end neural diarization. The steps still describe what has to be worked out.
Diarization vs speaker identification vs speaker recognition
These terms get mixed up. They are different jobs.
| Term | Question it answers | Needs voices saved in advance? |
|---|---|---|
| Speaker diarization | Who spoke when, with anonymous labels | No |
| Speaker identification | Which known person is this voice | Yes |
| Speaker verification | Is this voice the person it claims to be | Yes |
| Speech recognition | What words were said | No |
“Speaker recognition” is used loosely for identification and verification together. It is not the same as speech recognition, which is about words.
In practice, products combine them. Diarization separates the voices in a new recording. Then each group can be compared with voiceprints saved earlier, and if one matches, “Speaker 2” becomes a name.
Why does diarization get speakers wrong?
Because some audio does not contain enough information to tell voices apart. The usual causes:
- Overlapping speech. When two people talk at once, a single stretch of audio belongs to both. A system that assigns one speaker per moment has to drop one of them. The pyannote.metrics documentation notes that the standard metric takes overlapping speech into account, “potentially leading to increased missed detection” when a system has no overlap detection.
- Short turns. “Yes.” “Right.” “Mm-hm.” A fraction of a second is very little audio to build an embedding from, so these often get attached to the previous speaker.
- Distance from the microphone. A voice far from the phone arrives quieter, with more room echo. Its embedding drifts, and the same person can be split into two speakers, or merged with someone else at the far end.
- Similar voices. Two people with similar pitch and accent sit close together in embedding space, and clustering may merge them.
- Noise and changes in the room. A fan, a passing truck, or someone moving closer to the phone halfway through all change how a voice sounds to the model.
- Many speakers. The more people, the harder it is to pick the right number of groups, especially when some speak only once.
None of this is specific to one product. It is the nature of the problem, and it is why every transcript should be read with the speaker labels in mind.
What is diarization error rate?
Diarization error rate (DER) is the standard way to score a diarization system against a human-made reference. The pyannote.metrics documentation calls it “the de facto standard metric” and defines it as:
DER = (false alarm + missed detection + confusion) / total
- False alarm: time where non-speech was marked as speech.
- Missed detection: time where speech was marked as non-speech.
- Confusion: time given to the wrong speaker.
- Total: the sum of every speaker’s reference speech time.
Lower is better. Research systems are compared on shared datasets and public evaluations. The NIST Rich Transcription evaluation series, for instance, was set up to make transcriptions “more readable by humans and more useful for machines” and included metadata extraction tasks alongside speech-to-text.
Be careful with DER figures in marketing. A number measured on a benchmark of studio-quality recordings tells you little about a phone on a cafe table. The only test that counts is your own recordings.
How do you record so speakers are separated well?
You control more than the software does.
- One phone, in the middle. Place it so every voice is at a similar distance. A phone in front of you makes you loud and everyone else faint.
- Keep it still and clear. Not under papers, not next to a laptop fan, not in a pocket.
- Reduce noise. Close the door. Move away from the coffee machine.
- One person at a time. Overlap is the hardest case. A short pause between speakers helps.
- Let each person speak a full sentence early. A round of introductions gives the system clean audio for every voice.
- Use the right microphone mode. Modes that isolate the nearest voice are good for dictation and bad for groups.
And tell people first. Recording others may require their consent depending on where you are, so tell people before you record. See is it legal to record a meeting.
How Nuvi separates and remembers speakers
Nuvi, the AI note taker for iPhone, iPad and Apple Watch by Ege Beşe, returns a transcript with speakers separated. One voice makes a “Note” and several voices make a “Meeting”.
- Microphone modes. “Room” hears every voice around the phone and is the one for meetings. “Just me” reduces background noise and cleans up the voice closest to the phone, so people across the room may come out faint.
- Naming voices. You can name a speaker with “Save and remember voice”. Nuvi then recognizes that voice in later recordings, on the device. After a meeting it may ask “Which one is you?” or “Do you know these voices?”.
- Where voiceprints live. They are computed while the audio is transcribed, returned to the device and never stored on the server. They stay on the device, are not synced to iCloud and are excluded from backups. That also means remembered voices do not follow you to a second device.
- What can go wrong. Transcripts are AI output. They can be wrong or incomplete, and they can attribute words to the wrong speaker, for the reasons above. No accuracy figure is published for Nuvi, because a single number would not describe your room.
Transcription happens on servers the developer operates, not on the device, and needs an internet connection. The details are on privacy first and in AI note taker privacy. The speech model itself is covered in Whisper multilingual transcription.
Get Nuvi. The features page shows the transcript view, and the interviews and meetings pages show where speaker labels matter most.