Guide
How to handle a recording with multiple speakers
Automatic speaker labels are the most oversold feature in transcription. Here is what the technology actually does, and how to get an accurate multi-speaker transcript anyway.
"Automatic speaker identification" appears on the feature list of most transcription products. It is real technology, it is genuinely useful sometimes, and it is nowhere near as reliable as the marketing implies.
If you are transcribing a meeting, a panel or a group interview, it is worth knowing what you are actually getting.
What speaker diarisation does #
Diarisation is a separate step from transcription. Rather than recognising words, it clusters the audio by voice characteristics — pitch, timbre, cadence — and decides "this stretch sounds like the same person as that stretch". Then it labels the clusters Speaker 1, Speaker 2, and so on.
Two things follow from that description, and they explain almost every complaint about the feature.
It does not know who anyone is. It produces "Speaker 1", not "Dr Chen". Assigning names is always a human step; a tool that shows real names is either using metadata from a platform that knew who was unmuted, or you told it.
It works on voice similarity, not meaning. Two people with similar voices merge into one label. One person whose voice changes — leaning away from the mic, becoming animated, joining by phone — splits into two.
When it works, and when it does not #
Reasonably reliable: two people, distinct voices, one microphone each or a close shared mic, taking turns without much overlap. A two-person podcast recorded properly is close to the best case.
Unreliable: four or more people. A conference room with one recorder in the middle. Anyone on speakerphone. Similar voices — this is a real and awkward failure mode, and it disproportionately affects groups of people of the same gender and similar age, which is most meetings.
Essentially hopeless: crosstalk. When two people speak simultaneously, the audio contains both, and the clustering has to pick one. This is precisely where transcripts get attributed to the wrong person — and it is exactly the moment in an argument or a negotiation that you most want to be accurate about.
The failure that matters #
An error in transcription is usually obvious. You read a sentence that makes no sense and you go and check it.
An error in speaker attribution is invisible. The sentence is correct, well-formed, and attributed to the wrong person. Nothing about it looks wrong. If you build a quote, a meeting record or a legal note on it, the mistake propagates silently.
That asymmetry is why I would not build a workflow that depends on automatic labels being right.
What to do instead #
If the platform knew who was speaking, use that #
This is the one case with a genuinely better answer. Zoom, Teams and Meet know who is unmuted, so their own transcripts can label speakers from metadata rather than by guessing from audio. If accurate attribution is the priority and the meeting was cloud-recorded on a plan that includes transcription, that is the most reliable source available. Our meeting recordings guide covers the trade-offs.
Record so the problem is smaller #
- One microphone per person if the recording matters. Separate tracks make attribution trivial and eliminate diarisation entirely.
- A round-table where people are asked not to talk over each other produces a far better transcript than one where they do. Saying this at the start of a recorded meeting is normal and helps everyone.
- Have people identify themselves the first time they speak — "Priya here" — which gives you an anchor in the transcript.
Label it yourself, at the point of use #
For most work you do not need the whole transcript labelled. You need the six passages you are going to quote or act on labelled, correctly.
The practical method: work in a segment view with timestamps, find the passages that matter, play the audio for those lines, and write the name in as you confirm it. You are listening to those moments anyway to check the words. Labelling them costs nothing extra and is the only method that is actually accurate.
Use structure rather than labels #
For a lot of meeting work, a transcript with timestamps and no speaker labels is perfectly usable — you are searching for decisions and actions, not building a play script. Do not spend an afternoon fixing labels you never needed.
Where we stand #
Offline Transcription does not do automatic speaker identification. Being straightforward about why: the Whisper family does not do diarisation, adding it means a second model with the reliability characteristics described above, and shipping labels that are wrong in an invisible way is worse for the kind of work our users do than shipping no labels at all.
What it does provide is the workflow that makes manual labelling fast: every segment carries its timestamps, and you can play the audio behind any line to confirm who said it before you write a name next to it.
If automatic labels are a hard requirement for you, that is a legitimate reason to use a different tool, and I would rather say so than pretend otherwise.
The short version #
Diarisation clusters voices; it does not know people. It is decent for two distinct speakers and unreliable for groups, similar voices and crosstalk — and its errors are invisible in a way transcription errors are not. Record separate tracks if it matters, use platform metadata if you have it, and otherwise label the handful of passages you actually use, by ear, at the point of use.
Offline Transcription
Transcribe audio and video entirely on your own Mac. Nothing is uploaded, there is no account, and there is no per-minute meter.