How to identify who is speaking in a recording
A transcript without speaker labels is hard to use: you can read what was said, but not who committed to what. Speaker diarization fills that gap.
The short answer
Use a transcription tool that performs speaker diarization. It splits the audio into turns and groups them by voice, labeling them Speaker 1, Speaker 2 and so on. You then attach names to those labels. Better tools remember a named voice, so the next recording with the same person is labeled for you.
How diarization works, in plain terms
The software cuts the audio into short slices, turns each slice into a numeric fingerprint of the voice, and clusters slices that sound alike. Each cluster becomes a speaker. It doesn't know anyone's name; it only knows that these stretches of audio probably came from the same person.
Naming is a separate step. Either you label the clusters by hand, or the tool compares each cluster's fingerprint with voices you've named before and suggests a match.
Where it goes wrong
Expect to glance over the labels and fix a few, especially on short turns like "yeah" or "right".
- Crosstalk. When two people talk at once, one of them usually gets swallowed.
- Similar voices. Two people with similar pitch and accent can be merged into one speaker.
- One voice split in two. A person who moves away from the mic, or switches from calm to animated, can show up as two speakers.
- A single shared microphone in a big room makes everyone sound far away and alike. Separate mics, or one per side of a call, help the most.
How to do it in MediaFind
Speaker labels are part of the free plan, and all of it runs on your computer without an API key.
- Add your recordings to MediaFind. Each one is transcribed and diarized on your machine.
- Open a file: the transcript shows who said what, with timestamps.
- Give a speaker a name once. MediaFind matches that voice in your other recordings, and it also checks against a bundled gallery of well-known voices.
- Review suggestions: confirm a match or dismiss it. Unsure matches stay unnamed rather than guessed.
- Search by speaker to find everything one person said, or combine it with a topic, e.g. everything a named guest said about pricing.
Getting cleaner speaker labels
Most diarization errors start at the recording stage. A few habits help any tool:
- Record each side of a call on its own channel or device when you can.
- Keep people at a steady distance from the microphone.
- Ask people to introduce themselves at the start; it gives you a clean sample to name each voice from.
- Avoid music or TV in the background, which can be mistaken for an extra speaker.
Tips and limitations
Name voices from a clean stretch of speech rather than a noisy one. If one person was split into two speakers, naming both with the same name keeps searches together. Diarization is statistical, so double-check attributions before you quote someone in print.
Voice recognition only uses voices you name, plus the bundled well-known ones; nothing is sent anywhere. Recording people usually requires their consent, and naming voices doesn't change that. For meetings, see meetings; for interviews, transcribing research interviews.
Frequently asked questions
What is speaker diarization?
It's the step that splits a recording into speaker turns and groups them by voice, answering who spoke when. It labels speakers generically; naming them is a separate step.
Can software tell me the actual names of the speakers?
Only if it has heard them before. MediaFind recognizes voices you've named and a bundled set of well-known voices; anyone else stays as a numbered speaker until you name them.
Why does one person show up as two speakers?
Their voice changed enough, from distance, volume or emotion, that the clustering treated it as two people. Giving both labels the same name fixes it for search.
Is speaker identification a Pro feature in MediaFind?
No. Speaker labels and searching by speaker are included in the free plan. Face recognition is the Pro feature.
Search your whole library on your own computer.
Free for up to 10 files, with a 7-day Pro trial. No account, nothing uploaded.
Download for macOS View pricingKeep reading
What is speaker diarization?How software works out who spoke when in a recording, and why the labels are anonymous until you name them. How to transcribe interviews for qualitative research
Turn a stack of recorded interviews into speaker-labeled, timestamped transcripts ready for coding, without sending participant audio to a cloud service. Search every podcast episode for the quote you remember
Find the line a guest said three seasons ago, check who said it, and cut it into a promo clip without re-listening to the back catalog. Search every meeting you've recorded, and stop re-watching them
Get action items and decisions from each meeting, then ask across all of them, with no bot in the call and nothing uploaded. What are word-level timestamps?
Timing for every single word in a transcript, and why that precision matters for search, captions and editing.