Home · Glossary · What is word error rate (WER)?
Glossary

What is word error rate (WER)?

When a transcription tool claims to be accurate, the number behind it is usually word error rate. It is useful, but only if you know what it counts.

Definition: Word error rate (WER) measures transcription accuracy: the number of substituted, deleted and inserted words divided by the number of words in a correct reference transcript. Lower is better.

The formula

WER compares a machine transcript (the hypothesis) with a correct, human-checked transcript (the reference). The two are aligned word by word, choosing the alignment that needs the fewest edits, and each difference is counted as one of three errors:

  • Substitution (S): a word was replaced by a different one.
  • Deletion (D): a word in the reference is missing.
  • Insertion (I): an extra word appears that was not said.

A worked example

WER = (S + D + I) / N, where N is the number of words in the reference.

Reference: “the budget is due on friday” (6 words). Hypothesis: “the budget was due friday.” “is” became “was” (1 substitution) and “on” is missing (1 deletion). Nothing was added (0 insertions), so WER = (1 + 1 + 0) / 6, about 33%.

Because insertions are counted but N is not, WER can exceed 100%. A transcript that invents many words over a quiet or noisy stretch can have more errors than there were words in the reference. So WER is not simply “percent wrong,” and “accuracy = 100% minus WER” is only a rough shorthand.

Why one number misleads

  • Normalization changes the score. Whether “ten” vs “10,” “OK” vs “okay,” capitalization and punctuation count as errors can move WER a lot. Two published numbers are only comparable if they normalized the same way.
  • Every word weighs the same. Getting a client's name wrong costs exactly as much as missing an “um,” though one matters far more.
  • The test audio decides the result. Clean read speech scores far better than crosstalk, accents, phone lines or jargon. A low WER on a benchmark says little about your recordings.
  • Word boundaries vary by language. For languages written without spaces, character error rate (CER) is used instead.

Measuring it on your own audio

The most useful WER is one you measure on your own material. Pick two or three short clips that represent your usual recordings, correct their transcripts by hand to make references, and compare each tool's output using the formula above, normalizing both sides the same way. Then read the errors, not just the score: which names and terms were wrong?

MediaFind does not publish a WER figure. It lets you choose a transcription model that fits your machine, trading speed for accuracy, and a names & terms list keeps people and products spelled right, which targets exactly the errors WER treats as minor. See transcription, how we choose speech models and transcribing video offline on a Mac.

Frequently asked questions

Can word error rate be higher than 100%?

Yes. Insertions add to the error count but not to the number of reference words, so a transcript with many invented words can score above 100%.

What is a good word error rate?

It depends on the audio. Clean, single-speaker recordings can score far lower than noisy multi-speaker calls with the same tool, so compare tools on your own recordings.

What is the difference between WER and CER?

WER counts errors per word. Character error rate counts per character and is used for languages without spaces between words, or to judge spelling closeness.

Search your whole library on your own computer.

Free for up to 10 files, with a 7-day Pro trial. No account, nothing uploaded.

Download for macOS View pricing