Evaluation
What is word error rate and how should data buyers use it?
Updated 2026-08-01 · 4 min read

Short answer
Word error rate is the sum of substitutions, insertions and deletions divided by the number of reference words, so a WER of 15% means fifteen errors per hundred words. For data buyers, the number is only meaningful alongside three things: the evaluation set it was measured on, the transcription convention used for the reference, and the slice breakdown. A model at 12% aggregate WER that sits at 30% on one dialect is not a 12% model for the users in that region, and a WER computed against inconsistently transcribed references measures your annotation quality rather than your model.
Key takeaways
- WER without a named evaluation set is not a number, it is a claim.
- Reference transcription convention changes WER by several points on code-mixed Indian speech.
- Always read WER per slice: language, dialect, device, noise condition.
How it is computed
WER counts the minimum edits needed to turn the hypothesis into the reference, normalised by reference length. Insertions can push WER above 100%, which surprises people the first time they see it on noisy audio.
Character error rate is often more informative for Indian scripts with rich morphology, where a single inflectional difference marks a whole word wrong.
Why the reference matters as much as the model
If one transcriber writes an English loanword in Roman script and another writes it in Devanagari, the same audio yields two references and two WERs. On code-mixed Indian speech this alone can move the number by several points.
This is why the transcription style guide is a modelling artefact, not a clerical one, and why two-pass QA against that guide is part of the deliverable.

Reading it as a buyer
Ask three questions of any WER figure: which evaluation set, which convention, and what does the worst slice look like. A vendor or team that cannot answer all three is reporting a marketing number.
Setting a WER target in a contract
Targets are achievable when they are set against a named, frozen evaluation set with an agreed convention and a defined measurement script. Set against 'good quality audio', they are unenforceable and everyone knows it.
Frequently asked questions
What is word error rate and how should data buyers use it?
Word error rate is the sum of substitutions, insertions and deletions divided by the number of reference words, so a WER of 15% means fifteen errors per hundred words. For data buyers, the number is only meaningful alongside three things: the evaluation set it was measured on, the transcription convention used for the reference, and the slice breakdown. A model at 12% aggregate WER that sits at 30% on one dialect is not a 12% model for the users in that region, and a WER computed against inconsistently transcribed references measures your annotation quality rather than your model.
How it is computed?
WER counts the minimum edits needed to turn the hypothesis into the reference, normalised by reference length. Insertions can push WER above 100%, which surprises people the first time they see it on noisy audio.
Why the reference matters as much as the model?
If one transcriber writes an English loanword in Roman script and another writes it in Devanagari, the same audio yields two references and two WERs. On code-mixed Indian speech this alone can move the number by several points.
Reading it as a buyer?
Ask three questions of any WER figure: which evaluation set, which convention, and what does the worst slice look like. A vendor or team that cannot answer all three is reporting a marketing number.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.