aidataservices.inAI data collection · India

Evaluation

What is word error rate and how should data buyers use it?

Updated 2026-08-01 · 4 min read

AI team reviewing dataset dashboards — illustration for: What is word error rate and how should data buyers use it?

Short answer

Word error rate is the sum of substitutions, insertions and deletions divided by the number of reference words, so a WER of 15% means fifteen errors per hundred words. For data buyers, the number is only meaningful alongside three things: the evaluation set it was measured on, the transcription convention used for the reference, and the slice breakdown. A model at 12% aggregate WER that sits at 30% on one dialect is not a 12% model for the users in that region, and a WER computed against inconsistently transcribed references measures your annotation quality rather than your model.

Key takeaways

The argument at a glance1WER without a named evaluation set is not a number, it is a claim.2Reference transcription convention changes WER by several points on code-mixed Indian speech.3Always read WER per slice: language, dialect, device, noise condition.
  • WER without a named evaluation set is not a number, it is a claim.
  • Reference transcription convention changes WER by several points on code-mixed Indian speech.
  • Always read WER per slice: language, dialect, device, noise condition.

How it is computed

WER counts the minimum edits needed to turn the hypothesis into the reference, normalised by reference length. Insertions can push WER above 100%, which surprises people the first time they see it on noisy audio.

Character error rate is often more informative for Indian scripts with rich morphology, where a single inflectional difference marks a whole word wrong.

Why the reference matters as much as the model

If one transcriber writes an English loanword in Roman script and another writes it in Devanagari, the same audio yields two references and two WERs. On code-mixed Indian speech this alone can move the number by several points.

This is why the transcription style guide is a modelling artefact, not a clerical one, and why two-pass QA against that guide is part of the deliverable.

Two speakers recording natural conversational speech data — evaluation context for What is word error rate and how should data buyers use it
Two speakers recording natural conversational speech data

Reading it as a buyer

Ask three questions of any WER figure: which evaluation set, which convention, and what does the worst slice look like. A vendor or team that cannot answer all three is reporting a marketing number.

Setting a WER target in a contract

Targets are achievable when they are set against a named, frozen evaluation set with an agreed convention and a defined measurement script. Set against 'good quality audio', they are unenforceable and everyone knows it.

Frequently asked questions

What is word error rate and how should data buyers use it?

Word error rate is the sum of substitutions, insertions and deletions divided by the number of reference words, so a WER of 15% means fifteen errors per hundred words. For data buyers, the number is only meaningful alongside three things: the evaluation set it was measured on, the transcription convention used for the reference, and the slice breakdown. A model at 12% aggregate WER that sits at 30% on one dialect is not a 12% model for the users in that region, and a WER computed against inconsistently transcribed references measures your annotation quality rather than your model.

How it is computed?

WER counts the minimum edits needed to turn the hypothesis into the reference, normalised by reference length. Insertions can push WER above 100%, which surprises people the first time they see it on noisy audio.

Why the reference matters as much as the model?

If one transcriber writes an English loanword in Roman script and another writes it in Devanagari, the same audio yields two references and two WERs. On code-mixed Indian speech this alone can move the number by several points.

Reading it as a buyer?

Ask three questions of any WER figure: which evaluation set, which convention, and what does the worst slice look like. A vendor or team that cannot answer all three is reporting a marketing number.

Related reading

Turn this into a dataset specification

Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.

Request a dataset quote