Evaluation
How do you measure transcription quality in a delivered corpus?
Updated 2026-08-01 · 4 min read

Short answer
Measure it by sampling every batch, re-transcribing the sample independently, and computing error against an adjudicated reference — not by asking the vendor whether they checked. Report accuracy separately for overall words, numerals, named entities and code-mixed spans, because a single aggregate number hides the categories that break products. Set the threshold in the contract against a named convention, and specify that failing batches are re-transcribed rather than credited.
Key takeaways
- Sample and independently re-transcribe; self-reported QA is not measurement.
- Report per-category accuracy — numerals and entities behave differently from prose.
- Adjudicate disagreements; two annotators disagreeing means the reference is not fixed.
Sampling design
Random samples stratified by language, speaker, condition and transcriber, sized so the estimate is meaningful per stratum. Convenience samples of easy audio produce reassuring numbers and no information.
Adjudication
Where the independent transcription differs from the delivered one, a third native reviewer adjudicates against the style guide. Many apparent errors are guideline ambiguities, and those get fixed in the guideline, not blamed on the transcriber.

Categories to report
Overall word accuracy, numerals and units, named entities, code-mixed spans, and non-speech event tagging. Products fail on the specific categories, not on the average.
Contracting the threshold
State the threshold, the convention, the sampling method and the remedy. Without all four, a quality clause is decorative.
Frequently asked questions
How do you measure transcription quality in a delivered corpus?
Measure it by sampling every batch, re-transcribing the sample independently, and computing error against an adjudicated reference — not by asking the vendor whether they checked. Report accuracy separately for overall words, numerals, named entities and code-mixed spans, because a single aggregate number hides the categories that break products. Set the threshold in the contract against a named convention, and specify that failing batches are re-transcribed rather than credited.
Sampling design?
Random samples stratified by language, speaker, condition and transcriber, sized so the estimate is meaningful per stratum. Convenience samples of easy audio produce reassuring numbers and no information.
Adjudication?
Where the independent transcription differs from the delivered one, a third native reviewer adjudicates against the style guide. Many apparent errors are guideline ambiguities, and those get fixed in the guideline, not blamed on the transcriber.
Categories to report?
Overall word accuracy, numerals and units, named entities, code-mixed spans, and non-speech event tagging. Products fail on the specific categories, not on the average.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.