aidataservices.inAI data collection · India

Evaluation

How do you measure transcription quality in a delivered corpus?

Updated 2026-08-01 · 4 min read

Transcriber timestamping Indian language audio — illustration for: How do you measure transcription quality in a delivered corpus?

Short answer

Measure it by sampling every batch, re-transcribing the sample independently, and computing error against an adjudicated reference — not by asking the vendor whether they checked. Report accuracy separately for overall words, numerals, named entities and code-mixed spans, because a single aggregate number hides the categories that break products. Set the threshold in the contract against a named convention, and specify that failing batches are re-transcribed rather than credited.

Key takeaways

The argument at a glance1Sample and independently re-transcribe; self-reported QA is not measurement.2Report per-category accuracy — numerals and entities behave differently from prose.3Adjudicate disagreements; two annotators disagreeing means the reference is not fixed.
  • Sample and independently re-transcribe; self-reported QA is not measurement.
  • Report per-category accuracy — numerals and entities behave differently from prose.
  • Adjudicate disagreements; two annotators disagreeing means the reference is not fixed.

Sampling design

Random samples stratified by language, speaker, condition and transcriber, sized so the estimate is meaningful per stratum. Convenience samples of easy audio produce reassuring numbers and no information.

Adjudication

Where the independent transcription differs from the delivered one, a third native reviewer adjudicates against the style guide. Many apparent errors are guideline ambiguities, and those get fixed in the guideline, not blamed on the transcriber.

Structured dataset packages ready for delivery — evaluation context for How do you measure transcription quality in a delivered corpus
Structured dataset packages ready for delivery

Categories to report

Overall word accuracy, numerals and units, named entities, code-mixed spans, and non-speech event tagging. Products fail on the specific categories, not on the average.

Contracting the threshold

State the threshold, the convention, the sampling method and the remedy. Without all four, a quality clause is decorative.

Frequently asked questions

How do you measure transcription quality in a delivered corpus?

Measure it by sampling every batch, re-transcribing the sample independently, and computing error against an adjudicated reference — not by asking the vendor whether they checked. Report accuracy separately for overall words, numerals, named entities and code-mixed spans, because a single aggregate number hides the categories that break products. Set the threshold in the contract against a named convention, and specify that failing batches are re-transcribed rather than credited.

Sampling design?

Random samples stratified by language, speaker, condition and transcriber, sized so the estimate is meaningful per stratum. Convenience samples of easy audio produce reassuring numbers and no information.

Adjudication?

Where the independent transcription differs from the delivered one, a third native reviewer adjudicates against the style guide. Many apparent errors are guideline ambiguities, and those get fixed in the guideline, not blamed on the transcriber.

Categories to report?

Overall word accuracy, numerals and units, named entities, code-mixed spans, and non-speech event tagging. Products fail on the specific categories, not on the average.

Related reading

Turn this into a dataset specification

Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.

Request a dataset quote