aidataservices.inAI data collection · India

Annotation

How do you write annotation guidelines for Indian speech data?

Updated 2026-08-01 · 4 min read

Annotator labelling audio segments and speaker turns — illustration for: How do you write annotation guidelines for Indian speech data?

Short answer

A usable annotation guideline for Indian speech fixes six decisions in writing with worked examples: script and orthography per language, how embedded English is written, how numerals, dates and currency are rendered, how disfluencies and repairs are marked, how non-speech events are tagged, and how named entities are handled. Ambiguity in any of these produces inconsistent references, which inflates measured error and silently caps model quality. The guideline is a living document with a change log, because the first two weeks of production always surface cases nobody anticipated.

Key takeaways

The argument at a glance1Six decisions cover most inconsistency: script, English tokens, numerals, disfluencies, events, entities.2Worked examples beat rules — annotators pattern-match.3Version the guideline and re-QA earlier batches when a rule changes.
  • Six decisions cover most inconsistency: script, English tokens, numerals, disfluencies, events, entities.
  • Worked examples beat rules — annotators pattern-match.
  • Version the guideline and re-QA earlier batches when a rule changes.

The six decisions

Script and orthography per language, including how loanwords are spelt. Whether embedded English appears in Roman or native script. Whether numerals are written as digits or as spoken words. How filler words, false starts and repairs are transcribed or omitted. Which non-speech events are tagged and how. Whether named entities are marked and normalised.

Each decision should carry three worked examples: a clear case, a borderline case and a case that is explicitly out of scope.

Running the guideline in production

Collect annotator questions daily in the first two weeks, resolve them into new examples, and version the document. When a rule changes materially, re-QA the batches produced under the old rule rather than shipping a corpus with two conventions inside it.

Audio waveforms being prepared as ASR training data — annotation context for How do you write annotation guidelines for Indian speech data
Audio waveforms being prepared as ASR training data

Measuring adherence

Sample every batch for guideline adherence separately from accuracy. A transcriber can be accurate and non-compliant, and a corpus of accurate-but-inconsistent transcripts trains worse than a slightly less accurate consistent one.

Handing it to your team

Deliver the guideline with the corpus. Your team needs it to interpret the data, to extend the corpus later, and to write evaluation references that are comparable to the training references.

Frequently asked questions

How do you write annotation guidelines for Indian speech data?

A usable annotation guideline for Indian speech fixes six decisions in writing with worked examples: script and orthography per language, how embedded English is written, how numerals, dates and currency are rendered, how disfluencies and repairs are marked, how non-speech events are tagged, and how named entities are handled. Ambiguity in any of these produces inconsistent references, which inflates measured error and silently caps model quality. The guideline is a living document with a change log, because the first two weeks of production always surface cases nobody anticipated.

The six decisions?

Script and orthography per language, including how loanwords are spelt. Whether embedded English appears in Roman or native script. Whether numerals are written as digits or as spoken words. How filler words, false starts and repairs are transcribed or omitted. Which non-speech events are tagged and how. Whether named entities are marked and normalised.

Running the guideline in production?

Collect annotator questions daily in the first two weeks, resolve them into new examples, and version the document. When a rule changes materially, re-QA the batches produced under the old rule rather than shipping a corpus with two conventions inside it.

Measuring adherence?

Sample every batch for guideline adherence separately from accuracy. A transcriber can be accurate and non-compliant, and a corpus of accurate-but-inconsistent transcripts trains worse than a slightly less accurate consistent one.

Related reading

Turn this into a dataset specification

Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.

Request a dataset quote