Annotation
How do you write annotation guidelines for Indian speech data?
Updated 2026-08-01 · 4 min read

Short answer
A usable annotation guideline for Indian speech fixes six decisions in writing with worked examples: script and orthography per language, how embedded English is written, how numerals, dates and currency are rendered, how disfluencies and repairs are marked, how non-speech events are tagged, and how named entities are handled. Ambiguity in any of these produces inconsistent references, which inflates measured error and silently caps model quality. The guideline is a living document with a change log, because the first two weeks of production always surface cases nobody anticipated.
Key takeaways
- Six decisions cover most inconsistency: script, English tokens, numerals, disfluencies, events, entities.
- Worked examples beat rules — annotators pattern-match.
- Version the guideline and re-QA earlier batches when a rule changes.
The six decisions
Script and orthography per language, including how loanwords are spelt. Whether embedded English appears in Roman or native script. Whether numerals are written as digits or as spoken words. How filler words, false starts and repairs are transcribed or omitted. Which non-speech events are tagged and how. Whether named entities are marked and normalised.
Each decision should carry three worked examples: a clear case, a borderline case and a case that is explicitly out of scope.
Running the guideline in production
Collect annotator questions daily in the first two weeks, resolve them into new examples, and version the document. When a rule changes materially, re-QA the batches produced under the old rule rather than shipping a corpus with two conventions inside it.

Measuring adherence
Sample every batch for guideline adherence separately from accuracy. A transcriber can be accurate and non-compliant, and a corpus of accurate-but-inconsistent transcripts trains worse than a slightly less accurate consistent one.
Handing it to your team
Deliver the guideline with the corpus. Your team needs it to interpret the data, to extend the corpus later, and to write evaluation references that are comparable to the training references.
Frequently asked questions
How do you write annotation guidelines for Indian speech data?
A usable annotation guideline for Indian speech fixes six decisions in writing with worked examples: script and orthography per language, how embedded English is written, how numerals, dates and currency are rendered, how disfluencies and repairs are marked, how non-speech events are tagged, and how named entities are handled. Ambiguity in any of these produces inconsistent references, which inflates measured error and silently caps model quality. The guideline is a living document with a change log, because the first two weeks of production always surface cases nobody anticipated.
The six decisions?
Script and orthography per language, including how loanwords are spelt. Whether embedded English appears in Roman or native script. Whether numerals are written as digits or as spoken words. How filler words, false starts and repairs are transcribed or omitted. Which non-speech events are tagged and how. Whether named entities are marked and normalised.
Running the guideline in production?
Collect annotator questions daily in the first two weeks, resolve them into new examples, and version the document. When a rule changes materially, re-QA the batches produced under the old rule rather than shipping a corpus with two conventions inside it.
Measuring adherence?
Sample every batch for guideline adherence separately from accuracy. A transcriber can be accurate and non-compliant, and a corpus of accurate-but-inconsistent transcripts trains worse than a slightly less accurate consistent one.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.