aidataservices.inAI data collection · India

Collection guides

How do you collect Hindi speech data for ASR training?

Updated 2026-08-01 · 5 min read

Speaker recording scripted prompts for a speech data collection project — illustration for: How do you collect Hindi speech data for ASR training?

Short answer

Collect Hindi speech data by fixing the corpus specification first — 500–2,000 hours of audio from 1,000–3,000 native speakers, split across Khari Boli, Awadhi, Braj dialects and balanced for gender, age and recording condition. Recruit in Uttar Pradesh, Bihar where the target varieties are actually spoken, record to one written protocol (16 kHz telephony or 48 kHz studio), transcribe in Devanagari with a documented convention for urban hindi speech is hinglish in practice. expect 15-40% english tokens in spontaneous speech: numbers, brands, technology terms, and whole clause switches. any hindi corpus that excludes english tokens will not match production traffic., and accept the corpus against measured WER and metadata completeness rather than studio hours consumed.

Key takeaways

The argument at a glance1Hindi has roughly 528 million speakers across Uttar Pradesh, Bihar, Madhya Pradesh; a corpus that samples only one state w…2Budget 500–2,000 hours for a first production ASR corpus, and at least 1,000–3,000 distinct speakers to avoid speaker over…3The hardest part is not recording — it is recruitment, dialect quotas, consent and transcription consistency across cities…
  • Hindi has roughly 528 million speakers across Uttar Pradesh, Bihar, Madhya Pradesh; a corpus that samples only one state will not generalise.
  • Budget 500–2,000 hours for a first production ASR corpus, and at least 1,000–3,000 distinct speakers to avoid speaker overfitting.
  • The hardest part is not recording — it is recruitment, dialect quotas, consent and transcription consistency across cities.

Step 1 — Write the Hindi corpus specification

A specification is the contract everything else runs against. For Hindi it must state the hours or speaker count, the dialect split across Khari Boli, Awadhi, Braj, Bhojpuri-influenced Hindi, the demographic quotas, the recording condition, the transcription convention and the acceptance criteria.

Where teams get this wrong is by specifying hours without specifying speakers. Two hundred hours from eighty speakers trains a model that recognises eighty voices. The speaker count is the variable that governs generalisation, and for Hindi we recommend 1,000–3,000 distinct participants.

Step 2 — Set quotas before recruiting

Quotas are set before the first session and tracked daily. Retrofitting a quota after 60% of collection is complete usually means discarding data, because the remaining pool cannot correct the imbalance.

Quota dimensionTypical targetReason
Speakers1,000–3,000Enough distinct voices that the model learns Hindi phonology rather than a handful of speakers
Gender50 / 50Pitch and formant range differ; unbalanced cohorts bias recognition
Age18–25: 30%, 26–40: 40%, 41–60: 30%Older speakers keep conservative Hindi forms younger urban speakers have dropped
RegionUttar Pradesh, Bihar, Madhya PradeshCovers 7 recognised dialect varieties
ConditionStudio / quiet room / field / telephonyMatch the acoustic profile of your deployment
Annotators writing prompts and responses for LLM training data — collection guides context for How do you collect Hindi speech data for ASR training
Annotators writing prompts and responses for LLM training data

Step 3 — Design the Hindi prompt script

Scripted prompts must be phonetically balanced for Hindi, covering Four-way stop contrast (voiced/voiceless x aspirated/unaspirated) that collapses in models trained on English-first acoustic units, Retroflex series ट ठ ड ढ ण routinely mis-mapped to alveolar /t/ /d/ by imported lexicons, Nasalisation (chandrabindu) is phonemic and is dropped by most off-the-shelf G2P front-ends in sufficient density. Generic translated English scripts produce corpora that miss exactly the contrasts an ASR model struggles with.

Alongside scripted material, collect spontaneous speech. Public Hindi corpora skew heavily towards read newspaper text from educated urban speakers in Delhi and NCR. Rural Bihar and eastern UP speech, elderly speakers, and low-literacy speakers reading prompts aloud are largely absent. Spontaneous data is where the disfluencies, hesitations and natural prosody live, and models trained only on read speech degrade sharply on real users.

Step 4 — Recording protocol and capture chain

  • 48 kHz / 24-bit studio capture where the deployment is app or device audio; 8 kHz narrowband captured over a real telephony path where the deployment is a contact centre
  • Documented microphone and interface chain per studio so files from different cities are interchangeable
  • Automated checks for clipping, DC offset, noise floor and silence ratio on ingest
  • Speaker metadata recorded at session time — dialect, district, age band, gender, education, device
  • Written consent in the speaker's own language, retained for audit and covering commercial model training

Step 5 — Transcription and annotation in Devanagari

Transcription is where Hindi corpora most often fail acceptance. Inconsistent Devanagari vs romanised spelling for the same English loan word Nukta characters (क़ ख़ ग़ ज़ फ़) applied inconsistently by transcribers Numerals: whether to write digits, Devanagari numerals, or spelled-out words must be fixed in the style guide up front Honorific verb forms create long agreement chains that annotators shorten unless the guide forbids it

Urban Hindi speech is Hinglish in practice. Expect 15-40% English tokens in spontaneous speech: numbers, brands, technology terms, and whole clause switches. Any Hindi corpus that excludes English tokens will not match production traffic. Decide the convention in advance — native script throughout, Roman for embedded English, or a tagged hybrid — and publish it as a style guide with worked examples. Two-pass QA by a second native reviewer measures against that guide rather than against personal preference.

Step 6 — Acceptance and delivery

Acceptance should be measurable: transcript accuracy sampled per batch, metadata completeness at 100%, audio validation pass rate, and quota adherence within an agreed tolerance. Deliver in your ingest format — WAV plus JSON or TSV manifests, with speaker IDs preserved and a consent register attached.

Roll delivery in batches rather than one final handover. Batch delivery lets your team catch a format mismatch in week two instead of week ten, and lets training start before collection ends.

Recruitment reality in Hindi-speaking regions

Largest recruitment pool in the network. A 1,000-speaker Hindi cohort with balanced gender and 18-45 age bands is typically fielded across four cities to avoid a single-city accent bias.

Our studio network covers Delhi, Lucknow, Jaipur, which is what makes dialect quotas achievable without contracting a separate vendor per state.

Frequently asked questions

How many hours of Hindi speech data do I need for a usable ASR model?

500–2,000 hours is the usual first production corpus for Hindi, on top of any pretrained multilingual base. Fine-tuning an existing multilingual model can show measurable gains from 50–100 hours if the data matches your deployment acoustics.

How many speakers should a Hindi dataset have?

1,000–3,000 distinct native speakers. Speaker diversity matters more than raw hours once you are past the first hundred hours.

Which Hindi dialects should be covered?

At minimum Khari Boli, Awadhi, Braj. Which ones dominate your quota depends on where your users are, not on which dialect is considered standard.

How is code-mixing handled in Hindi transcripts?

Urban Hindi speech is Hinglish in practice. Expect 15-40% English tokens in spontaneous speech: numbers, brands, technology terms, and whole clause switches. Any Hindi corpus that excludes English tokens will not match production traffic. We fix the convention in the style guide before collection and QA against it, because inconsistent code-mix handling is a common cause of silent WER inflation.

How long does a Hindi collection take?

A 100–300 hour Hindi programme typically runs 3–6 weeks from signed scope to final delivery, with rolling batches from week two.

Related reading

Turn this into a dataset specification

Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.

Request a dataset quote