ASR datasets · ଓଡ଼ିଆ
Odia ASR Training Data
Transcribed speech corpora built to train and evaluate automatic speech recognition, with verbatim transcription, timestamps and per-token language tagging where code-mixing occurs. This page covers how that works specifically for Odia, where retains a distinct retroflex ଳ and a full retroflex series.

- Language
- Odia (or-IN)
- Dialects covered
- 5
- Typical programme
- 100-500 hours
- Cities
- Bhubaneswar, Cuttack, Sambalpur
What changes when the language is Odia
The service specification stays constant across languages; the linguistics do not. For Odia, three things drive the design of a asr datasets programme.
- Retains a distinct retroflex ଳ and a full retroflex series
- Dialect spread: Cuttack-Bhubaneswar standard, Sambalpuri (Kosli), Ganjami, Baleswari
- Urban Odia mixes Hindi and English; western Odisha mixes Chhattisgarhi and Sambalpuri forms.
Technical specification
| Parameter | Standard |
|---|---|
| Audio | 16 kHz or 48 kHz PCM WAV |
| Transcription | Verbatim, including disfluencies, false starts and fillers |
| Timestamps | Utterance level by default; word level on request |
| Tagging | Noise, overlap, unintelligible, foreign-language and code-switch tags |
| Normalisation | Raw and normalised text columns delivered separately |
| Split | Train/dev/test splits with no speaker leakage across splits |

Odia cohort design
Western Odisha recruitment requires local field partners; remote-only recruitment yields an all-coastal cohort.
| Dimension | Typical split | Why it matters for Odia |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Odia forms that younger urban speakers have lost |
| Region | Odisha / parts of Jharkhand, West Bengal, Chhattisgarh and Andhra Pradesh and others | Dialect spread across 5 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Process
- Style guide authored per language, covering numerals, loanwords, script and disfluency rules
- Transcriber calibration round with inter-annotator agreement measurement
- First-pass transcription
- Second-pass native review
- Automated consistency checks against the style guide
- Split generation with speaker-disjoint partitions
Odia-specific quality rules
- Sambalpuri normalised into coastal Odia
- Unicode confusables between Odia and Bengali characters when transcribers reuse tooling
- Inconsistent handling of tribal-language loanwords
Agreement is measured, not assumed. We report word-level agreement on a held-out sample so you can judge label quality before training on it.
Deliverables
- Audio plus aligned transcripts (JSON/TSV, or your schema)
- Style guide as delivered documentation
- Inter-annotator agreement report
- Speaker-disjoint train/dev/test splits
Worked example
A representative Odia asr datasets engagement: 100 hours from 300 speakers, 50/50 gender, ages 18-45, spread across Bhubaneswar, Cuttack, Sambalpur, recorded to the specification above and delivered in WAV with a per-utterance manifest.
Timeline: Transcription adds roughly 1-2 weeks per 100 hours after recording, per language.
Where this data is missing today
Odia is one of the least-resourced major Indian languages. Sambalpuri and Ganjami are effectively absent from public data.
Frequently asked
How much does Odia asr datasets cost?
Priced per delivered hour or unit against a written spec. The cost drivers for Odia are dialect spread, demographic narrowness and recording condition, in that order.
Which Odia dialects are included?
By default Cuttack-Bhubaneswar standard, Sambalpuri (Kosli), Ganjami, Baleswari and others, tagged per speaker. You can also commission a single-dialect corpus if you are targeting one region.
Can you deliver Odia data in our format?
Yes. Audio plus aligned transcripts (JSON/TSV, or your schema) is the default, but naming, schema and directory structure follow your pipeline.
How long does a Odia programme take?
Transcription adds roughly 1-2 weeks per 100 hours after recording, per language.
Request a Odia asr datasets quote
Hours, speakers, dialects, deadline. Send what you have.