Code-Switching ASR · বাংলা
Bengali Data for Code-Switching ASR
Recognising speech that switches between an Indian language and English several times per sentence. In Bengali, the binding constraint is usually dialect coverage and code-mixing, not raw hours.

- Language
- Bengali
- Primary metric
- Switch-point accuracy
- Typical volume
- 500-2,000 hours
Data profile required
- Genuinely code-mixed spontaneous speech
- Per-token language ID labels
- A fixed rule for script of English tokens
What Bengali adds to the requirement
- Inherent vowel is realised as /ɔ/ or /o/, which breaks G2P rules copied from Devanagari-based systems
- No phonemic distinction between শ ষ স in most speech despite three orthographic characters
- Dialects to cover: Kolkata standard (Rarhi), Sylheti-influenced, Rangpuri / North Bengal, Medinipuri
- Kolkata professional speech mixes English heavily; rural West Bengal much less. A single 'Bengali' dataset without register tags conflates two very different acoustic and lexical distributions.

Metrics to track
- Switch-point accuracy
- Mixed-utterance WER
- Language ID token accuracy
Failure modes
- Concatenating monolingual data and calling it code-mixed
- Leaving script conventions to individual annotators
For Bengali specifically: Indian Bengali is under-collected relative to Bangladeshi Bengali, and North Bengal and Tripura varieties are almost entirely missing.
Recommended cohort
Tag every speaker as Indian Bengali and record district of origin; mixing in Bangladeshi speech without tags is a common and costly dataset defect.
| Dimension | Typical split | Why it matters for Bengali |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Bengali forms that younger urban speakers have lost |
| Region | West Bengal / Tripura / Assam (Barak Valley) and others | Dialect spread across 5 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Suggested programme shape
Start with an evaluation set of 100 speakers spread across every Bengali dialect in scope, collected before training data. Then field 500-2,000 hours of training data from disjoint speakers.
This ordering is what makes the improvement measurable rather than assumed.
Frequently asked
Is there usable public Bengali data for code-switching asr?
Indian Bengali is under-collected relative to Bangladeshi Bengali, and North Bengal and Tripura varieties are almost entirely missing.
How many Bengali speakers do we need?
1,000-3,000 speakers for a training corpus, plus a disjoint evaluation cohort covering each dialect. Speaker count matters more than hours for generalisation.
Can you run this across multiple languages at once?
Yes. Multi-language programmes run to one master specification so per-language results stay comparable.
Scope Bengali data for code-switching asr
Send the target metric and the languages in scope.