বাংলা · Bengali · bn-IN
Bengali AI Training Data Collection
We collect Bengali speech and language data for AI training across West Bengal, Tripura, Assam (Barak Valley) and beyond, covering 5 dialect varieties rather than a single prestige standard.

- Speakers
- 97M
- Script
- Bengali
- Studio cities
- 5
- Typical programme
- 500-2,000 hours
Why Bengali breaks generic speech models
Bengali is not a variant of a language your model already handles. The specific properties below are the ones that show up as errors in production.
- Inherent vowel is realised as /ɔ/ or /o/, which breaks G2P rules copied from Devanagari-based systems
- No phonemic distinction between শ ষ স in most speech despite three orthographic characters
- Consonant clusters simplify in colloquial speech in ways read prompts never capture
Dialects we cover
Bengali has 5 varieties that matter for data collection. Collecting only the standard variety produces a model that works in the capital and fails everywhere else.
- Kolkata standard (Rarhi)
- Sylheti-influenced
- Rangpuri / North Bengal
- Medinipuri
- Bangal varieties

Code-mixing reality
Kolkata professional speech mixes English heavily; rural West Bengal much less. A single 'Bengali' dataset without register tags conflates two very different acoustic and lexical distributions.
Specify your expected English-token ratio in the brief. It is cheaper to collect the right mix than to filter the wrong one afterwards.
What is missing from public data
Indian Bengali is under-collected relative to Bangladeshi Bengali, and North Bengal and Tripura varieties are almost entirely missing.
This is the practical reason to commission collection rather than assemble open corpora: the gap in the public data is exactly the part your users occupy.
Transcription rules that decide whether the data is usable
These are the Bengali-specific decisions that go into the style guide before any transcriber starts work.
- Three sibilant characters chosen inconsistently for the same sound
- Bangladeshi vs Indian Bengali orthographic conventions mixed within one dataset
- Verb conjugation register (cholit vs sadhu) normalised by transcribers
Recording types available
- Cholit-bhasha spontaneous conversation
- Read prompts covering cluster simplification
- Regional dialect sets from North Bengal and Tripura
Recommended cohort structure
Tag every speaker as Indian Bengali and record district of origin; mixing in Bangladeshi speech without tags is a common and costly dataset defect.
| Dimension | Typical split | Why it matters for Bengali |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Bengali forms that younger urban speakers have lost |
| Region | West Bengal / Tripura / Assam (Barak Valley) and others | Dialect spread across 5 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Where we record
Bengali sessions run in Kolkata, Siliguri, Durgapur, Agartala, Asansol. Multi-city fielding is the default because single-city cohorts carry an audible accent signature.
The numbers we hold ourselves to
- 100% of delivered files pass automated technical QA for SNR, clipping, duration and silence
- 5-25% of files pass a second native-speaker content review, stratified by city, dialect and transcriber, and escalating to 100% on any batch that fails the agreed threshold
- Accepted yield runs 85-90% for scripted speech, 60-70% for spontaneous, 55-65% for conversational and 50-60% for telephony
- Default cohort quotas: 50/50 gender, with age bands at 30% (18-25), 40% (26-40) and 30% (41-60)
- 48 kHz / 24-bit capture, delivered as 16-bit PCM WAV, with studio sessions held below a -50 dBFS noise floor
- First response within one working day; a scoped, fixed quote within two to three
These are the figures a delivery is measured against, not aspirations. A batch that misses them is re-recorded at our cost rather than repaired.
Frequently asked
How many Bengali speakers can you field?
1,000-3,000 speakers for a standard programme, fielded across 5 cities. Tag every speaker as Indian Bengali and record district of origin; mixing in Bangladeshi speech without tags is a common and costly dataset defect.
Do you transcribe Bengali in Bengali?
Yes, and we can deliver romanised or dual-script transcripts alongside. Three sibilant characters chosen inconsistently for the same sound is one of the rules fixed in the style guide before work begins.
Can you collect dialect-specific Bengali data?
Yes. Every speaker is tagged with region and dialect, so you can evaluate per variety after delivery instead of discovering the imbalance in production.
What about Bengali-English code-mixing?
Kolkata professional speech mixes English heavily; rural West Bengal much less. A single 'Bengali' dataset without register tags conflates two very different acoustic and lexical distributions.
Request a Bengali dataset quote
Tell us the hours, speaker count and dialect spread you need. You get a scope, a timeline and a price.