অসমীয়া · Assamese (Eastern Nagari) · as-IN
Assamese AI Training Data Collection
We collect Assamese speech and language data for AI training across Assam, Arunachal Pradesh, parts of Nagaland and Meghalaya and beyond, covering 4 dialect varieties rather than a single prestige standard.

- Speakers
- 15M
- Script
- Assamese (Eastern Nagari)
- Studio cities
- 5
- Typical programme
- 100-500 hours
Why Assamese breaks generic speech models
Assamese is not a variant of a language your model already handles. The specific properties below are the ones that show up as errors in production.
- Assamese has the voiceless velar fricative /x/, unique among major Indian languages and routinely mis-modelled
- No retroflex-dental contrast in the way Hindi has it, so Hindi-derived phone sets over-generate
- ৰ and ৱ characters are specific to Assamese and are frequently substituted with Bengali equivalents in tooling
Dialects we cover
Assamese has 4 varieties that matter for data collection. Collecting only the standard variety produces a model that works in the capital and fails everywhere else.
- Kamrupi
- Goalparia
- Upper Assam (Sibsagar standard)
- Barak Valley contact varieties

Code-mixing reality
Assamese speech mixes Hindi, English and Bengali, with substantial contact influence in Barak Valley and tea-garden communities.
Specify your expected English-token ratio in the brief. It is cheaper to collect the right mix than to filter the wrong one afterwards.
What is missing from public data
Extremely low-resource. Almost no spontaneous Assamese speech data exists publicly, and non-standard dialects have none.
This is the practical reason to commission collection rather than assemble open corpora: the gap in the public data is exactly the part your users occupy.
Transcription rules that decide whether the data is usable
These are the Assamese-specific decisions that go into the style guide before any transcriber starts work.
- Bengali characters substituted for Assamese ৰ / ৱ
- Goalparia and Kamrupi forms standardised to Sibsagar Assamese
- Tea-garden community speech excluded because transcribers cannot handle it
Recording types available
- Spontaneous conversation across Upper and Lower Assam
- Read prompts covering /x/ and Assamese-specific graphemes
- Government-service and banking domain utterances
Recommended cohort structure
Expect longer fielding times and higher per-hour cost than for Hindi or Marathi; the speaker pool with transcription-grade literacy is smaller.
| Dimension | Typical split | Why it matters for Assamese |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Assamese forms that younger urban speakers have lost |
| Region | Assam / Arunachal Pradesh / parts of Nagaland and Meghalaya and others | Dialect spread across 4 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Where we record
Assamese sessions run in Guwahati, Jorhat, Dibrugarh, Silchar, Tezpur. Multi-city fielding is the default because single-city cohorts carry an audible accent signature.
The numbers we hold ourselves to
- 100% of delivered files pass automated technical QA for SNR, clipping, duration and silence
- 5-25% of files pass a second native-speaker content review, stratified by city, dialect and transcriber, and escalating to 100% on any batch that fails the agreed threshold
- Accepted yield runs 85-90% for scripted speech, 60-70% for spontaneous, 55-65% for conversational and 50-60% for telephony
- Default cohort quotas: 50/50 gender, with age bands at 30% (18-25), 40% (26-40) and 30% (41-60)
- 48 kHz / 24-bit capture, delivered as 16-bit PCM WAV, with studio sessions held below a -50 dBFS noise floor
- First response within one working day; a scoped, fixed quote within two to three
These are the figures a delivery is measured against, not aspirations. A batch that misses them is re-recorded at our cost rather than repaired.
Frequently asked
How many Assamese speakers can you field?
300-800 speakers for a standard programme, fielded across 5 cities. Expect longer fielding times and higher per-hour cost than for Hindi or Marathi; the speaker pool with transcription-grade literacy is smaller.
Do you transcribe Assamese in Assamese (Eastern Nagari)?
Yes, and we can deliver romanised or dual-script transcripts alongside. Bengali characters substituted for Assamese ৰ / ৱ is one of the rules fixed in the style guide before work begins.
Can you collect dialect-specific Assamese data?
Yes. Every speaker is tagged with region and dialect, so you can evaluate per variety after delivery instead of discovering the imbalance in production.
What about Assamese-English code-mixing?
Assamese speech mixes Hindi, English and Bengali, with substantial contact influence in Barak Valley and tea-garden communities.
Request a Assamese dataset quote
Tell us the hours, speaker count and dialect spread you need. You get a scope, a timeline and a price.