aidataservices.inAI data collection · India

অসমীয়া · Assamese (Eastern Nagari) · as-IN

Assamese AI Training Data Collection

We collect Assamese speech and language data for AI training across Assam, Arunachal Pradesh, parts of Nagaland and Meghalaya and beyond, covering 4 dialect varieties rather than a single prestige standard.

Request a dataset quoteReply within one working day
Diverse Indian speakers waiting for multilingual data collection sessions — Assamese AI Training Data Collection
Speakers
15M
Script
Assamese (Eastern Nagari)
Studio cities
5
Typical programme
100-500 hours
01

Why Assamese breaks generic speech models

Assamese is not a variant of a language your model already handles. The specific properties below are the ones that show up as errors in production.

  • Assamese has the voiceless velar fricative /x/, unique among major Indian languages and routinely mis-modelled
  • No retroflex-dental contrast in the way Hindi has it, so Hindi-derived phone sets over-generate
  • ৰ and ৱ characters are specific to Assamese and are frequently substituted with Bengali equivalents in tooling
15M speakers · 4 recognised varietiesKamrupiGoalpariaUpper Assam (Sibsagar standar…Barak Valley contact varietiesRecruited acrossAssam, Arunachal Pradesh, parts of Nagaland and MeghalayaStudio citiesGuwahati, Jorhat, Dibrugarh, SilcharDialect follows geography. A single-city cohort carries an audible accent signaturethat a national model will not generalise past.
02

Dialects we cover

Assamese has 4 varieties that matter for data collection. Collecting only the standard variety produces a model that works in the capital and fails everywhere else.

  • Kamrupi
  • Goalparia
  • Upper Assam (Sibsagar standard)
  • Barak Valley contact varieties
Audio waveforms being prepared as ASR training data — supporting assamese ai training data collection
Audio waveforms being prepared as ASR training data
03

Code-mixing reality

Assamese speech mixes Hindi, English and Bengali, with substantial contact influence in Barak Valley and tea-garden communities.

Specify your expected English-token ratio in the brief. It is cheaper to collect the right mix than to filter the wrong one afterwards.

04

What is missing from public data

Extremely low-resource. Almost no spontaneous Assamese speech data exists publicly, and non-standard dialects have none.

This is the practical reason to commission collection rather than assemble open corpora: the gap in the public data is exactly the part your users occupy.

05

Transcription rules that decide whether the data is usable

These are the Assamese-specific decisions that go into the style guide before any transcriber starts work.

  • Bengali characters substituted for Assamese ৰ / ৱ
  • Goalparia and Kamrupi forms standardised to Sibsagar Assamese
  • Tea-garden community speech excluded because transcribers cannot handle it
06

Recording types available

  • Spontaneous conversation across Upper and Lower Assam
  • Read prompts covering /x/ and Assamese-specific graphemes
  • Government-service and banking domain utterances
07

Recommended cohort structure

Expect longer fielding times and higher per-hour cost than for Hindi or Marathi; the speaker pool with transcription-grade literacy is smaller.

DimensionTypical splitWhy it matters for Assamese
Gender50 / 50Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition
Age18-25: 30%, 26-40: 40%, 41-60: 30%Older speakers retain conservative Assamese forms that younger urban speakers have lost
RegionAssam / Arunachal Pradesh / parts of Nagaland and Meghalaya and othersDialect spread across 4 recognised varieties
EducationMixed, including below-graduatePrompt-reading fluency correlates with education and skews prosody
ConditionStudio / quiet room / fieldMatch the noise profile of your deployment
08

Where we record

Assamese sessions run in Guwahati, Jorhat, Dibrugarh, Silchar, Tezpur. Multi-city fielding is the default because single-city cohorts carry an audible accent signature.

09

The numbers we hold ourselves to

  • 100% of delivered files pass automated technical QA for SNR, clipping, duration and silence
  • 5-25% of files pass a second native-speaker content review, stratified by city, dialect and transcriber, and escalating to 100% on any batch that fails the agreed threshold
  • Accepted yield runs 85-90% for scripted speech, 60-70% for spontaneous, 55-65% for conversational and 50-60% for telephony
  • Default cohort quotas: 50/50 gender, with age bands at 30% (18-25), 40% (26-40) and 30% (41-60)
  • 48 kHz / 24-bit capture, delivered as 16-bit PCM WAV, with studio sessions held below a -50 dBFS noise floor
  • First response within one working day; a scoped, fixed quote within two to three

These are the figures a delivery is measured against, not aspirations. A batch that misses them is re-recorded at our cost rather than repaired.

Frequently asked

How many Assamese speakers can you field?

300-800 speakers for a standard programme, fielded across 5 cities. Expect longer fielding times and higher per-hour cost than for Hindi or Marathi; the speaker pool with transcription-grade literacy is smaller.

Do you transcribe Assamese in Assamese (Eastern Nagari)?

Yes, and we can deliver romanised or dual-script transcripts alongside. Bengali characters substituted for Assamese ৰ / ৱ is one of the rules fixed in the style guide before work begins.

Can you collect dialect-specific Assamese data?

Yes. Every speaker is tagged with region and dialect, so you can evaluate per variety after delivery instead of discovering the imbalance in production.

What about Assamese-English code-mixing?

Assamese speech mixes Hindi, English and Bengali, with substantial contact influence in Barak Valley and tea-garden communities.

Request a Assamese dataset quote

Tell us the hours, speaker count and dialect spread you need. You get a scope, a timeline and a price.

Request a dataset quote