aidataservices.inAI data collection · India

Catalogue · অসমীয়া

Ready-made Assamese speech datasets

Corpora already recorded in Assamese and available under licence, plus what to check before buying existing data instead of commissioning a build.

Request a dataset quoteReply within one working day
Voice artist recording training data for an AI voice model — Ready-made Assamese speech datasets
Language
Assamese (as-IN)
Licence
Perpetual, commercial, model-training permitted
Availability
Subject to consent scope; confirmed on enquiry
Lead time
Days, not weeks
01

Available corpus types

CorpusTypical bandAnnotationBest for
Assamese scripted speech100-500 hoursVerbatim + timestampsPrompt-read speech from a phonetically balanced script, the baseline corpus type for ASR and TTS training.
Assamese spontaneous speech100-500 hoursVerbatim + timestampsUnscripted monologue on prompted topics, which carries the disfluencies and prosody that scripted data never produces.
Assamese conversational speech100-500 hoursVerbatim + speaker turns + timestampsTwo-party conversation recorded on separate channels, with overlap and turn-taking preserved.
Assamese telephony speech50-300 hoursVerbatim + timestampsNarrowband call audio captured over a telephony path, matching what a deployed contact-centre model actually receives.

Availability is confirmed per enquiry. Some material can only be licensed within the consent scope the speakers originally agreed to, and we will not stretch that to close a sale.

02

Ready-made versus commissioned

Ready-madeCommissioned build
Lead timeDays3-16 weeks
Cost per hourLowerHigher
Cohort controlFixed by the original specYours
Dialect quotasAs recordedDesigned
ExclusivityNon-exclusiveExclusive if you want it
Best forBaselines, prototypes, augmentationProduction models
Studio-grade voice recording session for text-to-speech training data — supporting ready-made assamese speech datasets
Studio-grade voice recording session for text-to-speech training data
03

What to check before licensing existing data

  • Consent scope: does it permit commercial model training, and by a third party?
  • Dialect distribution: a corpus that is 90% one urban variety will not generalise across Assam, Arunachal Pradesh, parts of Nagaland and Meghalaya
  • Recording conditions: studio-only data underperforms badly on noisy deployments
  • Transcription convention: Bengali characters substituted for Assamese ৰ / ৱ will otherwise surface as tokenisation noise
  • Overlap: check whether the corpus already sits in the public sets you trained on
04

Where the public corpora fall short

Extremely low-resource. Almost no spontaneous Assamese speech data exists publicly, and non-standard dialects have none.

This is usually the reason buyers move from licensing to commissioning: the free and cheap material covers the easy half of the language.

05

What a usable Assamese corpus has to cover

  • Dialect spread across Kamrupi, Goalparia, Upper Assam (Sibsagar standard), Barak Valley contact varieties rather than a single prestige variety
  • The phonetic contrasts that Assamese models actually get wrong: Assamese has the voiceless velar fricative /x/, unique among major Indian languages and routinely mis-modelled
  • Code-mixed speech transcribed rather than discarded — Assamese speech mixes Hindi, English and Bengali, with substantial contact influence in Barak Valley and tea-garden communities.
  • Recording geography covering Guwahati, Jorhat, Dibrugarh, Silchar, since dialect follows location
  • Utterance types beyond read prompts: Spontaneous conversation across Upper and Lower Assam, Read prompts covering /x/ and Assamese-specific graphemes, Government-service and banking domain utterances

Expect longer fielding times and higher per-hour cost than for Hindi or Marathi; the speaker pool with transcription-grade literacy is smaller.

Frequently asked

Can I buy an existing Assamese dataset today?

Where a corpus exists within an appropriate consent and licence scope, yes — usually within days. We confirm availability against your intended use before quoting.

Is the licence exclusive?

Ready-made corpora are licensed non-exclusively. Exclusivity is available on commissioned builds.

Can we augment a ready-made corpus with new recordings?

That is the most common pattern: licence the base, then commission the dialects, conditions or domain vocabulary it lacks.

Do you provide a data sheet?

Yes — speaker counts, demographic distribution, condition mix, annotation convention and known limitations, before you commit.

Related pages

Check Assamese availability

Tell us the hours, corpus type and intended use. We confirm what is licensable now and what needs recording.

Request a dataset quote