Languages
Indian language speech and language data collection
Each language is collected in its own region by native-speaker recruiters, with dialect quotas set before the first session. Below is what we cover today, with speaker base, dialect spread and typical delivery window.

- Public Hindi corpora skew heavily towards read newspaper text from educated urban speakers in Delhi and NCR. Rural Bihar and eastern UP speech, elderly speakers, and low-literacy speakers reading prompts aloud are largely absent. Dialect coverage: Khari Boli, Awadhi, Braj.528M speakers · DevanagariHindi Speech Data Collection
- Available Marathi speech data is dominated by standard Puneri read speech. Vidarbha, Marathwada, and coastal Konkan varieties are severely under-collected, which is exactly where deployed voice products lose accuracy. Dialect coverage: Standard (Puneri), Varhadi (Vidarbha), Marathwadi.99M speakers · DevanagariMarathi Speech Data Collection
- Almost all public Tamil audio is literary read speech from news or scripture. Genuine colloquial Tamil, especially southern and Kongu varieties, is the single biggest gap for anyone building Tamil voice products. Dialect coverage: Chennai (Madras Bashai), Kongu (Coimbatore), Madurai.82M speakers · TamilTamil Speech Data Collection
- Coastal Andhra read speech dominates. Telangana rural and Rayalaseema speech is thin, despite Hyderabad being the largest deployment market. Dialect coverage: Telangana, Coastal Andhra (Godavari), Rayalaseema.96M speakers · TeluguTelugu Speech Data Collection
- Mysuru/Bengaluru standard dominates. North Karnataka (Dharwad, Kalaburagi) and coastal Mangaluru speech are barely represented in any public corpus. Dialect coverage: Bangalore urban, Mysuru (standard literary), Dharwad / North Karnataka.59M speakers · KannadaKannada Speech Data Collection
- Indian Bengali is under-collected relative to Bangladeshi Bengali, and North Bengal and Tripura varieties are almost entirely missing. Dialect coverage: Kolkata standard (Rarhi), Sylheti-influenced, Rangpuri / North Bengal.97M speakers · BengaliBengali Speech Data Collection
- Very little spontaneous Gujarati audio exists publicly; nearly all of it is Ahmedabad read speech. Dialect coverage: Standard (Amdavadi), Surti, Kathiyawadi.55M speakers · GujaratiGujarati Speech Data Collection
- Central Kerala news-reading dominates public data. Malabar and southern varieties, and fast conversational speech generally, are missing. Dialect coverage: Thiruvananthapuram, Kochi (central), Malabar / Kozhikode.35M speakers · MalayalamMalayalam Speech Data Collection
- Tonal variation is essentially unmodelled in public Punjabi data, and Malwai/Doabi rural speech is scarce. Dialect coverage: Majhi (standard), Malwai, Doabi.33M speakers · GurmukhiPunjabi Speech Data Collection
- Odia is one of the least-resourced major Indian languages. Sambalpuri and Ganjami are effectively absent from public data. Dialect coverage: Cuttack-Bhubaneswar standard, Sambalpuri (Kosli), Ganjami.38M speakers · OdiaOdia Speech Data Collection
- Extremely low-resource. Almost no spontaneous Assamese speech data exists publicly, and non-standard dialects have none. Dialect coverage: Kamrupi, Goalparia, Upper Assam (Sibsagar standard).15M speakers · Assamese (Eastern Nagari)Assamese Speech Data Collection
- Indian Urdu specifically, and Dakhini in particular, are absent from public data dominated by Pakistani Urdu broadcast speech. Dialect coverage: Dakhini (Hyderabad), Lucknawi, Dehlvi.51M speakers · Perso-Arabic (Nastaliq)Urdu Speech Data Collection
- Almost no public corpus contains genuine intra-sentential Hindi-English switching with per-token language tags. This is the highest-value gap for anyone building Indian conversational AI. Dialect coverage: Delhi corporate Hinglish, Mumbai Bambaiya, Call-centre register.350M speakers · Devanagari + LatinHinglish Speech Data Collection
- Commercial English ASR is trained overwhelmingly on US and UK speech. Indian English accent data with substrate-language tagging is the fastest way to close the accuracy gap for Indian deployments. Dialect coverage: North Indian (Hindi-substrate), Maharashtrian, South Indian (Tamil/Telugu/Kannada/Malayalam substrate).130M speakers · LatinIndian English Speech Data Collection
Need a language not listed?
We regularly stand up collection for lower-resource Indian languages and dialects. Tell us the language and the volume.