मराठी · Devanagari · mr-IN
Marathi AI Training Data Collection
We collect Marathi speech and language data for AI training across Maharashtra, Goa, parts of Karnataka and beyond, covering 6 dialect varieties rather than a single prestige standard.

- Speakers
- 99M
- Script
- Devanagari
- Studio cities
- 6
- Typical programme
- 500-2,000 hours
Why Marathi breaks generic speech models
Marathi is not a variant of a language your model already handles. The specific properties below are the ones that show up as errors in production.
- Retains the retroflex lateral ळ, which has no Hindi or English equivalent and is frequently substituted with ल by non-native transcribers
- Affricates च and ज have both alveolar and palatal realisations depending on the word, a distinction lost in Devanagari orthography
- Pitch and vowel-length differences between Puneri and Varhadi shift the acoustic distribution enough to hurt cross-dialect ASR
Dialects we cover
Marathi has 6 varieties that matter for data collection. Collecting only the standard variety produces a model that works in the capital and fails everywhere else.
- Standard (Puneri)
- Varhadi (Vidarbha)
- Marathwadi
- Konkani-influenced coastal Marathi
- Ahirani
- Malvani

Code-mixing reality
Mumbai and Pune speech mixes Marathi, Hindi, and English in the same sentence. Marathi-only recordings collected in Pune under-represent the Mumbai reality of tri-lingual switching.
Specify your expected English-token ratio in the brief. It is cheaper to collect the right mix than to filter the wrong one afterwards.
What is missing from public data
Available Marathi speech data is dominated by standard Puneri read speech. Vidarbha, Marathwada, and coastal Konkan varieties are severely under-collected, which is exactly where deployed voice products lose accuracy.
This is the practical reason to commission collection rather than assemble open corpora: the gap in the public data is exactly the part your users occupy.
Transcription rules that decide whether the data is usable
These are the Marathi-specific decisions that go into the style guide before any transcriber starts work.
- ळ vs ल substitution by Hindi-trained transcribers
- Anusvara placement varies between conservative and modern orthography
- Varhadi verb endings get 'corrected' to standard forms unless the guide explicitly requires verbatim dialect transcription
Recording types available
- Phonetically balanced read prompts covering ळ, ण and retroflex clusters
- Agricultural, banking, and government-scheme domain utterances
- Spontaneous two-party conversation with natural Hindi/English switching
- Regional dialect elicitation sets recorded in Nagpur and Aurangabad
Recommended cohort structure
A representative Marathi cohort should be split roughly 40% western Maharashtra, 25% Vidarbha, 20% Marathwada, 15% Konkan rather than concentrated in Pune.
| Dimension | Typical split | Why it matters for Marathi |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Marathi forms that younger urban speakers have lost |
| Region | Maharashtra / Goa / parts of Karnataka and others | Dialect spread across 6 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Where we record
Marathi sessions run in Mumbai, Pune, Nagpur, Nashik, Aurangabad, Kolhapur. Multi-city fielding is the default because single-city cohorts carry an audible accent signature.
The numbers we hold ourselves to
- 100% of delivered files pass automated technical QA for SNR, clipping, duration and silence
- 5-25% of files pass a second native-speaker content review, stratified by city, dialect and transcriber, and escalating to 100% on any batch that fails the agreed threshold
- Accepted yield runs 85-90% for scripted speech, 60-70% for spontaneous, 55-65% for conversational and 50-60% for telephony
- Default cohort quotas: 50/50 gender, with age bands at 30% (18-25), 40% (26-40) and 30% (41-60)
- 48 kHz / 24-bit capture, delivered as 16-bit PCM WAV, with studio sessions held below a -50 dBFS noise floor
- First response within one working day; a scoped, fixed quote within two to three
These are the figures a delivery is measured against, not aspirations. A batch that misses them is re-recorded at our cost rather than repaired.
Frequently asked
How many Marathi speakers can you field?
1,000-3,000 speakers for a standard programme, fielded across 6 cities. A representative Marathi cohort should be split roughly 40% western Maharashtra, 25% Vidarbha, 20% Marathwada, 15% Konkan rather than concentrated in Pune.
Do you transcribe Marathi in Devanagari?
Yes, and we can deliver romanised or dual-script transcripts alongside. ळ vs ल substitution by Hindi-trained transcribers is one of the rules fixed in the style guide before work begins.
Can you collect dialect-specific Marathi data?
Yes. Every speaker is tagged with region and dialect, so you can evaluate per variety after delivery instead of discovering the imbalance in production.
What about Marathi-English code-mixing?
Mumbai and Pune speech mixes Marathi, Hindi, and English in the same sentence. Marathi-only recordings collected in Pune under-represent the Mumbai reality of tri-lingual switching.
Request a Marathi dataset quote
Tell us the hours, speaker count and dialect spread you need. You get a scope, a timeline and a price.