Dataset specification · Evaluation cohort
500 speakers of Marathi Telephony Speech
An evaluation cohort build of 500 speakers of Marathi telephony speech. Speaker-count-driven work: verification, diarisation and accent robustness, where breadth matters more than hours. Speech captured over a real telephony path — a genuine call through the network, not studio audio downsampled afterwards to imitate one.

- Volume
- 500 speakers
- Scale
- Evaluation cohort
- Per speaker
- 10–20 minutes of accepted audio per speaker
- Accepted yield
- 50–60% of recorded time is accepted
How Marathi telephony speech is captured
Speech captured over a real telephony path — a genuine call through the network, not studio audio downsampled afterwards to imitate one.
Scenario-driven calls between a caller and an agent or IVR flow, each leg recorded separately, across a deliberate spread of handsets and network conditions.
Yield at this style: 50–60% of recorded time is accepted. Dropped calls, network artefacts beyond tolerance and unusable legs are discarded. Real network conditions are the point of the style and also its main cost.
The specification
| Field | Value |
|---|---|
| Language | Marathi (mr-IN, Devanagari) |
| Volume | 500 speakers |
| Equivalent | 250 hours at 30 minutes per speaker |
| Speech type | Telephony Speech |
| Per speaker | 10–20 minutes of accepted audio per speaker |
| Sample rate | 8 kHz narrowband, matching what a deployed contact-centre model actually receives |
| Codec | G.711 and AMR-NB captured explicitly, with the codec recorded per call in the manifest |
| Legs | Caller and agent recorded on separate legs, never as a mixed call recording |
| Network conditions | Handset type, network carrier and packet-loss events logged per call |
| Dialects | Standard (Puneri), Varhadi (Vidarbha), Marathwadi, Konkani-influenced coastal Marathi |
| Transcription | Verbatim, native-speaker, second-pass reviewed |

Designing the call flows for Marathi
- Call scenarios drawn from real contact-centre intents: balance enquiry, complaint, booking change, escalation
- IVR flows scripted with deliberate mis-entry and barge-in paths, since those are where deployed systems fail
- Handset spread specified up front — low-end Android, feature phone, landline — because handset variance is a real acoustic axis
- Background conditions varied on purpose: street, vehicle, indoor, since real callers are rarely in quiet rooms
- Built against Marathi specifically: Retains the retroflex lateral ळ, which has no Hindi or English equivalent and is frequently substituted with ल by non-native transcribers
- Code-mixing handled explicitly rather than edited out — Mumbai and Pune speech mixes Marathi, Hindi, and English in the same sentence. Marathi-only recordings collected in Pune under-represent the Mumbai reality of tri-lingual switching.
Running an evaluation cohort Marathi build
Six to eight weeks. Speaker-count targets front-load recruitment, so the coordinator team is proportionally larger than an equivalent hours-based build.
Four batches, each one a demographically complete slice rather than a convenient chunk, so early batches are usable for training on their own.
A representative Marathi cohort should be split roughly 40% western Maharashtra, 25% Vidarbha, 20% Marathwada, 15% Konkan rather than concentrated in Pune.
| Parameter | At this volume |
|---|---|
| Cities | Three to four — Mumbai, Pune, Nagpur, Nashik, Aurangabad |
| Studios | Four rooms |
| Recruiters | Four coordinators |
| Audio yield | ~250 hours at 30 minutes per speaker |
| Sessions per day | 25–30 |
| Team | 1 programme lead, 4 coordinators, 8 engineers, 14 transcribers, 2 QA leads |
Cohort design
At 500 speakers quotas are enforced per dialect and reconciled fortnightly. Aggregate demographics are reported per batch so drift is visible while there is still time to correct it.
Recruitment must be stratified by handset and carrier as well as by dialect, which adds a screening axis the other styles do not have.
| Dimension | Typical split | Why it matters for Marathi |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Marathi forms that younger urban speakers have lost |
| Region | Maharashtra / Goa / parts of Karnataka and others | Dialect spread across 6 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Marathi-specific considerations
- Retains the retroflex lateral ळ, which has no Hindi or English equivalent and is frequently substituted with ल by non-native transcribers
- Mumbai and Pune speech mixes Marathi, Hindi, and English in the same sentence. Marathi-only recordings collected in Pune under-represent the Mumbai reality of tri-lingual switching.
- Available Marathi speech data is dominated by standard Puneri read speech. Vidarbha, Marathwada, and coastal Konkan varieties are severely under-collected, which is exactly where deployed voice products lose accuracy.
Quality gates for telephony speech
- Codec and sample rate verified per file, rejecting any studio audio that has been downsampled to fake a telephony path
- Echo and double-talk checked on both legs
- DTMF events verified against the call log
- Level normalisation applied per leg, since handset output levels vary far more than studio microphones do
- ळ vs ल substitution by Hindi-trained transcribers
- Anusvara placement varies between conservative and modern orthography
100% technical QA, 15% content QA, plus mandatory duplicate-speaker detection across cities.
What goes wrong on telephony speech sessions
- Studio audio downsampled to 8 kHz and passed off as telephony, which has none of the codec or packet-loss characteristics that matter
- Echo and double-talk that make the agent leg unusable
- Carrier and handset monoculture, producing a corpus that only represents one acoustic path
- Over-clean recordings from participants who move somewhere quiet to take the call, defeating the purpose
Risks at evaluation cohort in Marathi
Cost at this band is driven by: Cohort breadth and screening depth, not recorded hours.
- Duplicate speakers across cities inflate the apparent cohort and quietly corrupt speaker-disjoint splits
- Recruiting for breadth tempts coordinators toward the easiest available demographic, which needs weekly quota audit
- Marathi carries 6 recognised varieties across Maharashtra, Goa, parts of Karnataka, so the quota matrix is wider than the headline volume suggests
- Available Marathi speech data is dominated by standard Puneri read speech. Vidarbha, Marathwada, and coastal Konkan varieties are severely under-collected, which is exactly where deployed voice products lose accuracy.
Deliverables
- WAV audio to your naming convention, with the per-file manifest
- Verbatim Devanagari transcripts with utterance-level timestamps
- Call metadata: duration, codec, handset class, carrier, packet-loss events
- Intent label per call and per turn
- DTMF and hold, transfer and barge-in events
- Agent and caller leg identifiers
- Per-speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates, rejection reasons and agreement statistics
- Speaker-disjoint train / dev / test splits on request
What this trains, and what it does not
- Contact-centre and IVR ASR
- Voice bots operating over the phone network
- Intent classification on narrowband audio
- Robustness to codec and packet loss
Narrowband telephony audio is the wrong input for TTS or any wideband model — the frequency content simply is not there. Use it for models that will be deployed on a phone line and nothing else.
Frequently asked
Is 500 speakers of Marathi enough?
Enough for reliable per-dialect evaluation and for training speaker-verification and diarisation systems.
Why telephony speech rather than another speech type?
Contact-centre and IVR ASR, Voice bots operating over the phone network, Intent classification on narrowband audio are what this style is the right input for. Narrowband telephony audio is the wrong input for TTS or any wideband model — the frequency content simply is not there. Use it for models that will be deployed on a phone line and nothing else.
How long does an evaluation cohort Marathi build take?
Six to eight weeks. Speaker-count targets front-load recruitment, so the coordinator team is proportionally larger than an equivalent hours-based build. Four batches, each one a demographically complete slice rather than a convenient chunk, so early batches are usable for training on their own.
What does 500 speakers of Marathi telephony speech cost?
Quoted per delivered hour against this specification. At this band the drivers are cohort breadth and screening depth, not recorded hours. Send the spec and you get one fixed figure.
How much QA is applied at this volume?
100% technical QA, 15% content QA, plus mandatory duplicate-speaker detection across cities.
Quote this Marathi dataset
500 speakers, telephony speech, Marathi — evaluation cohort. Adjust anything and send it.