Sample · मराठी
Free Marathi speech data sample
Thirty minutes of Marathi audio, recorded to your specification with transcripts and full metadata, at no cost. It exists so you can judge us on files rather than on claims.

- Sample size
- ~30 minutes
- Cost
- Free
- Turnaround
- 5-8 working days
- Formats
- WAV 48 kHz + transcript + metadata
What the sample contains
- Marathi audio from at least four speakers with a mixed gender and age spread
- Coverage of Standard (Puneri), Varhadi (Vidarbha), Marathwadi where your spec calls for dialect breadth
- Verbatim transcripts in Devanagari, with a romanised parallel transcript on request
- Per-file metadata: speaker ID, age band, gender, district, device, environment and SNR
- A short QA report showing how the files were checked before they were sent
How to specify it
| Decision | Options | Default if you do not say |
|---|---|---|
| Speech type | Scripted / spontaneous / conversational / telephony | Scripted plus a spontaneous block |
| Environment | Studio / quiet room / field / call path | Studio and quiet room, split |
| Sample rate | 48 kHz, 16 kHz, 8 kHz | 48 kHz mastered, 16 kHz derived |
| Dialects | Standard (Puneri), Varhadi (Vidarbha), Marathwadi | Standard variety plus one regional |
| Annotation | Verbatim / timestamps / speaker turns / event tags | Verbatim with timestamps |
The sample is only useful if it looks like the data you would actually buy. Send the spec you would put in a purchase order and we record against that.

What to test it against
- Run it through your existing ASR or TTS pipeline and read the error breakdown by speaker and dialect
- Check how Marathi code-mixing is transcribed: Mumbai and Pune speech mixes Marathi, Hindi, and English in the same sentence. Marathi-only recordings collected in Pune under-represent the Mumbai reality of tri-lingual switching.
- Check the transcription convention against your tokeniser — ळ vs ल substitution by Hindi-trained transcribers is the usual failure point
- Confirm the metadata is complete enough to filter and stratify by
What a Marathi sample usually exposes
Marathi has roughly 99 million speakers across Maharashtra, Goa, parts of Karnataka, and models trained on generic multilingual data typically fail on the same things each time.
Available Marathi speech data is dominated by standard Puneri read speech. Vidarbha, Marathwada, and coastal Konkan varieties are severely under-collected, which is exactly where deployed voice products lose accuracy.
- Phonetics your acoustic model may not have seen: Retains the retroflex lateral ळ, which has no Hindi or English equivalent and is frequently substituted with ल by non-native transcribers
- Utterance types worth sampling separately: Phonetically balanced read prompts covering ळ, ण and retroflex clusters, Agricultural, banking, and government-scheme domain utterances, Spontaneous two-party conversation with natural Hindi/English switching
- Recording geography in the sample: Mumbai, Pune, Nagpur
What happens after
Nothing automatic. If the sample works, we scope the full build against the same specification and quote a fixed price. If it does not, tell us what failed — that feedback is more useful to us than a polite no.
A representative Marathi cohort should be split roughly 40% western Maharashtra, 25% Vidarbha, 20% Marathwada, 15% Konkan rather than concentrated in Pune.
Frequently asked
Is the Marathi sample really free?
Yes. There is no charge and no obligation. We cap it at roughly 30 minutes because that is enough to judge quality without becoming an unpaid production run.
Can we use the sample data commercially?
The sample is licensed for evaluation. Full commercial rights and IP transfer come with a paid delivery, including a paid pilot.
How long does it take?
Five to eight working days from an agreed specification, depending on how narrow the speaker cohort is.
Can we specify our own script?
Yes. Send your prompts, wake words or domain vocabulary and we record those instead of our standard set.
Related pages
Request your free Marathi sample
Send the specification you would buy against. Thirty minutes of audio, transcripts and metadata come back within a week.