Pricing · ಕನ್ನಡ
Kannada speech data: what it costs
Per-hour cost bands for Kannada speech data collection, why 5 dialect varieties change the number, and worked budgets at three common volumes.

- Language
- Kannada (kn-IN)
- Rate band
- ₹4,200 – ₹8,200 / hour
- Speakers
- 59M
- Studio cities
- 5
Why Kannada prices the way it does
Kannada has roughly 59 million speakers concentrated in Karnataka, parts of Maharashtra, Tamil Nadu and Andhra Pradesh, which sets how quickly a cohort can be recruited. Recruitment speed, not recording time, is the dominant cost in almost every programme.
Screen Bengaluru participants for native fluency and years of Karnataka residence; otherwise the cohort drifts towards second-language Kannada.
- Bengaluru is a migration city: Kannada speech there is mixed with English, Hindi, Tamil and Telugu. Native-only Kannada cohorts recruited in Bengaluru are hard to fill without screening for years of residence.
- Mysuru/Bengaluru standard dominates. North Karnataka (Dharwad, Kalaburagi) and coastal Mangaluru speech are barely represented in any public corpus.
Worked budgets
| Volume | Typical speakers | Indicative range | Timeline |
|---|---|---|---|
| 100 hours | 200 at 30 min | ₹4,20,000 – ₹8,20,000 | 3-6 weeks |
| 500 hours | 1,000 at 30 min | ₹19,32,000 – ₹37,72,000 | 6-10 weeks |
| 1,000 hours | 2,000 at 30 min | ₹35,70,000 – ₹69,70,000 | 10-16 weeks |
Unit rates fall with volume because setup, script design and recruiter onboarding are amortised, not because quality is relaxed.

Cost drivers specific to this language
- Dialect spread: covering Bangalore urban, Mysuru (standard literary), Dharwad / North Karnataka, Mangaluru coastal rather than one prestige variety adds recruitment cost but is what makes the corpus usable in production
- Script and transcription: Kannada transcription needs native reviewers, and Northern lexical items replaced with standard equivalents is the usual source of rework
- Phonetics: North Karnataka speech has markedly different intonation and lexicon from Mysuru standard, which requires reviewers trained on the language rather than generic annotators
Add-ons and their pricing
| Layer | Unit | Indicative |
|---|---|---|
| Verbatim transcription | per audio hour | ₹900 – ₹2,400 |
| Speaker diarisation and turn labels | per audio hour | ₹600 – ₹1,500 |
| Event and noise tagging | per audio hour | ₹400 – ₹1,100 |
| Romanised parallel transcript | per audio hour | ₹500 – ₹1,200 |
| Speaker-disjoint train/dev/test splits | one-off | Included |
How to get the number down without hurting the model
- Widen the age bands before you widen the dialect spread — dialect coverage is what determines production accuracy
- Use quiet-room capture where deployment audio is not studio-clean anyway
- Order transcription in a second phase once the audio passes acceptance
- Run a 10-20 hour pilot; specification errors caught there are the cheapest ones you will ever fix
Frequently asked
What is the per-hour rate for Kannada speech data?
Indicatively ₹4,200 to ₹8,200 per delivered hour for standard scripted or spontaneous capture with verbatim transcription. Narrow cohorts and studio-only capture sit at the top of that band.
Is Kannada more expensive than Hindi?
Somewhat. Smaller recruitment pools mean more effort per speaker, which shows up as a higher per-hour rate rather than a longer timeline.
Do you quote in INR or USD?
Either. Bands here are in INR; international clients are usually invoiced in USD at a fixed contract rate.
What is included in the quoted rate?
Recruitment, consent capture, recording, QA, transcription if ordered, metadata, delivery packaging and a perpetual licence with full IP transfer.
Related pages
Price a Kannada dataset
Tell us hours, speakers and dialect spread for Kannada and you get a fixed price against it.