Dataset specification · Evaluation scale
100 hours of Kannada Conversational Speech
An evaluation scale build of 100 hours of Kannada conversational speech. Building a benchmark or evaluation set that is large enough to trust, or fine-tuning a narrow domain. Two speakers hold an unscripted conversation seeded with a scenario, recorded on separate channels so overlap and turn-taking survive into the delivered files.

- Volume
- 100 hours
- Scale
- Evaluation scale
- Per speaker
- 20–30 minutes of accepted audio per speaker, in pairs
- Accepted yield
- 55–65% of recorded time is accepted
How Kannada conversational speech is captured
Two speakers hold an unscripted conversation seeded with a scenario, recorded on separate channels so overlap and turn-taking survive into the delivered files.
A 45–60 minute paired session with both speakers on isolated microphones, either in adjacent treated rooms or split-mic in one room with bleed measured and logged.
Yield at this style: 55–65% of recorded time is accepted. Overlap regions, crosstalk bleed and one-sided stretches all cost delivered time. The lowest-yield studio style we run, and priced accordingly.
The specification
| Field | Value |
|---|---|
| Language | Kannada (kn-IN, Kannada) |
| Volume | 100 hours |
| Equivalent | 200 speakers at 30 minutes each, or 100 speakers at one hour each |
| Speech type | Conversational Speech |
| Per speaker | 20–30 minutes of accepted audio per speaker, in pairs |
| Channels | Two, one per speaker, never mixed down before delivery |
| Channel isolation | Bleed measured per session and logged; sessions over threshold are re-recorded |
| Overlap | Preserved, timestamped and labelled rather than edited out |
| File granularity | Per-channel session WAV plus a turn-level manifest with speaker IDs |
| Dialects | Bangalore urban, Mysuru (standard literary), Dharwad / North Karnataka, Mangaluru coastal |
| Transcription | Verbatim, native-speaker, second-pass reviewed |

Designing the scenarios for Kannada
- Scenario seeds rather than scripts — a disagreement to resolve, a plan to make, an experience to compare
- Pairing designed deliberately: familiar pairs produce natural interruption, stranger pairs produce polite turn-taking, and you need both
- Scenarios that invite disagreement, because agreeable conversation produces almost no overlap to train on
- Register mixed across pairs so the corpus is not uniformly formal
- Built against Kannada specifically: North Karnataka speech has markedly different intonation and lexicon from Mysuru standard
- Code-mixing handled explicitly rather than edited out — Bengaluru is a migration city: Kannada speech there is mixed with English, Hindi, Tamil and Telugu. Native-only Kannada cohorts recruited in Bengaluru are hard to fill without screening for years of residence.
Running an evaluation scale Kannada build
Four to six weeks. Two cities running in parallel means recruitment and recording overlap rather than queue behind each other.
Three batches at roughly two-week intervals, each one a self-contained, speaker-disjoint slice you can evaluate independently.
Screen Bengaluru participants for native fluency and years of Karnataka residence; otherwise the cohort drifts towards second-language Kannada.
| Parameter | At this volume |
|---|---|
| Cities | Two, chosen for dialect contrast — Bengaluru, Mysuru |
| Studios | Two rooms running in parallel |
| Recruiters | Two coordinators |
| Speakers | ~200–240 |
| Sessions per day | 12–16 across both cities |
| Team | 2 coordinators, 4 engineers, 6 transcribers, 1 QA lead |
Cohort design
At 100 hours the cohort is deliberately simplified: two or three dialect groups rather than the full spread, with quotas enforced in aggregate. A build this size cannot support per-cell targets and should not claim to.
Effectively doubles recruitment load, since speakers are booked in matched pairs and a single drop-out cancels the whole session. Plan on 20–25% over-recruitment.
| Dimension | Typical split | Why it matters for Kannada |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Kannada forms that younger urban speakers have lost |
| Region | Karnataka / parts of Maharashtra, Tamil Nadu and Andhra Pradesh and others | Dialect spread across 5 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Kannada-specific considerations
- North Karnataka speech has markedly different intonation and lexicon from Mysuru standard
- Bengaluru is a migration city: Kannada speech there is mixed with English, Hindi, Tamil and Telugu. Native-only Kannada cohorts recruited in Bengaluru are hard to fill without screening for years of residence.
- Mysuru/Bengaluru standard dominates. North Karnataka (Dharwad, Kalaburagi) and coastal Mangaluru speech are barely represented in any public corpus.
Quality gates for conversational speech
- Channel bleed measured on every session and rejected above the agreed threshold
- Diarisation labels verified against the isolated channels rather than inferred from the mix
- Turn boundaries and overlap regions checked by a native listener
- Speaking-time balance per pair audited, so a dominant speaker does not silently halve the session's value
- Northern lexical items replaced with standard equivalents
- Inconsistent transliteration of English technical terms
100% technical QA, 25% content QA, escalating to full review on any batch that fails the agreed error threshold.
What goes wrong on conversational speech sessions
- Crosstalk bleed that makes clean per-speaker training targets impossible to recover afterwards
- One speaker dominating, leaving a pair that delivers half the expected audio
- Unnatural politeness between strangers, producing clean but unrepresentative turn-taking
- Scheduling attrition — both speakers have to show up, so no-show rates compound rather than add
Risks at evaluation scale in Kannada
Cost at this band is driven by: Dialect spread is the main driver at this band — two contrasting cities cost more than two convenient ones; Turnaround compression, if you need it inside four weeks.
- Two-city cohorts can hide a dialect gap that only appears when the model meets a third region
- Transcription consistency between two city teams needs an explicit convention document or the batches will not match
- Kannada carries 5 recognised varieties across Karnataka, parts of Maharashtra, Tamil Nadu and Andhra Pradesh, so the quota matrix is wider than the headline volume suggests
- Mysuru/Bengaluru standard dominates. North Karnataka (Dharwad, Kalaburagi) and coastal Mangaluru speech are barely represented in any public corpus.
Deliverables
- WAV audio to your naming convention, with the per-file manifest
- Verbatim Kannada transcripts with utterance-level timestamps
- Speaker-turn segmentation with start and end timestamps
- Overlap regions marked with participating speaker IDs
- Backchannel and interruption markers
- Per-pair relationship metadata: familiar or stranger
- Per-speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates, rejection reasons and agreement statistics
- Speaker-disjoint train / dev / test splits on request
What this trains, and what it does not
- Speaker diarisation
- Meeting and multi-party ASR
- Turn-taking and endpointing for voice agents
- Speaker separation and target-speaker extraction
Two-party conversation does not generalise to multi-party meetings with four or more speakers, where overlap statistics change substantially. Specify that case separately.
Frequently asked
Is 100 hours of Kannada enough?
A solid evaluation set, and enough to fine-tune an existing multilingual model on a narrow domain. Still short of what a general production model needs.
Why conversational speech rather than another speech type?
Speaker diarisation, Meeting and multi-party ASR, Turn-taking and endpointing for voice agents are what this style is the right input for. Two-party conversation does not generalise to multi-party meetings with four or more speakers, where overlap statistics change substantially. Specify that case separately.
How long does an evaluation scale Kannada build take?
Four to six weeks. Two cities running in parallel means recruitment and recording overlap rather than queue behind each other. Three batches at roughly two-week intervals, each one a self-contained, speaker-disjoint slice you can evaluate independently.
What does 100 hours of Kannada conversational speech cost?
Quoted per delivered hour against this specification. At this band the drivers are dialect spread is the main driver at this band — two contrasting cities cost more than two convenient ones and turnaround compression, if you need it inside four weeks. Send the spec and you get one fixed figure.
How much QA is applied at this volume?
100% technical QA, 25% content QA, escalating to full review on any batch that fails the agreed error threshold.
Quote this Kannada dataset
100 hours, conversational speech, Kannada — evaluation scale. Adjust anything and send it.