Dataset specification · Pilot cohort
100 speakers of Malayalam Conversational Speech
A pilot cohort build of 100 speakers of Malayalam conversational speech. Establishing whether a cohort definition is recruitable at all before it is scaled. Two speakers hold an unscripted conversation seeded with a scenario, recorded on separate channels so overlap and turn-taking survive into the delivered files.

- Volume
- 100 speakers
- Scale
- Pilot cohort
- Per speaker
- 20–30 minutes of accepted audio per speaker, in pairs
- Accepted yield
- 55–65% of recorded time is accepted
How Malayalam conversational speech is captured
Two speakers hold an unscripted conversation seeded with a scenario, recorded on separate channels so overlap and turn-taking survive into the delivered files.
A 45–60 minute paired session with both speakers on isolated microphones, either in adjacent treated rooms or split-mic in one room with bleed measured and logged.
Yield at this style: 55–65% of recorded time is accepted. Overlap regions, crosstalk bleed and one-sided stretches all cost delivered time. The lowest-yield studio style we run, and priced accordingly.
The specification
| Field | Value |
|---|---|
| Language | Malayalam (ml-IN, Malayalam) |
| Volume | 100 speakers |
| Equivalent | 50 hours at 30 minutes per speaker |
| Speech type | Conversational Speech |
| Per speaker | 20–30 minutes of accepted audio per speaker, in pairs |
| Channels | Two, one per speaker, never mixed down before delivery |
| Channel isolation | Bleed measured per session and logged; sessions over threshold are re-recorded |
| Overlap | Preserved, timestamped and labelled rather than edited out |
| File granularity | Per-channel session WAV plus a turn-level manifest with speaker IDs |
| Dialects | Thiruvananthapuram, Kochi (central), Malabar / Kozhikode, Thrissur |
| Transcription | Verbatim, native-speaker, second-pass reviewed |

Designing the scenarios for Malayalam
- Scenario seeds rather than scripts — a disagreement to resolve, a plan to make, an experience to compare
- Pairing designed deliberately: familiar pairs produce natural interruption, stranger pairs produce polite turn-taking, and you need both
- Scenarios that invite disagreement, because agreeable conversation produces almost no overlap to train on
- Register mixed across pairs so the corpus is not uniformly formal
- Built against Malayalam specifically: One of the most consonant-dense Indian languages; long geminates and clusters raise word error rates sharply
- Code-mixing handled explicitly rather than edited out — Manglish is standard in urban and professional speech, with heavy English noun and verb insertion.
Running a pilot cohort Malayalam build
Three to four weeks, almost entirely determined by how narrow the screening criteria are.
A single delivery with the full demographic report, since the cohort report is the point of a build this size.
Budget higher transcription effort per audio hour for Malayalam than for Hindi; speech rate and morphology make it slower to annotate.
| Parameter | At this volume |
|---|---|
| Cities | One — Kochi, Thiruvananthapuram |
| Studios | A single treated room |
| Recruiters | One coordinator |
| Audio yield | ~50 hours at 30 minutes per speaker |
| Sessions per day | 6–8 |
| Team | 1 coordinator, 2 engineers, 3 transcribers |
Cohort design
At 100 speakers the cohort is deliberately simplified: two or three dialect groups rather than the full spread, with quotas enforced in aggregate. A build this size cannot support per-cell targets and should not claim to.
Effectively doubles recruitment load, since speakers are booked in matched pairs and a single drop-out cancels the whole session. Plan on 20–25% over-recruitment.
| Dimension | Typical split | Why it matters for Malayalam |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Malayalam forms that younger urban speakers have lost |
| Region | Kerala / Lakshadweep / Puducherry (Mahe) and others | Dialect spread across 5 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Malayalam-specific considerations
- One of the most consonant-dense Indian languages; long geminates and clusters raise word error rates sharply
- Manglish is standard in urban and professional speech, with heavy English noun and verb insertion.
- Central Kerala news-reading dominates public data. Malabar and southern varieties, and fast conversational speech generally, are missing.
Quality gates for conversational speech
- Channel bleed measured on every session and rejected above the agreed threshold
- Diarisation labels verified against the isolated channels rather than inferred from the mix
- Turn boundaries and overlap regions checked by a native listener
- Speaking-time balance per pair audited, so a dominant speaker does not silently halve the session's value
- Old vs new script (chillu characters, Unicode normalisation) mixed within a dataset
- Fast speech leads to dropped-word transcription errors without a second-pass QA
100% content QA, plus per-speaker demographic verification against the screening record.
What goes wrong on conversational speech sessions
- Crosstalk bleed that makes clean per-speaker training targets impossible to recover afterwards
- One speaker dominating, leaving a pair that delivers half the expected audio
- Unnatural politeness between strangers, producing clean but unrepresentative turn-taking
- Scheduling attrition — both speakers have to show up, so no-show rates compound rather than add
Risks at pilot cohort in Malayalam
Cost at this band is driven by: Screening narrowness dominates; the recording itself is a minor cost at this size.
- Narrow criteria can make a cohort unrecruitable at any price, which is exactly what this band exists to find out
- 100 speakers cannot support per-dialect evaluation — treat the results as directional
- Malayalam carries 5 recognised varieties across Kerala, Lakshadweep, Puducherry (Mahe), so the quota matrix is wider than the headline volume suggests
- Central Kerala news-reading dominates public data. Malabar and southern varieties, and fast conversational speech generally, are missing.
Deliverables
- WAV audio to your naming convention, with the per-file manifest
- Verbatim Malayalam transcripts with utterance-level timestamps
- Speaker-turn segmentation with start and end timestamps
- Overlap regions marked with participating speaker IDs
- Backchannel and interruption markers
- Per-pair relationship metadata: familiar or stranger
- Per-speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates, rejection reasons and agreement statistics
- Speaker-disjoint train / dev / test splits on request
What this trains, and what it does not
- Speaker diarisation
- Meeting and multi-party ASR
- Turn-taking and endpointing for voice agents
- Speaker separation and target-speaker extraction
Two-party conversation does not generalise to multi-party meetings with four or more speakers, where overlap statistics change substantially. Specify that case separately.
Frequently asked
Is 100 speakers of Malayalam enough?
Enough to prove a cohort definition works and to build a small speaker-verification or accent test set. Not enough for per-dialect statistics.
Why conversational speech rather than another speech type?
Speaker diarisation, Meeting and multi-party ASR, Turn-taking and endpointing for voice agents are what this style is the right input for. Two-party conversation does not generalise to multi-party meetings with four or more speakers, where overlap statistics change substantially. Specify that case separately.
How long does a pilot cohort Malayalam build take?
Three to four weeks, almost entirely determined by how narrow the screening criteria are. A single delivery with the full demographic report, since the cohort report is the point of a build this size.
What does 100 speakers of Malayalam conversational speech cost?
Quoted per delivered hour against this specification. At this band the drivers are screening narrowness dominates; the recording itself is a minor cost at this size. Send the spec and you get one fixed figure.
How much QA is applied at this volume?
100% content QA, plus per-speaker demographic verification against the screening record.
Quote this Malayalam dataset
100 speakers, conversational speech, Malayalam — pilot cohort. Adjust anything and send it.