Dataset specification · Pilot cohort
100 speakers of Punjabi Conversational Speech
A pilot cohort build of 100 speakers of Punjabi conversational speech. Establishing whether a cohort definition is recruitable at all before it is scaled. Two speakers hold an unscripted conversation seeded with a scenario, recorded on separate channels so overlap and turn-taking survive into the delivered files.

- Volume
- 100 speakers
- Scale
- Pilot cohort
- Per speaker
- 20–30 minutes of accepted audio per speaker, in pairs
- Accepted yield
- 55–65% of recorded time is accepted
How Punjabi conversational speech is captured
Two speakers hold an unscripted conversation seeded with a scenario, recorded on separate channels so overlap and turn-taking survive into the delivered files.
A 45–60 minute paired session with both speakers on isolated microphones, either in adjacent treated rooms or split-mic in one room with bleed measured and logged.
Yield at this style: 55–65% of recorded time is accepted. Overlap regions, crosstalk bleed and one-sided stretches all cost delivered time. The lowest-yield studio style we run, and priced accordingly.
The specification
| Field | Value |
|---|---|
| Language | Punjabi (pa-IN, Gurmukhi) |
| Volume | 100 speakers |
| Equivalent | 50 hours at 30 minutes per speaker |
| Speech type | Conversational Speech |
| Per speaker | 20–30 minutes of accepted audio per speaker, in pairs |
| Channels | Two, one per speaker, never mixed down before delivery |
| Channel isolation | Bleed measured per session and logged; sessions over threshold are re-recorded |
| Overlap | Preserved, timestamped and labelled rather than edited out |
| File granularity | Per-channel session WAV plus a turn-level manifest with speaker IDs |
| Dialects | Majhi (standard), Malwai, Doabi, Puadhi |
| Transcription | Verbatim, native-speaker, second-pass reviewed |

Designing the scenarios for Punjabi
- Scenario seeds rather than scripts — a disagreement to resolve, a plan to make, an experience to compare
- Pairing designed deliberately: familiar pairs produce natural interruption, stranger pairs produce polite turn-taking, and you need both
- Scenarios that invite disagreement, because agreeable conversation produces almost no overlap to train on
- Register mixed across pairs so the corpus is not uniformly formal
- Built against Punjabi specifically: Punjabi is tonal: high, low and level tones distinguish words, and tone is not marked in Gurmukhi orthography
- Code-mixing handled explicitly rather than edited out — Punjabi speech mixes Hindi and English freely, with strong diaspora influence in urban registers.
Running a pilot cohort Punjabi build
Three to four weeks, almost entirely determined by how narrow the screening criteria are.
A single delivery with the full demographic report, since the cohort report is the point of a build this size.
Cover all three historic regions (Majha, Malwa, Doaba); tone realisation differs measurably between them.
| Parameter | At this volume |
|---|---|
| Cities | One — Amritsar, Ludhiana |
| Studios | A single treated room |
| Recruiters | One coordinator |
| Audio yield | ~50 hours at 30 minutes per speaker |
| Sessions per day | 6–8 |
| Team | 1 coordinator, 2 engineers, 3 transcribers |
Cohort design
At 100 speakers the cohort is deliberately simplified: two or three dialect groups rather than the full spread, with quotas enforced in aggregate. A build this size cannot support per-cell targets and should not claim to.
Effectively doubles recruitment load, since speakers are booked in matched pairs and a single drop-out cancels the whole session. Plan on 20–25% over-recruitment.
| Dimension | Typical split | Why it matters for Punjabi |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Punjabi forms that younger urban speakers have lost |
| Region | Punjab / Haryana / Delhi and others | Dialect spread across 5 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Punjabi-specific considerations
- Punjabi is tonal: high, low and level tones distinguish words, and tone is not marked in Gurmukhi orthography
- Punjabi speech mixes Hindi and English freely, with strong diaspora influence in urban registers.
- Tonal variation is essentially unmodelled in public Punjabi data, and Malwai/Doabi rural speech is scarce.
Quality gates for conversational speech
- Channel bleed measured on every session and rejected above the agreed threshold
- Diarisation labels verified against the isolated channels rather than inferred from the mix
- Turn boundaries and overlap regions checked by a native listener
- Speaking-time balance per pair audited, so a dominant speaker does not silently halve the session's value
- Tone is unrepresented in text, so pronunciation lexicons must be built from audio, not from spelling
- Shahmukhi vs Gurmukhi script decisions must be fixed per project
100% content QA, plus per-speaker demographic verification against the screening record.
What goes wrong on conversational speech sessions
- Crosstalk bleed that makes clean per-speaker training targets impossible to recover afterwards
- One speaker dominating, leaving a pair that delivers half the expected audio
- Unnatural politeness between strangers, producing clean but unrepresentative turn-taking
- Scheduling attrition — both speakers have to show up, so no-show rates compound rather than add
Risks at pilot cohort in Punjabi
Cost at this band is driven by: Screening narrowness dominates; the recording itself is a minor cost at this size.
- Narrow criteria can make a cohort unrecruitable at any price, which is exactly what this band exists to find out
- 100 speakers cannot support per-dialect evaluation — treat the results as directional
- Punjabi carries 5 recognised varieties across Punjab, Haryana, Delhi, so the quota matrix is wider than the headline volume suggests
- Tonal variation is essentially unmodelled in public Punjabi data, and Malwai/Doabi rural speech is scarce.
Deliverables
- WAV audio to your naming convention, with the per-file manifest
- Verbatim Gurmukhi transcripts with utterance-level timestamps
- Speaker-turn segmentation with start and end timestamps
- Overlap regions marked with participating speaker IDs
- Backchannel and interruption markers
- Per-pair relationship metadata: familiar or stranger
- Per-speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates, rejection reasons and agreement statistics
- Speaker-disjoint train / dev / test splits on request
What this trains, and what it does not
- Speaker diarisation
- Meeting and multi-party ASR
- Turn-taking and endpointing for voice agents
- Speaker separation and target-speaker extraction
Two-party conversation does not generalise to multi-party meetings with four or more speakers, where overlap statistics change substantially. Specify that case separately.
Frequently asked
Is 100 speakers of Punjabi enough?
Enough to prove a cohort definition works and to build a small speaker-verification or accent test set. Not enough for per-dialect statistics.
Why conversational speech rather than another speech type?
Speaker diarisation, Meeting and multi-party ASR, Turn-taking and endpointing for voice agents are what this style is the right input for. Two-party conversation does not generalise to multi-party meetings with four or more speakers, where overlap statistics change substantially. Specify that case separately.
How long does a pilot cohort Punjabi build take?
Three to four weeks, almost entirely determined by how narrow the screening criteria are. A single delivery with the full demographic report, since the cohort report is the point of a build this size.
What does 100 speakers of Punjabi conversational speech cost?
Quoted per delivered hour against this specification. At this band the drivers are screening narrowness dominates; the recording itself is a minor cost at this size. Send the spec and you get one fixed figure.
How much QA is applied at this volume?
100% content QA, plus per-speaker demographic verification against the screening record.
Quote this Punjabi dataset
100 speakers, conversational speech, Punjabi — pilot cohort. Adjust anything and send it.