Dataset specification · Pilot scale
50 hours of Indian English Telephony Speech
A pilot scale build of 50 hours of Indian English telephony speech. Proving that the specification survives contact with real speakers before anyone commits a budget to it. Speech captured over a real telephony path — a genuine call through the network, not studio audio downsampled afterwards to imitate one.

- Volume
- 50 hours
- Scale
- Pilot scale
- Per speaker
- 10–20 minutes of accepted audio per speaker
- Accepted yield
- 50–60% of recorded time is accepted
How Indian English telephony speech is captured
Speech captured over a real telephony path — a genuine call through the network, not studio audio downsampled afterwards to imitate one.
Scenario-driven calls between a caller and an agent or IVR flow, each leg recorded separately, across a deliberate spread of handsets and network conditions.
Yield at this style: 50–60% of recorded time is accepted. Dropped calls, network artefacts beyond tolerance and unusable legs are discarded. Real network conditions are the point of the style and also its main cost.
The specification
| Field | Value |
|---|---|
| Language | Indian English (en-IN, Latin) |
| Volume | 50 hours |
| Equivalent | 100 speakers at 30 minutes each, or 50 speakers at one hour each |
| Speech type | Telephony Speech |
| Per speaker | 10–20 minutes of accepted audio per speaker |
| Sample rate | 8 kHz narrowband, matching what a deployed contact-centre model actually receives |
| Codec | G.711 and AMR-NB captured explicitly, with the codec recorded per call in the manifest |
| Legs | Caller and agent recorded on separate legs, never as a mixed call recording |
| Network conditions | Handset type, network carrier and packet-loss events logged per call |
| Dialects | North Indian (Hindi-substrate), Maharashtrian, South Indian (Tamil/Telugu/Kannada/Malayalam substrate), Bengali-substrate |
| Transcription | Verbatim, native-speaker, second-pass reviewed |

Designing the call flows for Indian English
- Call scenarios drawn from real contact-centre intents: balance enquiry, complaint, booking change, escalation
- IVR flows scripted with deliberate mis-entry and barge-in paths, since those are where deployed systems fail
- Handset spread specified up front — low-end Android, feature phone, landline — because handset variance is a real acoustic axis
- Background conditions varied on purpose: street, vehicle, indoor, since real callers are rarely in quiet rooms
- Built against Indian English specifically: Retroflex realisation of /t/ and /d/
- Code-mixing handled explicitly rather than edited out — Indian English embeds Hindi and regional discourse markers, kinship terms, and food and place vocabulary that Western English lexicons lack.
Running a pilot scale Indian English build
Three to four weeks from signed specification. Recruitment is the critical path, not recording — the studio time itself is under two weeks.
Two batches: a 10-hour first batch in week two so you can run it through your pipeline while the rest is still recording, then the balance on completion.
Balance by substrate language, not by city alone, and tag each speaker so accent-band evaluation is possible after delivery.
| Parameter | At this volume |
|---|---|
| Cities | One, usually the densest pool for the language — Bengaluru, Delhi |
| Studios | A single treated room |
| Recruiters | One coordinator |
| Speakers | ~100–120 |
| Sessions per day | 6–8 |
| Team | 1 coordinator, 2 engineers, 3 transcribers |
Cohort design
At 50 hours the cohort is deliberately simplified: two or three dialect groups rather than the full spread, with quotas enforced in aggregate. A build this size cannot support per-cell targets and should not claim to.
Recruitment must be stratified by handset and carrier as well as by dialect, which adds a screening axis the other styles do not have.
| Dimension | Typical split | Why it matters for Indian English |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Indian English forms that younger urban speakers have lost |
| Region | Pan-India, with distinct regional accent bands and others | Dialect spread across 5 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Indian English-specific considerations
- Retroflex realisation of /t/ and /d/
- Indian English embeds Hindi and regional discourse markers, kinship terms, and food and place vocabulary that Western English lexicons lack.
- Commercial English ASR is trained overwhelmingly on US and UK speech. Indian English accent data with substrate-language tagging is the fastest way to close the accuracy gap for Indian deployments.
Quality gates for telephony speech
- Codec and sample rate verified per file, rejecting any studio audio that has been downsampled to fake a telephony path
- Echo and double-talk checked on both legs
- DTMF events verified against the call log
- Level normalisation applied per leg, since handset output levels vary far more than studio microphones do
- Indian-specific vocabulary flagged as errors by spellcheck-driven QA
- Numbers spoken in lakhs and crores mis-normalised into millions
100% content QA. At this volume there is no reason to sample, and a pilot exists precisely to surface specification problems.
What goes wrong on telephony speech sessions
- Studio audio downsampled to 8 kHz and passed off as telephony, which has none of the codec or packet-loss characteristics that matter
- Echo and double-talk that make the agent leg unusable
- Carrier and handset monoculture, producing a corpus that only represents one acoustic path
- Over-clean recordings from participants who move somewhere quiet to take the call, defeating the purpose
Risks at pilot scale in Indian English
Cost at this band is driven by: Fixed setup cost is amortised over very few hours, so the per-hour rate is at its highest here; Narrow demographic requirements bite hardest at small cohort sizes.
- The specification itself is usually the risk, not delivery — pilots exist to find the clauses that do not survive contact with real speakers
- A single-city cohort will not represent the language nationally, and reading pilot results as national is the most common mistake at this band
- Indian English carries 5 recognised varieties across Pan-India, with distinct regional accent bands, so the quota matrix is wider than the headline volume suggests
- Commercial English ASR is trained overwhelmingly on US and UK speech. Indian English accent data with substrate-language tagging is the fastest way to close the accuracy gap for Indian deployments.
Deliverables
- WAV audio to your naming convention, with the per-file manifest
- Verbatim Latin transcripts with utterance-level timestamps
- Call metadata: duration, codec, handset class, carrier, packet-loss events
- Intent label per call and per turn
- DTMF and hold, transfer and barge-in events
- Agent and caller leg identifiers
- Per-speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates, rejection reasons and agreement statistics
- Speaker-disjoint train / dev / test splits on request
What this trains, and what it does not
- Contact-centre and IVR ASR
- Voice bots operating over the phone network
- Intent classification on narrowband audio
- Robustness to codec and packet loss
Narrowband telephony audio is the wrong input for TTS or any wideband model — the frequency content simply is not there. Use it for models that will be deployed on a phone line and nothing else.
Frequently asked
Is 50 hours of Indian English enough?
Enough to validate a specification, benchmark a vendor and produce a small evaluation set. Not enough to move a production model's error rate.
Why telephony speech rather than another speech type?
Contact-centre and IVR ASR, Voice bots operating over the phone network, Intent classification on narrowband audio are what this style is the right input for. Narrowband telephony audio is the wrong input for TTS or any wideband model — the frequency content simply is not there. Use it for models that will be deployed on a phone line and nothing else.
How long does a pilot scale Indian English build take?
Three to four weeks from signed specification. Recruitment is the critical path, not recording — the studio time itself is under two weeks. Two batches: a 10-hour first batch in week two so you can run it through your pipeline while the rest is still recording, then the balance on completion.
What does 50 hours of Indian English telephony speech cost?
Quoted per delivered hour against this specification. At this band the drivers are fixed setup cost is amortised over very few hours, so the per-hour rate is at its highest here and narrow demographic requirements bite hardest at small cohort sizes. Send the spec and you get one fixed figure.
How much QA is applied at this volume?
100% content QA. At this volume there is no reason to sample, and a pilot exists precisely to surface specification problems.
Quote this Indian English dataset
50 hours, telephony speech, Indian English — pilot scale. Adjust anything and send it.