Dataset specification · Pilot scale
50 hours of Assamese Telephony Speech
A pilot scale build of 50 hours of Assamese telephony speech. Proving that the specification survives contact with real speakers before anyone commits a budget to it. Speech captured over a real telephony path — a genuine call through the network, not studio audio downsampled afterwards to imitate one.

- Volume
- 50 hours
- Scale
- Pilot scale
- Per speaker
- 10–20 minutes of accepted audio per speaker
- Accepted yield
- 50–60% of recorded time is accepted
How Assamese telephony speech is captured
Speech captured over a real telephony path — a genuine call through the network, not studio audio downsampled afterwards to imitate one.
Scenario-driven calls between a caller and an agent or IVR flow, each leg recorded separately, across a deliberate spread of handsets and network conditions.
Yield at this style: 50–60% of recorded time is accepted. Dropped calls, network artefacts beyond tolerance and unusable legs are discarded. Real network conditions are the point of the style and also its main cost.
The specification
| Field | Value |
|---|---|
| Language | Assamese (as-IN, Assamese (Eastern Nagari)) |
| Volume | 50 hours |
| Equivalent | 100 speakers at 30 minutes each, or 50 speakers at one hour each |
| Speech type | Telephony Speech |
| Per speaker | 10–20 minutes of accepted audio per speaker |
| Sample rate | 8 kHz narrowband, matching what a deployed contact-centre model actually receives |
| Codec | G.711 and AMR-NB captured explicitly, with the codec recorded per call in the manifest |
| Legs | Caller and agent recorded on separate legs, never as a mixed call recording |
| Network conditions | Handset type, network carrier and packet-loss events logged per call |
| Dialects | Kamrupi, Goalparia, Upper Assam (Sibsagar standard), Barak Valley contact varieties |
| Transcription | Verbatim, native-speaker, second-pass reviewed |

Designing the call flows for Assamese
- Call scenarios drawn from real contact-centre intents: balance enquiry, complaint, booking change, escalation
- IVR flows scripted with deliberate mis-entry and barge-in paths, since those are where deployed systems fail
- Handset spread specified up front — low-end Android, feature phone, landline — because handset variance is a real acoustic axis
- Background conditions varied on purpose: street, vehicle, indoor, since real callers are rarely in quiet rooms
- Built against Assamese specifically: Assamese has the voiceless velar fricative /x/, unique among major Indian languages and routinely mis-modelled
- Code-mixing handled explicitly rather than edited out — Assamese speech mixes Hindi, English and Bengali, with substantial contact influence in Barak Valley and tea-garden communities.
Running a pilot scale Assamese build
Three to four weeks from signed specification. Recruitment is the critical path, not recording — the studio time itself is under two weeks.
Two batches: a 10-hour first batch in week two so you can run it through your pipeline while the rest is still recording, then the balance on completion.
Expect longer fielding times and higher per-hour cost than for Hindi or Marathi; the speaker pool with transcription-grade literacy is smaller.
| Parameter | At this volume |
|---|---|
| Cities | One, usually the densest pool for the language — Guwahati, Jorhat |
| Studios | A single treated room |
| Recruiters | One coordinator |
| Speakers | ~100–120 |
| Sessions per day | 6–8 |
| Team | 1 coordinator, 2 engineers, 3 transcribers |
Cohort design
At 50 hours the cohort is deliberately simplified: two or three dialect groups rather than the full spread, with quotas enforced in aggregate. A build this size cannot support per-cell targets and should not claim to.
Recruitment must be stratified by handset and carrier as well as by dialect, which adds a screening axis the other styles do not have.
| Dimension | Typical split | Why it matters for Assamese |
|---|---|---|
| Gender | 50 / 50 | Pitch range differences change acoustic model behaviour; unbalanced cohorts bias recognition |
| Age | 18-25: 30%, 26-40: 40%, 41-60: 30% | Older speakers retain conservative Assamese forms that younger urban speakers have lost |
| Region | Assam / Arunachal Pradesh / parts of Nagaland and Meghalaya and others | Dialect spread across 4 recognised varieties |
| Education | Mixed, including below-graduate | Prompt-reading fluency correlates with education and skews prosody |
| Condition | Studio / quiet room / field | Match the noise profile of your deployment |
Assamese-specific considerations
- Assamese has the voiceless velar fricative /x/, unique among major Indian languages and routinely mis-modelled
- Assamese speech mixes Hindi, English and Bengali, with substantial contact influence in Barak Valley and tea-garden communities.
- Extremely low-resource. Almost no spontaneous Assamese speech data exists publicly, and non-standard dialects have none.
Quality gates for telephony speech
- Codec and sample rate verified per file, rejecting any studio audio that has been downsampled to fake a telephony path
- Echo and double-talk checked on both legs
- DTMF events verified against the call log
- Level normalisation applied per leg, since handset output levels vary far more than studio microphones do
- Bengali characters substituted for Assamese ৰ / ৱ
- Goalparia and Kamrupi forms standardised to Sibsagar Assamese
100% content QA. At this volume there is no reason to sample, and a pilot exists precisely to surface specification problems.
What goes wrong on telephony speech sessions
- Studio audio downsampled to 8 kHz and passed off as telephony, which has none of the codec or packet-loss characteristics that matter
- Echo and double-talk that make the agent leg unusable
- Carrier and handset monoculture, producing a corpus that only represents one acoustic path
- Over-clean recordings from participants who move somewhere quiet to take the call, defeating the purpose
Risks at pilot scale in Assamese
Cost at this band is driven by: Fixed setup cost is amortised over very few hours, so the per-hour rate is at its highest here; Narrow demographic requirements bite hardest at small cohort sizes.
- The specification itself is usually the risk, not delivery — pilots exist to find the clauses that do not survive contact with real speakers
- A single-city cohort will not represent the language nationally, and reading pilot results as national is the most common mistake at this band
- Assamese carries 4 recognised varieties across Assam, Arunachal Pradesh, parts of Nagaland and Meghalaya, so the quota matrix is wider than the headline volume suggests
- Extremely low-resource. Almost no spontaneous Assamese speech data exists publicly, and non-standard dialects have none.
Deliverables
- WAV audio to your naming convention, with the per-file manifest
- Verbatim Assamese (Eastern Nagari) transcripts with utterance-level timestamps
- Call metadata: duration, codec, handset class, carrier, packet-loss events
- Intent label per call and per turn
- DTMF and hold, transfer and barge-in events
- Agent and caller leg identifiers
- Per-speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates, rejection reasons and agreement statistics
- Speaker-disjoint train / dev / test splits on request
What this trains, and what it does not
- Contact-centre and IVR ASR
- Voice bots operating over the phone network
- Intent classification on narrowband audio
- Robustness to codec and packet loss
Narrowband telephony audio is the wrong input for TTS or any wideband model — the frequency content simply is not there. Use it for models that will be deployed on a phone line and nothing else.
Frequently asked
Is 50 hours of Assamese enough?
Enough to validate a specification, benchmark a vendor and produce a small evaluation set. Not enough to move a production model's error rate.
Why telephony speech rather than another speech type?
Contact-centre and IVR ASR, Voice bots operating over the phone network, Intent classification on narrowband audio are what this style is the right input for. Narrowband telephony audio is the wrong input for TTS or any wideband model — the frequency content simply is not there. Use it for models that will be deployed on a phone line and nothing else.
How long does a pilot scale Assamese build take?
Three to four weeks from signed specification. Recruitment is the critical path, not recording — the studio time itself is under two weeks. Two batches: a 10-hour first batch in week two so you can run it through your pipeline while the rest is still recording, then the balance on completion.
What does 50 hours of Assamese telephony speech cost?
Quoted per delivered hour against this specification. At this band the drivers are fixed setup cost is amortised over very few hours, so the per-hour rate is at its highest here and narrow demographic requirements bite hardest at small cohort sizes. Send the spec and you get one fixed figure.
How much QA is applied at this volume?
100% content QA. At this volume there is no reason to sample, and a pilot exists precisely to surface specification problems.
Quote this Assamese dataset
50 hours, telephony speech, Assamese — pilot scale. Adjust anything and send it.