Service
Speech Data Collection in India
Recruited-speaker speech corpora recorded to a written specification: scripted prompts, spontaneous monologue, or both, with full speaker metadata.

- Turnaround
- 3-6 weeks for 100-500 hours in a single language; multi-language programmes run in parallel.
- Languages
- 14 Indian languages + Indian English
- Delivery
- Audio files in the agreed format and naming convention
What you get
- Audio files in the agreed format and naming convention
- Per-utterance manifest (speaker ID, prompt ID, duration, condition)
- Speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates and rejection reasons
Technical specification
Every parameter below is written into the statement of work before recording begins. If your pipeline needs different values, they replace ours rather than being converted after delivery.
| Parameter | Standard |
|---|---|
| Sample rate | 48 kHz capture, delivered at 48/16 kHz as required |
| Bit depth | 24-bit capture, 16-bit PCM delivery |
| Format | WAV (PCM), one file per utterance or per session |
| Channels | Mono per speaker; multi-channel on request |
| Noise floor | Studio sessions below -50 dBFS; field sessions specified per project |
| Clipping | Zero tolerance; clipped takes are re-recorded, not repaired |

How the work runs
- Requirement lock: languages, hours, speaker count, demographic quotas, recording conditions
- Prompt design and linguistic review by native reviewers
- Speaker recruitment and screening against quota, with consent capture
- Recording sessions with real-time level and prompt-coverage monitoring
- Automated technical QA on every file (SNR, clipping, duration, silence)
- Native-speaker content QA on a defined sample, escalating to 100% on failure
- Packaging, manifest generation and delivery
Quality control
Every file passes automated technical checks. Content QA is sampled at 10% by default and raised per batch when the failure rate crosses the agreed threshold.
QA failures are remedied by re-collection, not by editing the delivered files. Repaired audio introduces artefacts that survive into your model.
Speaker and contributor sourcing
Speakers are recruited through studio-local networks and screened by native coordinators against your quota matrix before any recording time is booked.
Consent is captured per participant and mapped to file IDs, so provenance survives an external audit of your training data.
Timeline
Typical: 3-6 weeks for 100-500 hours in a single language; multi-language programmes run in parallel.
Staged delivery is available: first batches ship while later batches are still recording, so training can start early.
The numbers we hold ourselves to
- 100% of delivered files pass automated technical QA for SNR, clipping, duration and silence
- 5-25% of files pass a second native-speaker content review, stratified by city, dialect and transcriber, and escalating to 100% on any batch that fails the agreed threshold
- Accepted yield runs 85-90% for scripted speech, 60-70% for spontaneous, 55-65% for conversational and 50-60% for telephony
- Default cohort quotas: 50/50 gender, with age bands at 30% (18-25), 40% (26-40) and 30% (41-60)
- 48 kHz / 24-bit capture, delivered as 16-bit PCM WAV, with studio sessions held below a -50 dBFS noise floor
- First response within one working day; a scoped, fixed quote within two to three
These are the figures a delivery is measured against, not aspirations. A batch that misses them is re-recorded at our cost rather than repaired.
Commonly used for
- ASR training
- TTS training
- Speaker ID
- Accent adaptation
- Benchmark sets
Frequently asked
What is the minimum volume for speech data collection?
Programmes typically start around 50 hours or equivalent units per language. Smaller pilots are accepted when they lead into a larger build, because most of the setup cost is in specification and recruitment rather than recording time.
Can you work to our schema instead of yours?
Yes. Manifest fields, file naming, directory structure and label schema are set by you. Working to your schema from the start avoids a conversion pass that usually loses metadata.
Who owns the delivered data?
You do. Deliverables come with a perpetual, transferable licence and participant consent that covers model training and distribution of the resulting model.
How is pricing structured?
Per delivered hour or per unit, quoted against a written specification. Quotas, recording conditions and QA thresholds all move the price, which is why we quote from a spec rather than from a price list.
Get a quote for speech data collection
Send the specification you already have, or the rough shape of it, and you get a scoped quote with a timeline.