Pricing
Speech Data Collection: cost and pricing in India
Indicative rate bands for speech data collection, the five specification decisions that move the number most, and how a fixed quote is put together from a written spec.

- Pricing model
- Per delivered unit, fixed against spec
- Minimum
- ~50 hours or equivalent per language
- Turnaround
- 3-6 weeks for 100-500 hours in a single language; multi-language programmes run in parallel.
- Quote time
- One working day
What drives the price of speech data collection
Recruited-speaker speech corpora recorded to a written specification: scripted prompts, spontaneous monologue, or both, with full speaker metadata.
- Language tier: widely spoken languages have deeper recruitment pools and cost less per unit than low-resource ones
- Demographic narrowness: a 50/50 gender split across three age bands is routine; a narrow cohort multiplies recruitment effort
- Sample rate — 48 kHz capture, delivered at 48/16 kHz as required. Moving this off the standard changes the unit cost directly.
- Bit depth — 24-bit capture, 16-bit PCM delivery. Moving this off the standard changes the unit cost directly.
- Format — WAV (PCM), one file per utterance or per session. Moving this off the standard changes the unit cost directly.
- Channels — Mono per speaker; multi-channel on request. Moving this off the standard changes the unit cost directly.
- QA threshold and rework policy: a higher accept bar means more re-collection, which is priced in rather than absorbed later
Where the cost sits in speech data collection
Speakers are recruited through studio-local networks and screened by native coordinators against your quota matrix before any recording time is booked.
Turnaround: Typical: 3-6 weeks for 100-500 hours in a single language; multi-language programmes run in parallel.

Indicative rate bands
| Scope | Unit | Indicative band |
|---|---|---|
| Widely spoken language, standard spec | per delivered hour | ₹3,200 – ₹6,500 |
| Regional language, standard spec | per delivered hour | ₹4,200 – ₹8,200 |
| Low-resource language or narrow cohort | per delivered hour | ₹5,500 – ₹11,000 |
| Verbatim transcription add-on | per audio hour | ₹900 – ₹2,400 |
| Pilot batch (10-20 hours) | fixed | Quoted separately, credited against the full run |
These are bands, not a price list. Two projects with the same hour count can differ by 3x on cohort design alone, which is why every number here is replaced by a fixed figure once we see your specification.
What the price includes
- Audio files in the agreed format and naming convention
- Per-utterance manifest (speaker ID, prompt ID, duration, condition)
- Speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates and rejection reasons
How the quote is built
- Requirement lock: languages, hours, speaker count, demographic quotas, recording conditions
- Prompt design and linguistic review by native reviewers
- Speaker recruitment and screening against quota, with consent capture
- Recording sessions with real-time level and prompt-coverage monitoring
- Automated technical QA on every file (SNR, clipping, duration, silence)
- Native-speaker content QA on a defined sample, escalating to 100% on failure
- Packaging, manifest generation and delivery
Where budgets get wasted
The most expensive mistake in a data purchase is buying volume before the specification is settled. Re-collecting 200 hours because the noise profile did not match deployment costs more than the entire QA layer would have.
Every file passes automated technical checks. Content QA is sampled at 10% by default and raised per batch when the failure rate crosses the agreed threshold.
Best suited for
- ASR training
- TTS training
- Speaker ID
- Accent adaptation
- Benchmark sets
Frequently asked
How much does speech data collection cost per hour in India?
Most programmes land between ₹3,200 and ₹11,000 per delivered unit depending on language tier, cohort narrowness, recording condition and annotation depth. A written specification converts that band into a single fixed number.
Is there a minimum order?
Around 50 hours or equivalent per language for a production run. Pilots of 10-20 hours are accepted when they lead into a larger build, since most setup cost sits in specification and recruitment.
Do you charge for rejected audio?
No. You pay for delivered units that pass the agreed QA threshold. Rejected material is re-collected at our cost.
Can the price be fixed rather than estimated?
Yes. Once the specification is signed the price is fixed against it. Changes to quotas, conditions or annotation depth are re-quoted before work continues.
Related pages
Get a fixed price for speech data collection
Send the volume, languages and quality bar you have in mind and you get a scoped price within a working day.