Service explainers
What is speech data collection and how does it work?
Updated 2026-08-01 · 4 min read

Short answer
Recruited-speaker speech corpora recorded to a written specification: scripted prompts, spontaneous monologue, or both, with full speaker metadata. In practice the work runs as requirement lock: languages, hours, speaker count, demographic quotas, recording conditions, then prompt design and linguistic review by native reviewers, then speaker recruitment and screening against quota, with consent capture, and you receive audio files in the agreed format and naming convention, per-utterance manifest (speaker id, prompt id, duration, condition). Typical: 3-6 weeks for 100-500 hours in a single language; multi-language programmes run in parallel.
Key takeaways
- Every file passes automated technical checks. Content QA is sampled at 10% by default and raised per batch when the failure rate crosses the agreed threshold.
- Typical buyers: ASR training, TTS training, Speaker ID.
- Recruitment approach: Speakers are recruited through studio-local networks and screened by native coordinators against your quota matrix before any recording time is booked.
What the service covers
Recruited-speaker speech corpora recorded to a written specification: scripted prompts, spontaneous monologue, or both, with full speaker metadata.
- Audio files in the agreed format and naming convention
- Per-utterance manifest (speaker ID, prompt ID, duration, condition)
- Speaker metadata: age band, gender, region, dialect, education band
- Consent records mapped to speaker IDs
- QA report with pass rates and rejection reasons
Technical specification
These are defaults, not limits. Where your pipeline requires different values, they replace ours in the statement of work rather than being converted after delivery.
| Parameter | Standard |
|---|---|
| Sample rate | 48 kHz capture, delivered at 48/16 kHz as required |
| Bit depth | 24-bit capture, 16-bit PCM delivery |
| Format | WAV (PCM), one file per utterance or per session |
| Channels | Mono per speaker; multi-channel on request |
| Noise floor | Studio sessions below -50 dBFS; field sessions specified per project |
| Clipping | Zero tolerance; clipped takes are re-recorded, not repaired |

How the work runs
- Requirement lock: languages, hours, speaker count, demographic quotas, recording conditions
- Prompt design and linguistic review by native reviewers
- Speaker recruitment and screening against quota, with consent capture
- Recording sessions with real-time level and prompt-coverage monitoring
- Automated technical QA on every file (SNR, clipping, duration, silence)
- Native-speaker content QA on a defined sample, escalating to 100% on failure
- Packaging, manifest generation and delivery
Quality control and acceptance
Every file passes automated technical checks. Content QA is sampled at 10% by default and raised per batch when the failure rate crosses the agreed threshold.
Failures are remedied by re-collection rather than by editing delivered files, because repaired audio carries artefacts that survive into the trained model.
Who this is for
Recruitment for this service works as follows. Speakers are recruited through studio-local networks and screened by native coordinators against your quota matrix before any recording time is booked.
- ASR training
- TTS training
- Speaker ID
- Accent adaptation
- Benchmark sets
Timelines
Typical: 3-6 weeks for 100-500 hours in a single language; multi-language programmes run in parallel.
Multi-language programmes run in parallel rather than in sequence, so a five-language scope does not take five times as long.
Frequently asked questions
What is included in speech data collection?
Audio files in the agreed format and naming convention, Per-utterance manifest (speaker ID, prompt ID, duration, condition), Speaker metadata: age band, gender, region, dialect, education band, delivered against a written specification with acceptance criteria attached.
How long does speech data collection take?
Typical: 3-6 weeks for 100-500 hours in a single language; multi-language programmes run in parallel.
How is quality measured?
Every file passes automated technical checks. Content QA is sampled at 10% by default and raised per batch when the failure rate crosses the agreed threshold.
Which languages are supported?
Fifteen Indian languages plus Indian English, including Hindi, Marathi, Tamil, Telugu, Kannada.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.