Process
How long does a speech data collection project take?
Updated 2026-08-01 · 4 min read

Short answer
A 100–300 hour single-language Indian corpus typically takes three to six weeks from signed scope to final delivery, with rolling batches from week two. Multi-language programmes run in parallel, so five languages take closer to six to eight weeks than to five times as long. The variables that stretch timelines are rare dialect quotas that need field recruitment, narrow demographic windows, telephony capture that depends on participant availability, and annotation depth — verbatim timestamped transcription with speaker labels can take longer than the recording itself.
Key takeaways
- Recording is rarely the bottleneck; recruitment and annotation are.
- Rolling batch delivery starts around week two, so training need not wait for the end.
- Rare dialects and narrow demographics add weeks, not days.
A typical schedule
Week 0: scope and protocol sign-off. Week 1: recruitment, screening and studio scheduling. Weeks 2–4: recording with rolling QA and batch delivery. Weeks 4–6: annotation completion, final QA and delivery of manifests and consent register.
What compresses it
Parallel studios across cities, a wide demographic window, standard transcription depth, and a client team that can review batch one quickly. The last of these is more often the constraint than anyone expects.

What extends it
Rare dialect quotas requiring travel, participants who must meet narrow professional criteria, real telephony capture, word-level timestamps, and multi-pass annotation with entity tagging.
Planning around it
Sequence the programme so the first batch matches your highest-priority slice. If the schedule slips, you still have the data that matters most rather than an even spread of everything.
Frequently asked questions
How long does a speech data collection project take?
A 100–300 hour single-language Indian corpus typically takes three to six weeks from signed scope to final delivery, with rolling batches from week two. Multi-language programmes run in parallel, so five languages take closer to six to eight weeks than to five times as long. The variables that stretch timelines are rare dialect quotas that need field recruitment, narrow demographic windows, telephony capture that depends on participant availability, and annotation depth — verbatim timestamped transcription with speaker labels can take longer than the recording itself.
A typical schedule?
Week 0: scope and protocol sign-off. Week 1: recruitment, screening and studio scheduling. Weeks 2–4: recording with rolling QA and batch delivery. Weeks 4–6: annotation completion, final QA and delivery of manifests and consent register.
What compresses it?
Parallel studios across cities, a wide demographic window, standard transcription depth, and a client team that can review batch one quickly. The last of these is more often the constraint than anyone expects.
What extends it?
Rare dialect quotas requiring travel, participants who must meet narrow professional criteria, real telephony capture, word-level timestamps, and multi-pass annotation with entity tagging.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.