Turnaround
How long conversational speech data takes
Honest delivery schedules for conversational speech data, the stages that actually consume the calendar, and the three levers that safely compress a timeline.

- Standard
- 4-7 weeks for 100-300 hours in one language.
- Pilot
- 10-20 units in 1-2 weeks
- First batch
- Typically week 3
- Cadence
- Weekly batches after the first
Where the time actually goes
Recording is rarely the bottleneck. Specification agreement, speaker recruitment against quotas and native-reviewer QA are what set the calendar, and all three run in parallel once the spec is signed.
- Scenario design so conversations stay on-domain without becoming scripted
- Speaker pairing by familiarity level as specified
- Multi-channel recording with per-speaker isolation
- Diarisation-ready packaging
- Verbatim transcription with speaker turns and overlap tags
Indicative schedules
| Volume | Spec + setup | Collection | QA + delivery | Total |
|---|---|---|---|---|
| Pilot (10-20 units) | 3-5 days | 5-7 days | 3-4 days | 1-2 weeks |
| 100 units | 1 week | 2-4 weeks | 1 week | 4-6 weeks |
| 500 units | 1-2 weeks | 4-7 weeks | 1-2 weeks | 6-10 weeks |
| 1,000 units | 2 weeks | 7-11 weeks | 2 weeks | 10-16 weeks |
Batches are delivered as they pass QA, so you start training long before the final unit lands.

What safely compresses a timeline
- Parallel cities: the same cohort recruited across several studios at once rather than sequentially
- A signed specification on day one — most slipped schedules start as an unsettled spec
- Accepting rolling batches instead of a single final delivery
- Widening one non-critical quota (usually education band or district) while holding the ones that matter
What cannot be rushed
Recruitment of narrow cohorts, native review depth and consent administration have floors. Compressing them is how vendors end up delivering data that fails acceptance, which costs more calendar time than it saved.
Conversations are checked for naturalness, not just audio quality. Sessions that drift into reading are rejected.
What lands with each batch
- Per-speaker channels plus mixed reference
- Turn-level transcripts
- Overlap and backchannel tags
- Conversation metadata
Frequently asked
What is a realistic turnaround for conversational speech?
4-7 weeks for 100-300 hours in one language. for a standard production run, with a pilot possible in one to two weeks and first batches usually arriving in week three.
Can you hit a hard deadline?
Often, by running more cities in parallel and delivering rolling batches. We will say no rather than accept a date that forces a quality compromise.
Do we wait for everything before we can train?
No. Batches are delivered as they pass QA so training can start early and the spec can be adjusted while collection is still running.
What causes delays in practice?
Late specification sign-off, mid-project quota changes and narrow cohorts in low-density districts. All three are visible in the plan before we start.
Related pages
Need conversational speech by a date?
Send the deadline with the volume. You get a schedule that says what is achievable and what is not.