Buying guides
How do you write a speech data RFP?
Updated 2026-08-01 · 4 min read

Short answer
A speech data RFP should specify eight things: languages and dialect quotas, speaker count, minutes per speaker, demographic splits, recording condition and sample rate, speech type, annotation depth, and acceptance criteria with the metric that will be measured. Vendors who receive all eight return comparable fixed prices; vendors who receive 'we need Hindi speech data' return numbers that cannot be compared to each other. Add your delivery format and consent requirements and you have removed most of the reasons projects fail late.
Key takeaways
- Specify speakers and minutes separately — hours alone is not a specification.
- State the acceptance metric up front; it changes how vendors price QA.
- Ask for a paid pilot in your ingest format as part of the RFP.
The eight required fields
Languages and dialect quotas. Speaker count. Minutes per speaker. Demographic splits by gender, age and region. Recording condition, sample rate and bit depth. Speech type: scripted, spontaneous, conversational or telephony. Annotation depth: transcript only, timestamps, speaker labels, events, entities. Acceptance criteria with the measurement method.
Anything omitted becomes a vendor assumption, and every vendor assumes differently. That is why quotes for the same project vary by a factor of three.
Terms that protect you
IP assignment on delivery, no reuse or resale of the commissioned corpus, consent that explicitly covers commercial AI training, data handling and residency terms, and a defined remedy when a batch fails acceptance — re-collection rather than credit.

Ask for a pilot
Request 10–20 hours delivered in your ingest format before volume, priced at the volume rate. It surfaces format mismatches, annotation misreadings and acoustic problems while they are cheap to fix.
Red flags in responses
A quote with no acceptance criteria. A vendor who cannot describe their consent language. Price per hour with no speaker count attached. Claimed dialect coverage without a recruitment plan for the districts where those dialects are spoken.
Frequently asked questions
How do you write a speech data RFP?
A speech data RFP should specify eight things: languages and dialect quotas, speaker count, minutes per speaker, demographic splits, recording condition and sample rate, speech type, annotation depth, and acceptance criteria with the metric that will be measured. Vendors who receive all eight return comparable fixed prices; vendors who receive 'we need Hindi speech data' return numbers that cannot be compared to each other. Add your delivery format and consent requirements and you have removed most of the reasons projects fail late.
The eight required fields?
Languages and dialect quotas. Speaker count. Minutes per speaker. Demographic splits by gender, age and region. Recording condition, sample rate and bit depth. Speech type: scripted, spontaneous, conversational or telephony. Annotation depth: transcript only, timestamps, speaker labels, events, entities. Acceptance criteria with the measurement method.
Terms that protect you?
IP assignment on delivery, no reuse or resale of the commissioned corpus, consent that explicitly covers commercial AI training, data handling and residency terms, and a defined remedy when a batch fails acceptance — re-collection rather than credit.
Ask for a pilot?
Request 10–20 hours delivered in your ingest format before volume, priced at the volume rate. It surfaces format mismatches, annotation misreadings and acoustic problems while they are cheap to fix.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.