Recording specs
Should you collect telephony or studio speech data?
Updated 2026-08-01 · 4 min read

Short answer
Match the channel to the deployment. If your model serves phone calls, collect over a real telephony path at 8 kHz with genuine codec, handset and network effects; if it serves an app, a device or a TTS product, collect in a studio at 48 kHz. Downsampling studio audio to simulate telephony is the most common shortcut and it reliably underperforms, because band-limiting removes frequencies without adding the compression artefacts, packet loss and handset colouration a real call carries.
Key takeaways
- Simulated telephony data does not reproduce codec and network artefacts.
- Parallel studio and telephony captures of the same speakers enable channel-robustness training.
- Channel mismatch between training and deployment is a leading cause of production WER surprises.
What a real telephony capture involves
Speakers call into a recording bridge on their own handsets over their own networks, so the corpus carries the handset diversity, background environments and network conditions of the real user base.
This also brings the noise your model will actually meet: traffic, fans, television, family conversation, poor signal.
What studio capture is for
Controlled acoustics for TTS, wake-word and device-audio work, where the deliverable is clean wideband audio and the noise conditions are added deliberately rather than accidentally.

Parallel capture
Recording the same speakers simultaneously in-studio and over a telephony path produces aligned pairs. Those pairs are unusually valuable for channel-robustness training and for measuring exactly how much your model degrades on the narrowband path.
Specifying it
Write the channel into the specification explicitly, including whether telephony is real or simulated, the codec, and the handset mix. 'Phone-quality audio' is not a specification and will be interpreted differently by every vendor.
Frequently asked questions
Should you collect telephony or studio speech data?
Match the channel to the deployment. If your model serves phone calls, collect over a real telephony path at 8 kHz with genuine codec, handset and network effects; if it serves an app, a device or a TTS product, collect in a studio at 48 kHz. Downsampling studio audio to simulate telephony is the most common shortcut and it reliably underperforms, because band-limiting removes frequencies without adding the compression artefacts, packet loss and handset colouration a real call carries.
What a real telephony capture involves?
Speakers call into a recording bridge on their own handsets over their own networks, so the corpus carries the handset diversity, background environments and network conditions of the real user base.
What studio capture is for?
Controlled acoustics for TTS, wake-word and device-audio work, where the deliverable is clean wideband audio and the noise conditions are added deliberately rather than accidentally.
Parallel capture?
Recording the same speakers simultaneously in-studio and over a telephony path produces aligned pairs. Those pairs are unusually valuable for channel-robustness training and for measuring exactly how much your model degrades on the narrowband path.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.