Recording specs
How do you collect noisy and far-field speech data?
Updated 2026-08-01 · 4 min read

Short answer
Collect real noise rather than adding it later wherever the deployment allows. Real environments contribute noise, reverberation, speaker movement and Lombard effect — people speak differently in noise — and only the first of those can be simulated convincingly by mixing noise into clean audio. For far-field work, record at the microphone geometry and distances of the device, in rooms with the reverberation characteristics of the target environment, with parallel close-mic capture as a clean reference.
Key takeaways
- Lombard effect means noisy-environment speech is different speech, not clean speech plus noise.
- Record at the device's real microphone geometry for far-field models.
- Parallel close-mic reference tracks make noisy corpora far more useful.
Why augmentation is not enough
Mixing noise into clean recordings changes the signal but not the speaking behaviour. In real noise people raise volume, slow down, hyper-articulate and shift pitch. Models trained only on augmented data see none of that.
Far-field specifics
Distance, room reverberation and microphone array geometry dominate far-field performance. Record with the actual array where possible, at the distances users will stand, in rooms representative of the deployment — Indian homes, shop floors, car interiors.

Parallel references
A close-mic track recorded simultaneously gives an aligned clean reference for every noisy utterance. That pairing supports enhancement training, robustness evaluation and honest transcription of audio that would otherwise be ambiguous.
Specifying noise
Describe the environments, approximate SNR bands and the proportion of the corpus in each. 'Some noisy data' is not specifiable and will be delivered as whatever was convenient.
Frequently asked questions
How do you collect noisy and far-field speech data?
Collect real noise rather than adding it later wherever the deployment allows. Real environments contribute noise, reverberation, speaker movement and Lombard effect — people speak differently in noise — and only the first of those can be simulated convincingly by mixing noise into clean audio. For far-field work, record at the microphone geometry and distances of the device, in rooms with the reverberation characteristics of the target environment, with parallel close-mic capture as a clean reference.
Why augmentation is not enough?
Mixing noise into clean recordings changes the signal but not the speaking behaviour. In real noise people raise volume, slow down, hyper-articulate and shift pitch. Models trained only on augmented data see none of that.
Far-field specifics?
Distance, room reverberation and microphone array geometry dominate far-field performance. Record with the actual array where possible, at the distances users will stand, in rooms representative of the deployment — Indian homes, shop floors, car interiors.
Parallel references?
A close-mic track recorded simultaneously gives an aligned clean reference for every noisy utterance. That pairing supports enhancement training, robustness evaluation and honest transcription of audio that would otherwise be ambiguous.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.