Quality & compliance
Quality standards, consent and IP
A dataset is only useful if you can defend both its accuracy and its provenance. Both are contractual for us.

Recording standards
Studio collections are captured at 48 kHz / 24-bit or 16 kHz / 16-bit as specified, mono or multi-channel, with a documented microphone and interface chain. Telephony corpora are captured over a real narrowband path rather than downsampled from studio audio.
Every file is checked for clipping, noise floor, DC offset, silence ratio and channel separation before it enters the QA queue. Files outside tolerance are re-recorded, not patched.
Transcription and annotation QA
Transcripts are produced by native speakers of the target language and reviewed by a second native reviewer against a written style guide covering code-mixing, numerals, named entities, disfluencies and non-speech events.
Accuracy is sampled and reported per batch. Where you set a word error rate threshold in the acceptance criteria, the QA report shows measurement against it rather than an assurance.
Consent, privacy and IP
Every speaker signs informed consent in their own language covering the intended use, including commercial AI training. Consent records are retained and auditable.
Full IP and a perpetual, worldwide licence transfer to you on delivery unless a different arrangement is agreed. We do not resell delivered corpora and we do not retain them beyond the agreed retention window.
Personally identifiable content can be excluded at the prompt-design stage or redacted during annotation, depending on your compliance requirements.
Review our standards against yours
Send your acceptance criteria and compliance requirements. We will confirm what we can meet before you commit.