Process
From enquiry to delivered corpus
Every project runs the same four stages, whether it is 50 hours in one language or 5,000 across twelve.

1. Requirement intake
You send what you know: languages, hours or speaker counts, minutes per speaker, demographic splits, recording quality, speech type and deadline. Gaps are normal — we ask once and fill them in the scoping call.
If you are early and only have a model target, that is workable too. We translate 'our ASR degrades on Marathi call audio' into a corpus specification with numbers attached.
2. Scope and quote
We return a recruitment plan with dialect and demographic quotas, a recording protocol covering sampling rate, environment and capture chain, an annotation specification, and acceptance criteria.
Pricing is fixed against that scope. Larger programmes are invoiced against milestones, and most start with a paid pilot so you validate the format before volume.
3. Collection across India
Native-speaker recruiters source participants in the regions where the target dialects are actually spoken. Partner studios and mobile kits handle capture; every session follows one protocol so data from different cities is interchangeable.
Consent is collected in the speaker's own language, in writing, and retained for audit. Speaker metadata is recorded at session time, not reconstructed later.
4. QA and delivery
Audio passes automated validation for clipping, noise floor, silence ratio and channel integrity. Transcripts go through two-pass human QA with a native reviewer, measured against a written style guide.
Delivery is in your ingest format — WAV plus transcript, JSON or CSV manifests, speaker metadata, and a QA report showing measured accuracy against the acceptance criteria.
See the process against your spec
Send the requirement and we will map it onto these four stages with dates and numbers.