aidataservices.inAI data collection · India

Service

ASR Training Data in India

Transcribed speech corpora built to train and evaluate automatic speech recognition, with verbatim transcription, timestamps and per-token language tagging where code-mixing occurs.

Request a dataset quoteReply within one working day
Audio waveforms being prepared as ASR training data — ASR Training Data in India
Turnaround
Transcription adds roughly 1-2 weeks per 100 hours after recording, per language.
Languages
14 Indian languages + Indian English
Delivery
Audio plus aligned transcripts (JSON/TSV, or your schema)
01

What you get

  • Audio plus aligned transcripts (JSON/TSV, or your schema)
  • Style guide as delivered documentation
  • Inter-annotator agreement report
  • Speaker-disjoint train/dev/test splits
SpecRecruitRecordTranscribeQADeliverFails the accept bar → re-recorded, not repairedOne owner, no handoffs between vendors
02

Technical specification

Every parameter below is written into the statement of work before recording begins. If your pipeline needs different values, they replace ours rather than being converted after delivery.

ParameterStandard
Audio16 kHz or 48 kHz PCM WAV
TranscriptionVerbatim, including disfluencies, false starts and fillers
TimestampsUtterance level by default; word level on request
TaggingNoise, overlap, unintelligible, foreign-language and code-switch tags
NormalisationRaw and normalised text columns delivered separately
SplitTrain/dev/test splits with no speaker leakage across splits
Studio-grade voice recording session for text-to-speech training data — supporting asr training data in india
Studio-grade voice recording session for text-to-speech training data
03

How the work runs

  • Style guide authored per language, covering numerals, loanwords, script and disfluency rules
  • Transcriber calibration round with inter-annotator agreement measurement
  • First-pass transcription
  • Second-pass native review
  • Automated consistency checks against the style guide
  • Split generation with speaker-disjoint partitions
04

Quality control

Agreement is measured, not assumed. We report word-level agreement on a held-out sample so you can judge label quality before training on it.

QA failures are remedied by re-collection, not by editing the delivered files. Repaired audio introduces artefacts that survive into your model.

05

Speaker and contributor sourcing

Transcribers are native speakers of the target variety, not of a related standard language.

Consent is captured per participant and mapped to file IDs, so provenance survives an external audit of your training data.

06

Timeline

Transcription adds roughly 1-2 weeks per 100 hours after recording, per language.

Staged delivery is available: first batches ship while later batches are still recording, so training can start early.

07

The numbers we hold ourselves to

  • 100% of delivered files pass automated technical QA for SNR, clipping, duration and silence
  • 5-25% of files pass a second native-speaker content review, stratified by city, dialect and transcriber, and escalating to 100% on any batch that fails the agreed threshold
  • Accepted yield runs 85-90% for scripted speech, 60-70% for spontaneous, 55-65% for conversational and 50-60% for telephony
  • Default cohort quotas: 50/50 gender, with age bands at 30% (18-25), 40% (26-40) and 30% (41-60)
  • 48 kHz / 24-bit capture, delivered as 16-bit PCM WAV, with studio sessions held below a -50 dBFS noise floor
  • First response within one working day; a scoped, fixed quote within two to three

These are the figures a delivery is measured against, not aspirations. A batch that misses them is re-recorded at our cost rather than repaired.

08

Commonly used for

  • ASR training
  • ASR fine-tuning
  • Domain adaptation
  • WER benchmarking

Frequently asked

What is the minimum volume for asr datasets?

Programmes typically start around 50 hours or equivalent units per language. Smaller pilots are accepted when they lead into a larger build, because most of the setup cost is in specification and recruitment rather than recording time.

Can you work to our schema instead of yours?

Yes. Manifest fields, file naming, directory structure and label schema are set by you. Working to your schema from the start avoids a conversion pass that usually loses metadata.

Who owns the delivered data?

You do. Deliverables come with a perpetual, transferable licence and participant consent that covers model training and distribution of the resulting model.

How is pricing structured?

Per delivered hour or per unit, quoted against a written specification. Quotas, recording conditions and QA thresholds all move the price, which is why we quote from a spec rather than from a price list.

Get a quote for asr training data

Send the specification you already have, or the rough shape of it, and you get a scoped quote with a timeline.

Request a dataset quote