Compliance
Who owns a commissioned AI training dataset?
Updated 2026-08-01 · 4 min read

Short answer
In a properly drafted commission, you own it. The contract should assign all intellectual property in the recordings, transcripts and annotations to you on delivery, warrant that speaker consent covers commercial AI training by your organisation, and prohibit the vendor from reusing or reselling the corpus or any derivative of it. Where a vendor insists on retaining reuse rights in exchange for a lower price, that is a legitimate trade — but it must be an explicit decision, because a competitor training on your commissioned corpus erases the advantage you paid for.
Key takeaways
- IP assignment on delivery, not a licence, is the default to ask for.
- No-reuse and no-resale clauses protect the advantage you paid to create.
- Consent scope and IP assignment are separate protections; you need both.
Assignment vs licence
An assignment transfers ownership; a licence leaves it with the vendor. Licences are cheaper and appropriate for commodity corpora, but for data collected to your specification, assignment is the norm.
Reuse and exclusivity
Ask explicitly whether the vendor may reuse the corpus, derivatives, or the scripts and protocols. Some reuse — anonymised process learnings — is harmless; reuse of the audio is not.

Consent scope
Ownership of a recording does not grant the right to train on it if the speaker never consented to that use. Both protections must be in place, and the consent template should be an annexe to the contract.
Provenance for your customers
Enterprise buyers increasingly ask AI vendors where the training data came from. A clean chain — consent record, session log, assignment — is a commercial asset, not just legal hygiene.
Frequently asked questions
Who owns a commissioned AI training dataset?
In a properly drafted commission, you own it. The contract should assign all intellectual property in the recordings, transcripts and annotations to you on delivery, warrant that speaker consent covers commercial AI training by your organisation, and prohibit the vendor from reusing or reselling the corpus or any derivative of it. Where a vendor insists on retaining reuse rights in exchange for a lower price, that is a legitimate trade — but it must be an explicit decision, because a competitor training on your commissioned corpus erases the advantage you paid for.
Assignment vs licence?
An assignment transfers ownership; a licence leaves it with the vendor. Licences are cheaper and appropriate for commodity corpora, but for data collected to your specification, assignment is the norm.
Reuse and exclusivity?
Ask explicitly whether the vendor may reuse the corpus, derivatives, or the scripts and protocols. Some reuse — anonymised process learnings — is harmless; reuse of the audio is not.
Consent scope?
Ownership of a recording does not grant the right to train on it if the speaker never consented to that use. Both protections must be in place, and the consent template should be an annexe to the contract.
Related reading
Turn this into a dataset specification
Tell us the languages, speaker count and minutes. You get a written scope, a protocol and a fixed price within one working day.