Speech data licensing determines what a buyer may do with delivered audio and which rights and records support that use. Those questions should be resolved before collection, access, or model training begins. A later consent may support future processing if it is valid, but it does not retroactively cure collection or processing that lacked a valid legal basis at the time.
This guide is a procurement checklist, not legal advice. License scope, data-protection roles, lawful basis, consent language, biometric use, withdrawal handling, model-related rights, and exclusivity all depend on the project and jurisdiction. Use the questions below to brief qualified counsel and put the agreed answers in the signed contract.
Read the grant, not the license label
Exclusive and non-exclusive are only starting points. A non-exclusive grant may let a supplier license the same material to others; an exclusive grant restricts reuse only to the extent the contract says so. Define whether exclusivity covers the exact recordings, speakers, prompts, language, territory, period, use case, or derived material. Do not infer any of those points from a product-page label.
For every purchase, confirm the permitted purposes, territory, term, sublicensing and contractor access, retention and deletion duties, security requirements, onward transfer, audit rights, warranties, indemnities, and restrictions such as biometric identification or synthetic-voice generation. A research-only grant does not authorize commercial deployment unless the text says it does.
| Grant field | Verify before signing |
|---|---|
| Permitted uses | Training, fine-tuning, evaluation, benchmarking, and commercial deployment are named expressly; research-only wording is not stretched to production. |
| Model and derivative rights | Ownership and use of trained weights, outputs, synthetic speech, and derived datasets are allocated in the text rather than assumed. |
| Exclusivity scope | Exactly what exclusivity covers: the recordings, speakers, prompts, language, territory, period, use case, or derived material. |
| Sublicensing and access | Affiliates, contractors, and cloud processors are authorized where needed, with any notification or flow-down duties stated. |
| Retention and deletion | Retention limits, deletion triggers, certification duties, and what happens to models trained before termination. |
| Provenance evidence | Consent or notice records, the source chain, and the documentation actually delivered with the data, not just referenced in marketing. |
| Restrictions | Prohibited uses such as biometric identification, voice cloning, or resale are listed and reconciled with the intended roadmap. |
These grant fields apply whether you license an existing corpus or commission collection; a per-language configuration such as the Chinese speech dataset or Russian speech dataset page shows where Spirelight documents each of them in a project scope.
Score these fields across competing offers with the free speech data provider scorecard, whose rights and provenance criteria carry the heaviest default weights, and inspect the specimen data card to see one way provenance fields can be structured in delivery.
Spirelight's dataset pages describe custom collection configurations and planning inputs; they are not a promise of finished, immediately licensable inventory. Confirm feasibility, sample status, scope, rights, provenance evidence, delivery schedule, and final price in the written quote and contract.
Allocate model and derivative rights expressly
A data license should say which party may train, fine-tune, evaluate, deploy, and commercialize models using the delivery. It should separately address trained weights, outputs, transcripts, embeddings, synthetic speech, and any retained copies of source audio. The answer is contractual and project-specific; do not assume that ownership of a model, an output, or a synthetic voice follows automatically from access to the recordings.
Text-to-speech, voice cloning, speaker identification, and biometric authentication warrant explicit use restrictions and contributor-facing language. Ask counsel to align the supplier agreement, contributor notice or consent, and intended product use before training begins.
Lawful basis, permissions, and provenance are related but distinct
Provenance records where audio came from and how it moved through the supply chain. A license records contractual permission; free corpora need the same reading, as the LibriSpeech, LJSpeech, and Common Voice breakdown shows corpus by corpus. GDPR lawful basis governs processing of personal data. Consent may be the relevant lawful basis or an additional permission in some projects, but it is not the only possible Article 6 basis and a license from a supplier does not by itself establish GDPR compliance.
Determine the lawful basis and required notices before collection or reuse. If relying on consent, verify that it was freely given, specific, informed, unambiguous, and capable of being demonstrated, and that withdrawal can be handled. Link the applicable record to the contributor and files closely enough to support the intended audit trail. If another lawful basis is used, document that assessment and any purpose-compatibility analysis instead of presenting a consent form as a substitute.
A later consent can authorize future processing if it is valid, but it does not retrospectively legalize earlier collection or processing that lacked a valid basis. Remediation, deletion, and effects on an existing model are fact-specific questions for counsel.
Public availability is not the same as permission to train. For scraped, open, customer-call, or repurposed audio, verify copyright and database rights, contract terms, notices, lawful basis, purpose limitation, and any sector-specific recording rules for the actual source.
When voice data is biometric special-category data
A voice recording is personal data under the GDPR when it relates to an identified or identifiable person. The controller still needs an Article 6 lawful basis. Voice is not automatically special-category data merely because it is audio. Under Article 9, biometric data falls within the special categories when it results from specific technical processing and is used for the purpose of uniquely identifying a person.
If Article 9 applies, the controller needs both an Article 6 basis and an applicable Article 9(2) condition. Explicit consent is one Article 9(2) condition, not the only one. Authentication, speaker recognition, and voiceprint use therefore need a separate assessment from ordinary transcription. The official text is in GDPR Articles 4, 5, 6, 7, and 9 on EUR-Lex.
Ask where contributors were located, who acts as controller or processor, what basis and notices apply, whether cross-border transfers occur, which Article 9 condition applies if biometric identification is intended, and how access, objection, withdrawal, erasure, retention, and model-impact requests will be handled. UK processing requires a separate UK-law assessment.
What the signed agreement and evidence pack should cover
For the specific delivery, identify the files or collection scope, acceptance criteria, permitted and prohibited uses, territory, term, exclusivity, downstream access, retention, deletion, security, and any model-related rights. Require only warranties that match evidence the supplier can actually produce. The evidence pack may include the source and license chain, contributor notice or consent records where applicable, collection instructions, file-to-contributor mapping, transfer records, and agreed QA documentation.
Spirelight does not assume exclusivity, commercial-use scope, biometric permission, or model ownership from a catalogue page. These terms and the records available for a delivery must be confirmed project by project in the signed agreement. Use this guide to frame diligence, then have qualified counsel assess the applicable facts and jurisdictions.