"Free" and "free to use in a commercial product" are two different questions, and a lot of procurement research collapses them into one. LibriSpeech, LJSpeech, Common Voice, and AliMeeting are all free to download. Whether you can train and sell a model depends on the exact release licence and any other applicable source, platform, privacy, publicity, or personality rights—not on price alone.

This guide walks through the licence types you will actually meet, CC0, CC BY 4.0, CC BY-NC, CC BY-SA, and research-only terms, verifies what each of those four named corpora is actually licensed under, and lays out the risks a clean licence does not cover, alongside the cases where free data is genuinely the right call.

The short answer

Some open-data licences permit commercial exercise of the licensed rights, but that does not by itself establish complete legal clearance for a project. Others are not, and the fee, or lack of one, tells you nothing about which is which. The release licence is central, but commercial usability also depends on any separate source, platform, privacy, publicity, or personality rights and terms.

Take the query that sends people to this page more than any other: is there a free Greek ASR dataset cleared for commercial use? Mozilla lists Common Voice releases under CC0 unless otherwise specified; verify the current Greek release plus Mozilla Data Collective access, download, and forbidden-use terms before commercial use. Common Voice coverage differs by language and release; cite the dated release statistics and test whether its hours, validation status, speaker mix, and acoustics match the intended deployment. That gap, between legally clear and actually sufficient, runs through almost every free corpus, and it is what the rest of this guide works through.

The licences you will actually meet

Five licence patterns cover almost every open speech corpus you will run into. Knowing which one you are looking at, before a data pipeline gets built on top of it, is the entire exercise.

LicenceWhat it permits for a shipped commercial modelObligation it carries
CC0 / public-domain dedicationReview the exact release, jurisdictional effect, source permissions, and intended use.The stated dedication may remove copyright conditions; other rights and project duties still need assessment.
CC BY 4.0Review the exact release and whether the intended data and model uses fall within the grant.Attribution and notice conditions apply when Licensed Material, including modified material, is Shared; assess whether the planned distribution triggers them.
CC BY-NCNon-commercial use only. CC BY-NC grants the licensed rights only for uses not primarily intended for or directed toward commercial advantage or monetary compensation; whether a particular training or model use qualifies is fact-specific.The NonCommercial restriction applies to exercise of the licensed rights; attribution conditions apply when the material is Shared.
CC BY-SACC BY-SA permits commercial exercise of the licensed rights; if Adapted Material is Shared, the ShareAlike conditions apply. Whether a trained model is Adapted Material is jurisdiction-specific.Attribution and share-alike, which is hard to reconcile with a proprietary model.
Research-only or custom termsUsually restricted to non-commercial research, evaluation, or benchmarking; read the specific grant.Whatever the custom agreement states. Assume nothing beyond it.

The two that cause the most damage are CC BY-NC and CC BY-SA, because both look permissive at a glance and are not. Our guide to speech data licensing goes deeper into how these terms interact with model weights, which is a genuinely unsettled question even among IP lawyers.

The major open corpora and what their licences allow

Four corpora account for most of the free ASR dataset search traffic, and each sits in a different place on the table above. For the obligations in practice, what free for commercial use really means for LibriSpeech, LJSpeech, and Common Voice works through the same releases one by one.

LibriSpeech, hosted on OpenSLR as resource SLR12, is released under CC BY 4.0. It contains roughly 1,000 hours of read English speech drawn from LibriVox public-domain audiobooks, split into clean and noisier training, development, and test sets. LibriSpeech is published under CC BY 4.0, which permits commercial exercise of the licensed rights; assess the planned training and model distribution, and satisfy attribution when the licence requires it.

LJSpeech is public domain in the US: "no restrictions on its use," in the creator's own words. It is a single English speaker reading roughly 24 hours across 13,100 clips, drawn from seven public-domain non-fiction books and recorded by the LibriVox project. Attribution is requested, not required, which makes it one of the cleanest licences on this list, though its single-speaker, single-domain nature limits what it is good for beyond TTS work.

Mozilla Common Voice datasets are described by Mozilla as CC0 unless otherwise specified, but current access is through Mozilla Data Collective and a release page may impose separate conditions such as no re-identification or re-hosting. Review the exact release and download terms. It is also the widest in language coverage, spanning well over 100 languages, Greek and Vietnamese both included, because it is built by volunteer contributors reading and validating sentences rather than one fixed recording effort. Coverage differs by language and release, so check dated statistics, validation status, speaker mix, and acoustics rather than inferring fit from the contribution model.

AliMeeting, an Alibaba-released Mandarin meeting corpus hosted on OpenSLR as resource SLR119, is licensed CC BY-SA 4.0, not research-only as it is sometimes described. That is commercially usable, but the share-alike clause means a derivative you distribute may need to carry the same licence, a real complication if the derivative in question is a proprietary model. Our guide to open speech datasets catalogues these and other corpora in more depth, and our guide to ASR training data covers how corpora like these fit into a broader training pipeline.

The four risks free data carries into production

A license addresses only the rights within its grant and must be read for the exact release and intended use. It does not by itself resolve source permissions, privacy, publicity or personality rights, provenance, coverage, or deployment fit.

  • Attribution obligations. Under CC BY 4.0, attribution conditions apply when Licensed Material, including modified material, is Shared; assess separately whether and how a trained model triggers those conditions. If attribution lives only in a training script nobody maintains, it quietly falls out of your documentation and becomes a compliance gap the next time someone audits data lineage.
  • Share-alike contamination. CC BY-SA requires a compatible licence when Adapted Material is Shared; whether a trained model is Adapted Material is unresolved and jurisdiction-specific. Whether a trained model counts as a derivative work is genuinely unresolved, and legal teams read it differently. The conservative position, and the one most counsel land on, is to treat CC BY-SA data as carrying real risk for a proprietary model.
  • Unclear speaker consent. A Creative Commons licence governs copyright in the recording and transcript. It says nothing about whether the person speaking agreed to have their voice used to train a commercial product, a separate question the licence cannot answer either way.
  • No indemnity. Open datasets ship with no warranty and no one to call if a rights claim surfaces later. A commercial vendor contract typically gives you a counterparty and some contractual protection; a public download gives you neither.

When free data is genuinely the right call

None of this makes free data a poor choice. It makes it the right choice for a narrower set of jobs than the search volume around it suggests.

Prototyping is the clearest case. Before you know whether an architecture or approach works at all, spending a procurement cycle on licensed data is premature. LibriSpeech or Common Voice will show whether the model trains before you spend a cent.

Benchmarking is another. LibriSpeech's test-clean and test-other sets are the standard yardstick for English ASR word error rate precisely because they are free, fixed, and widely reported against, which makes your numbers comparable to everyone else's.

Augmentation is the third, and the one teams underuse. Blending a slice of Common Voice or LibriSpeech into a larger licensed set can add acoustic diversity, different microphones, rooms, and accents, more cheaply than commissioning more of the same condition from scratch.

All three are legitimate engineering practice, not a workaround. Commercial usability depends on the exact grant and facts; internal or research use is not automatically permitted, and shipping does not by itself determine every licence condition.

When to compare another source

When a model moves beyond internal testing, verify more than downloadability: the exact release and version, publisher and source chain, license terms, source and platform permissions, applicable privacy basis and notices, attribution or share-alike duties, coverage, and deployment fit. Open, licensed, and custom sources can each fit when their evidence supports the intended use.

Coverage differs by language and release. Cite dated statistics and test usable hours, validation status, speaker mix, dialects, and acoustics against the deployment before sourcing the documented gap. The buying AI training data guide provides the comparison framework.

A public corpus can be appropriate for prototyping, evaluation, training, or augmentation when its exact terms and measured fit support the project. A paid license or collection does not establish rights, provenance, indemnity, or production depth automatically; verify the specific contract and delivery.