Almost nobody in this category publishes a price, so the most common question about speech training data is also the least answered. This page gives the arithmetic: what drives the per-hour rate, what Spirelight charges, five worked examples at real catalogue prices, and how to tell whether two quotes are even measuring the same product.

Why nobody publishes a price

Search for what speech training data costs and you will find quote forms. The stated reason is that every collection is bespoke, which is true but not the whole story: an opaque price is also easier to defend, and a buyer who cannot compare is a buyer who negotiates from behind. The practical effect is that teams cannot budget a project without first entering a sales process, so a lot of reasonable projects never get scoped at all.

Speech data is not actually that hard to price. It is sold per audio hour, and a small number of factors move that rate in predictable directions. What follows is how the arithmetic works, and what Spirelight charges, which is published on every dataset page.

What the market publishes for speech data collection

Most speech data collection is sold through quote forms. You describe the project, a salesperson calls, and the price arrives attached to a pitch. The public numbers that do exist are few enough to fit in one table, and they are worth knowing before you enter that process, because they anchor what a reasonable quote looks like.

LXT is the notable vendor that publishes collection pricing: $15,000 for focused domain datasets of 50 to 200 hours, up to $150,000 or more for multi-accent collections past 1,000 hours. Entry pricing at the large Asian data vendors is published around $20,000 per purchase. Dataset marketplaces list raw uncleaned US call-center audio from about $25 per hour, but read that row carefully: no consent documentation, no transcripts, no collection at all. It is a different product, and we come back to why that matters.

SourcePublished priceWhat it covers
LXT, focused domain dataset$15,00050 to 200 hours, collected to spec
LXT, large multi-accent collection$150,000+1,000+ hours across accent cells
Large Asian data vendorsAround $20,000Published entry pricing per purchase
Dataset marketplacesFrom about $25 per hourRaw uncleaned US call-center audio, no consent documentation, no transcripts
Most other vendorsQuote onlyPrice revealed after a sales process
Spirelight, worked exampleAbout $9,00050-hour scripted single-language studio collection with time-coded transcripts

That is close to the entire public record. Everyone else prices by quote, which means the burden of sanity-checking a number falls on you. The rest of this page is the checklist for doing that.

What moves the price per hour

  • Language availability. The single biggest driver, and it is about recruiting rather than recording. Finding fifty verified native speakers of a widely spoken language is routine; finding fifty of a low-resource language, in-region, willing to be recorded, is the whole job. Rare languages carry a real premium and any vendor quoting the same rate across all languages is averaging somewhere.
  • Order size. Contracting, delivery setup, and quality sampling cost roughly the same whether you buy ten hours or a thousand, so the per-hour rate falls as volume rises. This is why minimums exist, and why a small order legitimately costs more per hour rather than being marked up out of spite.
  • Recording conditions. Remote recording on a contributor's own device is the cheapest path. Studio capture, far-field microphone arrays, in-car rigs, and telephony re-recording each add setup cost and reduce how many usable hours a session produces.
  • Transcription depth. An untranscribed audio hour is a fraction of the cost of one with verified verbatim transcripts, timestamps, and speaker labels. Quotes that omit transcription are not comparable to quotes that include it, and this is the most common way two prices look different when they are measuring different products.
  • License breadth. Non-exclusive commercial use is the standard. Exclusivity, resale rights, or a guarantee that the data will never be licensed to a competitor multiply the price, because the vendor is giving up all future revenue on that asset.

What it costs at Spirelight

Every dataset page publishes its own per-hour rate. Across the catalogue those rates run $60 to $95 per audio hour at full scale, for spontaneous conversational speech delivered with verified transcripts and metadata under a commercial license. The rate rises as the order gets smaller:

  • 100 hours and above: $60 to $95 per hour. This is the list rate.
  • 25 to 99 hours: $80 to $130 per hour.
  • 10 to 24 hours: $100 to $155 per hour. 10 hours is the minimum order.

For off-the-shelf orders, the volume tiers and worked examples for 10, 25, and 100 hour purchases are in the small speech datasets guide, and every dataset page on the catalogue shows its own per-hour rate.

Worked examples

Each figure below is the real catalogue price, given as a range because the rate depends on the language. Everything includes verified transcripts, speaker metadata, and a commercial license.

  1. Evaluation set, 10 hours: $1,000 to $1,550. Enough to benchmark an existing model on a dialect and find out whether you have a problem worth spending on. This is the cheapest useful purchase in speech data and usually the right first one.
  2. First fine-tune, 25 hours: $2,000 to $3,250. Twenty hours to adapt on and five held out for evaluation from the same recording conditions, which is what makes a before-and-after comparison honest.
  3. Domain adaptation, 100 hours: $6,000 to $9,500. Enough speaker diversity that the model does not overfit to a handful of voices, at the full-scale per-hour rate.
  4. Production corpus, 500 hours: $30,000 to $47,500. A serious single-language asset, the size at which most teams stop buying and start thinking about whether to collect.
  5. Large multi-language build, 2,000 hours: $120,000 to $190,000. At this scale the conversation is usually about collection schedules and exclusivity rather than catalogue purchase.

Custom collection, where audio is recorded to your specification rather than licensed from an existing set, is priced on the same per-hour basis with the recording conditions factored in. If a language you need is already in the catalogue, licensing hours from it is faster and cheaper than commissioning new recording.

For custom collection, a worked example. A 50-hour scripted single-language studio collection, delivered with time-coded transcripts, runs about $9,000. The comparable published rate for a focused domain dataset in that range is $15,000, which puts the Spirelight price 40% below it. That gap is typical rather than cherry-picked: comparable collections usually land 35 to 40 percent below quoted competitor rates. What a collection engagement includes end to end, from spec to delivery, is covered in the data collection provider checklist.

Why the discount is structural, not promotional

A 35 to 40 percent gap against quoted rates should make you suspicious. Discounts that size are usually a loss leader, a quality cut, or an introductory rate that expires. This one is none of those. It comes from three middlemen that most vendors bill through and we do not.

  • Owned studio locations. Most vendors rent recording facilities per project and pass the rental through with a margin on top. Spirelight records in its own studios, so there is no facility line on the quote. How studio and on-site recording work in practice is its own guide.
  • Automated recruiting into a standing crowd. Recruiting through an agency adds a per-participant fee and weeks of billed lead time. A standing contributor crowd with automated matching removes both: speakers who fit the demographic cells are already registered and vetted before the project starts.
  • Moderators on staff. Session moderators hired as daily contractors carry a markup on every recording day. Staff moderators cost what they cost, with no contractor layer.

Each removed middleman is a percentage, and percentages compound. That is the whole explanation. The output spec, the consent documentation, and the QA process are the same as in a full-price collection, because none of the removed costs were quality costs.

Compare like for like or not at all

The $25 per hour marketplace audio and a collected, consented, transcribed dataset are not the same product at two prices. They are different products at different risk levels. Raw scraped or resold recordings ship with no consent documentation, no transcripts, and no answer when a regulator, a customer, or an acquirer asks where the training data came from.

Consented, documentation-backed audio costs more per hour precisely because most of the cost sits in the consent, the transcription, and the records. Dividing both price tags by hours and comparing the results is meaningless arithmetic. And the risk side is not hypothetical: if you deploy in Europe, provenance and documentation are compliance inputs, as the EU AI Act speech data guide lays out.

How to compare two quotes

Most apparent price differences are product differences. Before comparing numbers, pin down five things for each quote so you are measuring the same thing:

  1. Is transcription included, and verified by whom? Machine transcription with no human pass is a different product at a different price.
  2. Is it conversational or read speech? Spontaneous dialogue costs more to collect and is worth more for most production use cases. Scripted prompt data is cheaper for a reason.
  3. What license, and can you keep the model? Confirm in writing that models trained on the data remain yours, especially at small volumes where some vendors restrict this.
  4. Is consent documented per speaker? Data without a provenance record is cheaper to produce and considerably more expensive to own once you have to account for what your model was trained on.
  5. What is the minimum order? A low per-hour rate attached to a 500 hour minimum is not a low price if you need twenty hours.

Where the money is usually wasted

In our experience the expensive mistake is rarely the rate. It is buying the wrong hours: a large corpus of clean read speech when the deployment is noisy phone calls, or broad language coverage when the failure is concentrated in one accent. Both are avoidable by buying a small evaluation set first and measuring before committing. The measurement costs four figures and routinely changes what the five or six figure order should have been.