Speech-data quotes are hard to compare because language, speaker criteria, recording conditions, transcripts, QA, rights, and acceptance rules vary. This guide gives buyers a practical cost model, five scope scenarios, and a like-for-like worksheet. Any displayed planning figure is only useful when its assumptions and evidence are stated; a written quote confirms the deliverables and final price.
Why speech-data prices are hard to compare
Many custom collections are quoted from a project specification. That is reasonable when language, speaker cells, recording conditions, transcripts, QA, rights, schedule, and acceptance criteria differ, but it makes headline rates difficult to compare.
Audio hours are one pricing unit, not a complete specification. Ask every supplier to itemize the same inputs and to define a usable delivered hour. Treat a public planning figure as indicative only when its assumptions, date, and evidence are stated; the written quote controls.
Build a comparable market view
Public prices age quickly and often describe different products. Record the source URL, publication date, currency, tax treatment, minimum order, usable-hour definition, and included deliverables for every figure. If any of those fields is missing, use the number only as a lead for a current written quote.
| Quote field | What to record | Why it matters |
|---|---|---|
| Collection status | Finished inventory, custom collection, or feasibility target | Changes availability and schedule |
| Audio scope | Language, speaker cells, speech type, recording conditions, usable hours | Defines what is being measured |
| Annotation and QA | Transcript type, labels, reviewer coverage, acceptance thresholds | Separates audio-only from verified deliverables |
| Rights and records | Source, permissions, permitted uses, privacy records, retention, restrictions | Determines whether the data fits the intended use |
| Commercial terms | Minimum, setup fees, volume tiers, currency, taxes, schedule, payment terms | Turns a headline rate into a comparable total |
Send the same specification to each supplier and compare written assumptions. A marketplace listing, a finished corpus license, and a custom collection are different procurement routes; evaluate each on fit and documented evidence rather than category-level assumptions.
What moves the price per hour
- Language availability. Recruitment effort can vary by language, dialect, location, eligibility criteria, and target speaker count. Ask for the feasibility assumptions and any recruiting contingency in the quote.
- Order size. Fixed setup work and volume assumptions can change the effective rate. Compare the minimum order, setup fees, volume tiers, and usable-hour definition rather than assuming a universal volume discount.
- Recording conditions. Remote, studio, far-field, in-car, and telephony setups carry different equipment, supervision, privacy, and rejection-rate assumptions. Ask suppliers to state the recording method and which setup costs are included.
- Transcription depth. Transcript type, timestamps, speaker labels, event tags, reviewer coverage, and acceptance thresholds all affect scope. A quote for audio only is not comparable with one that includes verified annotations.
- License breadth. Permitted uses, territory, term, redistribution, model-related rights, retention, and exclusivity are contract-specific. Compare the same rights package and ask the supplier to price any requested restrictions separately.
How Spirelight scopes collection pricing
Spirelight prices a custom speech-data collection from the confirmed specification, not from a universal shelf catalogue. Language, speaker profile, recording environment, consent and usage rights, annotation, QA, volume, and schedule all affect the written quote.
- Start with the intended model and evaluation target. That determines which speakers, conditions, labels, and holdouts belong in scope.
- Use published page figures only as planning inputs. They are not a promise of current availability or a final quote.
- Confirm the project minimum and volume band. These vary with the configuration and are documented in the proposal.
For custom collection planning, the small speech datasets guide explains pilot sizes. Collection configuration pages describe possible targets and may show evidence-gated planning inputs; they are not proof of inventory, availability, or final price.
Five scope scenarios for budgeting
These are scope scenarios, not quoted prices or guaranteed outcomes. Obtain current like-for-like quotes for the specified language, speakers, conditions, deliverables, rights, and schedule.
- 10-hour evaluation scenario. A narrow held-out benchmark for a defined dialect or condition. Confirm representativeness and the decision threshold.
- 25-hour adaptation pilot. Reserve a separately held-out split and define what improvement would justify expansion.
- 100-hour single-domain scenario. Broaden speaker and acoustic cells only after the pilot identifies where coverage is missing.
- 500-hour production scenario. Require a cell-level allocation, staged acceptance, delivery schedule, and change-control process.
- 2,000-hour multilingual scenario. Compare country-by-country feasibility, reviewer capacity, rights, schedule, and dependencies rather than applying one blended rate blindly.
Custom collection and licensing are different routes. Compare them against the same target conditions, usable-hour definition, annotations, rights, QA, and schedule. A collection configuration is not proof of finished inventory; confirm whether applicable data or a custom collection is available before comparing routes.
A 50-hour scripted, single-language studio scenario with time-coded transcripts still requires a written quote. Do not extrapolate a vendor-wide price or discount from any worked example. Match language, speaker cells, conditions, transcripts, metadata, QA, rights, acceptance criteria, and schedule line by line. The data collection provider checklist covers what to include end to end.
Why like-for-like quotes matter
A large apparent price gap should make you inspect the scope. Recording location, recruiting, moderation, transcription, QA, rights, and acceptance criteria can turn two headline prices into different products.
- Recording location. A quote may include owned, partner, or rented facilities. Ask which site is proposed, who operates it, and whether equipment, access, supervision, insurance, and facility charges are included. How studio and on-site recording work in practice is its own guide.
- Recruitment method. Recruitment may be direct or subcontracted. Ask who sources each speaker cell, how eligibility is verified, which fees and lead-time assumptions are included, and who remains responsible for acceptance.
- Moderation model. Moderators may be employees or contractors. Ask who moderates each session, how qualifications and responsibility are assigned, and whether supervision and review costs are included.
Each operating choice can affect price. Compare written deliverables, responsibilities, evidence, and acceptance criteria rather than a claimed percentage discount.
Compare procurement routes like for like
A marketplace listing, a licensed corpus, and a commissioned collection may include different audio conditions, annotations, source records, rights, and QA. Do not infer those fields from the supplier category or price. Verify them for the specific asset and intended use. Before requesting quotes, compare AI training data companies across these routes so the shortlist matches the product you actually need.
For European deployments, provenance and documentation can be relevant compliance inputs depending on the system, role, and data use. The EU AI Act speech data guide explains the project-specific assessment and the records to request.
How to compare two quotes
Most apparent price differences are product differences. Before comparing numbers, pin down five things for each quote so you are measuring the same thing:
- Is transcription included, and verified by whom? Machine transcription with no human pass is a different product at a different price.
- Is it conversational or read speech? Match speech type to the deployment task, and ask how prompting, moderation, and usable-hour acceptance affect the quote.
- What rights apply? Confirm permitted uses, model-related rights, retention, transfer, territory, term, and restrictions in the contract.
- What provenance and data-protection records exist? Check source, permissions, license scope, applicable privacy notices and lawful basis, and consent records when consent is relied on.
- What is the minimum order? A low per-hour rate attached to a 500 hour minimum is not a low price if you need twenty hours.
Where the money is usually wasted
A common budget risk is buying hours that do not match deployment: clean read speech for noisy calls, or broad language coverage when the measured gap is one accent. A representative evaluation pilot can test the need before a larger commitment; compare its quoted cost with the decision value and define the success threshold first.