Choosing an AI data collection provider means checking how the work is actually run, not comparing identical-looking capability pages. For speech, the useful evidence covers speaker sourcing, consent, recording controls, transcription rules, QA gates, and the chain linking each delivered file to its rights.

This provider-evaluation guide gives you a practical checklist and explains how to compare quote mechanics on a like-for-like basis. The commercial overview and quote destination live on our speech data collection service page; this page is for evaluating any provider, Spirelight included.

What to evaluate in a speech data collection provider

A data collection service does the work that happens before a model ever sees a training example: it finds the right people, captures data from them under defined conditions, and turns the raw capture into something a training pipeline can consume. For speech and audio, the delivery is rarely just audio files. A serious provider ships transcripts, speaker metadata, QA reports, and consent documentation as parts of the same project.

The category runs wider than speech. Vendors in this space also collect images, video, text, and sensor data. This guide sticks to speech and audio, because that is our work, and because speech collection has operational problems, recruiting native speakers, running recording sessions, checking transcripts, that generalist services routinely underestimate.

The three collection formats

Almost every speech collection project uses one of three formats, and the format drives cost more than any other single decision.

Scripted recording puts a fixed prompt in front of a speaker: wake words, voice commands, digit strings, read sentences. It is the cheapest format because sessions are predictable and transcription is nearly free, since the script is the transcript. It is the right choice for wake word models, command grammars, and pronunciation coverage.

Prompted tasks give speakers a scenario instead of a script: order a pizza, ask for directions, dispute a charge. Speakers phrase the task their own way, so you get natural variation inside a bounded domain. Transcription becomes real work here, which raises the cost per hour.

Spontaneous conversation is two people actually talking. It is the hardest format to run and the most valuable to own, because it contains the disfluencies, interruptions, and overlapping speech that production ASR fails on. If your model transcribes real conversations, some portion of your training data has to look like this. Our guide to custom speech data collection walks through scoping a project format by format.

In-studio, on-site, or remote

The second decision is where the recording happens. In-studio collection gives you controlled acoustics, calibrated microphones, and a moderator in the room who catches problems while the speaker is still in the chair. On-site collection brings the rig to the environment you are modeling, a car cabin, a factory floor, a clinic, so the noise in the data is the noise your product will hear. Remote collection uses participants' own phones and headsets, which trades acoustic control for device diversity and geographic reach.

None of the three is better in general. Wake word and command data usually wants studio control. Environment-specific models want on-site capture. Broad ASR robustness often benefits from the messiness of remote devices. We cover the studio and field methodology in detail in our guide to on-site speech data recording.

How to evaluate a provider

Capability pages all read the same, so evaluate on mechanics. Five checks separate providers quickly.

  • Consent and license chain. Ask to see the consent form participants actually sign and how each recording links back to a signed consent. If the answer is a summary instead of a document, the chain probably does not exist in auditable form.
  • QA gates during collection, not after. A provider that checks audio while sessions are running can re-record a failed batch the same week. A provider that QAs after delivery hands you the dispute instead of the fix. Ask when the first quality gate fires.
  • Named sourcing for languages and accents. A claim of 100 languages is a claim about a spreadsheet. Ask where the speakers for your specific languages come from: an owned contributor pool, a partner agency, or a crowd platform the vendor resells.
  • Turnaround mechanics. Not the promised date, the process: batch sizes, what triggers a re-record, and who absorbs the cost when a batch fails QA.
  • Pricing transparency. Most providers quote-wall everything. Published pricing is worth reading as a signal of operational confidence: a vendor that knows its unit economics can afford to print them.

What a collection quote is actually made of

Every collection quote, whatever the line items say, decomposes into five parts: recruiting, facility, moderation, QA, and margin. Recruiting pays to find and schedule the right speakers. Facility pays for rooms and equipment. Moderation pays the people who run sessions. QA pays for the checking and re-recording loop. Margin is what remains.

How a provider sources each part sets its floor price. A vendor that rents studios passes the rental plus a markup into your quote. A vendor that recruits through agencies pays a per-participant fee and bills you the lead time. A vendor that staffs moderation per project pays contractor premiums and passes those through too.

We built Spirelight to strip those layers out. We own our studio locations, so no facility rental margin enters the quote. Recruiting runs automated into a standing multilingual contributor crowd, so there is no per-participant agency fee and no recruitment lead time billed. Moderators are on staff, not hired per project. That is why comparable collections typically land 35 to 40 percent below the rates competitors quote: the work is the same, the cost structure is not.

A worked example with published prices

Published prices are rare in this market, which makes the few that exist useful anchors. LXT, one of the few competitors that publishes prices, lists $15,000 for focused domain datasets of 50 to 200 hours and $150,000 and up for multi-accent collections of 1,000 hours or more.

CollectionPriceSource
Focused domain dataset, 50 to 200 hours$15,000LXT published rate
Multi-accent collection, 1,000+ hours$150,000 and upLXT published rate
50-hour scripted single-language studio collection, time-coded transcripts includedAbout $9,000Spirelight, roughly 40% less

The caveat is like-for-like scope. An hour of scripted studio audio, an hour of transcribed spontaneous conversation, and an hour of raw remote capture are different products at different prices, so compare specs, not headline numbers. Our guide to speech data collection cost breaks down what actually moves the per-hour rate.

Check the shelf before you commission

Custom collection is the expensive route, and often not the necessary one. Our dataset catalogue covers roughly 60 languages with per-hour rates published on every page and a minimum order of 10 hours, so pricing a licensed alternative takes minutes, not a sales cycle. If a catalogue dataset covers part of your spec, license that part and commission only the remainder.

The economics of small licensed orders, and what the entry tiers actually buy, are covered in our guide to small speech datasets.

What a Spirelight delivery includes

We run collection in about 60 languages, from scripted studio sessions to spontaneous conversation, with transcription and QA handled in-house as part of the same project rather than subcontracted afterward. Collection is run from the EU under GDPR from the first consent form, not retrofitted for compliance at delivery.

A delivery brief should state which consent, speaker metadata, transcription, and QA records are in scope. Do not assume a universal documentation bundle: name the artifacts your legal and model teams need in the quote, then verify them before delivery. Our guide to the EU AI Act and speech data provides the checklist.