Speech and Audio Data Collection Services for AI

Spirelight scopes custom speech and audio data collection for AI teams, on-site in controlled environments or remotely through vetted contributors. The project brief defines target languages, speaker profiles, recording conditions, consent and rights, annotation, QA, delivery artifacts, and acceptance criteria. Collection pages show planning rates where available; the written proposal confirms feasibility and final price.

Buyer answer

What a custom collection scope confirms.

Spirelight scopes commissioned speech and audio collection, not resale of an existing corpus. Before recording starts, the written scope confirms the speaker matrix and recruitment feasibility, the capture protocol for the production channel, consent and rights coverage for the intended model use, QA and acceptance rules, and the delivery artifacts and pricing basis for the specific brief.

Scope fieldWhat the written proposal should confirm
Speaker matrix and recruitmentLanguage, dialect, accent, demographic and device quotas, recruitment method, verification, and confirmed feasibility for each requested locale.
Capture protocol and channel realismScripted or spontaneous design, prompts, equipment, environments, noise conditions, and how the recording setup reproduces the production channel.
Consent, rights, and provenanceConsent wording, permitted uses including model training and evaluation, retention, transfers, exclusivity, and the provenance records that link each recording to its permission.
QA and acceptanceAutomated and human checks, sampling method, rejection rules, the acceptance test the buyer runs, and what triggers rework or replacement.
Delivery, schedule, and pricing basisAudio specs, transcripts, metadata schema, manifests, batch cadence, any pilot stage, schedule assumptions, and the pricing unit behind the quote.

Comparing several vendors against the same brief? Use the free speech data provider scorecard to score the evidence, and the dataset specification worksheet to freeze the brief itself.

What we collect

From scripted prompts to real-world, in-the-wild audio.

A custom collection can cover scripted, spontaneous, wake-word, far-field, in-car, and other defined scenarios. The written scope confirms speaker criteria, recording conditions, feasibility, rights, QA, and acceptance requirements.

Scripted speech

Read prompts, sentences, and command sets for controlled coverage of phonemes, vocabulary, and domain terms, recorded to a consistent studio or device spec.

Spontaneous and conversational

Natural monologue, dialogue, and multi-party conversation for models that need real, unscripted speech with disfluencies, overlaps, and turn-taking.

Wake-word

Targeted wake-word and short-command capture with speaker, distance, device, positive-example, and negative-example requirements defined in the brief.

Far-field

Distant-microphone and room-scale recording for smart speakers and voice interfaces that must work across a room, not just into a handset.

In-car

Cabin and in-vehicle audio can be scoped with road, engine, HVAC, device, and passenger-speech conditions relevant to the intended evaluation.

Custom scenarios

Have a scenario of your own? The brief can define custom prompts, environments, devices, staffing, security, and acceptance criteria, subject to feasibility.

By model type

Collection for ASR, TTS, and voice AI.

The right recording design depends on what you are training. Each project defines its own speaker, prompt, device, environment, rights, QA, and delivery requirements.

ASR

For speech recognition, a brief can prioritize speaker and accent diversity, realistic noise, vocabulary coverage, and evaluation slices relevant to the intended deployment.

TTS

For text to speech, a collection can be scoped for studio conditions, selected voice profiles, tone and pacing requirements, and an agreed high-sample-rate format.

Voice AI and assistants

For assistants and voice interfaces, a scope can include wake-word, command, and conversational data across agreed devices and environments.

Collection and audio and speech annotation services can be combined in one project when the scope defines both recording and labeling requirements.

On-site collection

On-site and controlled-environment recording, run as an operation.

Some briefs cannot be collected remotely: the channel, the acoustics, or the verification requirements demand a controlled environment with a moderator in the room. Spirelight scopes on-site collection around the deployment channel, with demographic quotas, recruitment, rig, and session protocol confirmed in the written proposal.

Treated rooms and studios

Repeatable acoustics for scripted and read capture, with fixed microphone positions and per-session room-tone baselines. Vocal-booth conditions where TTS-grade recording requires them.

Far-field and device rooms

Furnished rooms with controlled noise sources and microphone arrays at device-realistic positions, for assistants and devices that listen across a room rather than into a handset.

In-car and cabin capture

Microphone positions fixed to the production locations in the cabin, recorded parked and on defined drive profiles, with HVAC, speed, and window states logged as metadata.

Meetings and multi-party

Table-top arrays with a close-talk reference channel per participant, and session tasks designed to elicit the natural overlap and turn-taking meeting models need.

Telephony loop

Audio routed through a real phone channel when production listens at 8 kHz, so the training data carries the codec and bandwidth the model will actually hear.

Moderated sessions

A trained moderator briefs the speaker, verifies identity and permissions at the door, listens live, re-records failed items on the spot, and logs every deviation for QA.

The method is public. Our on-site speech data collection playbook covers demographic quota design, recording setups by industry, session failure modes, and cost drivers, and the participant recruitment guide covers how the right speakers end up in the chair.

Network and consent

Recruitment, consent, and rights defined before recording.

A collection brief should name the target speakers, permitted uses, consent and provenance records, retention, transfers, and acceptance criteria. Spirelight confirms those requirements in the written project scope.

Targeted recruitment

Recruitment criteria can include language, dialect, age, gender, and location. The agreed speaker matrix and feasibility are documented before launch.

Explicit consent

The project scope defines the consent language and the provenance records required to connect recordings with the applicable permission.

Clear licensing

Permitted uses, exclusivity, retention, model-related rights, delivery records, and other license terms are stated in the project contract.

Languages and dialects

Language and dialect recruitment built around your target market.

Where native-speaker recruitment is required, the written brief defines language, dialect, regional variety, speaker profile, and verification method. Feasibility is confirmed before commitment.

Browse the language-specific collection configurations for planning rates and capacity, then confirm feasibility, sample status, rights, schedule, and final price against your brief.

Formats and delivery

Delivered as structured training data, not a pile of files.

The written scope defines the schema and the audio, transcript, metadata, manifest, and handoff artifacts included in delivery.

Audio specs

Audio format, sample rate, bit depth, channel layout, and any conversion rules are agreed for the capture method and target pipeline.

Metadata and transcripts

Speaker, device, condition, transcript, and timestamp fields can be delivered in JSON, CSV, or another agreed schema.

Handoff

The written handoff plan defines delivery channel, batch cadence, checksums, identifiers, and the provenance records required for the project.

Why Spirelight

A speech specialist, not a generic data collection company.

Spirelight focuses this service on speech and voice. Buyers should compare the named recruitment method, recording protocol, QA evidence, rights, and acceptance test for their own brief.

Speech-only focus

Voice data collection is our whole business, so our prompts, rigs, and QA are built for audio quality rather than adapted from a generic labeling tool.

Consent requirements in scope

The project contract defines required consent language, permitted uses, license terms, retention, and provenance records. Buyers should verify those artifacts and review legal risk for their intended use.

In-production QA

Where in-production review is included, the QA plan defines which audio and metadata batches are sampled, what is checked, and what triggers escalation or rework.

See how collection fits alongside transcription, annotation, and QA on our services overview. Still scoping a project? Our guide to custom speech data collection covers scripted vs. spontaneous recording, a brief template, speaker recruitment, and how in-production QA can be scoped.

Request a quote

Tell us what you need to record and we will design it.

Send us your languages, speaker profiles, audio volume, recording conditions, and intended use. We will assess feasibility and respond with the information needed to scope a written proposal.

Data collection FAQ

What is speech and audio data collection for AI?

Speech and audio data collection is the process of recording real human voices under controlled conditions to build training data for AI models. It covers recruiting the right speakers, capturing scripted or spontaneous audio to a defined spec, obtaining consent, and delivering the recordings with metadata and transcripts. The result is a dataset your team can train ASR, TTS, or voice AI models on.

How is a custom collection scoped and priced?

A collection is scoped around speakers, languages, audio volume, recording conditions, annotation, QA, rights, and delivery artifacts. Published page figures are planning inputs; the written proposal confirms the pricing unit, assumptions, project minimum, schedule, and final price.

How do you scope language coverage?

We scope recruitment to the language, dialect, regional accent, and speaker profile in your project brief. We confirm recruitment feasibility for the requested locale before the collection is agreed.

How do you handle consent and licensing?

The written project scope defines the required consent language, permitted uses, provenance records, retention, transfers, and license terms before collection begins. The applicable legal basis, contracting roles, and GDPR requirements are assessed for the specific project.

What is the typical timeline?

Timing depends on recruitment, language, recording setup, volume, QA, and acceptance requirements. The written proposal states any pilot, batch schedule, review points, and delivery dates before launch.

What does the dataset catalogue contain?

The catalogue presents language-specific custom collection configurations and planning inputs. It is not a promise of finished inventory. Sample status, recruitment feasibility, recording conditions, rights, schedule, and final price are confirmed against the buyer's brief.

Can you run on-site speech data collection in a controlled environment?

Yes. On-site collection is scoped as moderated sessions in treated rooms, far-field device rooms, vehicle cabins, meeting-style rooms, or through a telephony loop, with the rig matched to the production channel. The written scope confirms location, staffing, equipment, demographic quotas, recruitment, consent records, QA, and schedule for the specific project. The full method is in our on-site collection playbook.

What does an AI data collection service actually do?

It produces training data that does not exist yet, rather than reselling data that does. The work is a chain: defining the spec, recruiting speakers who match it, capturing audio under controlled conditions, obtaining and recording consent, transcribing and annotating, then running quality control before delivery. Vendors differ in how much of that chain they own. Ask which steps are done in-house and which are subcontracted, because handoffs can create gaps in consent documentation and QA accountability.

How much do AI data collection services cost?

Collection cost depends on language and speaker recruitment, recording setup, volume, annotation, QA, rights, and schedule. Spirelight publishes planning rates and project minimums where available; the written proposal confirms what is included and the final price.

How do I choose an AI data collection company?

Compare named recruitment methods, the consent and rights required for your use, recording and annotation protocols, QA evidence, acceptance criteria, delivery artifacts, and applicable case studies. Ask whether a sample matched to your brief is currently available, and compare quotes against the same specification.

Can you collect speech data for a small project?

A small pilot can be considered when it can answer a defined evaluation question. Feasibility, the project minimum, sample status, deliverables, rights, and price are confirmed for the requested configuration before work begins.

How should buyers assess GDPR and EU AI Act requirements for speech data?

Voice recordings can be personal data, and biometric processing may trigger additional requirements. The buyer and provider should document their roles, lawful basis, purpose, consent where applicable, retention, transfers, security, provenance, and required AI Act records for the specific project. Denmark-based operations do not create blanket compliance, so the buyer should review the written scope with counsel.

Ready to start? Request a quote or add audio and speech annotation to your project.