AI Data Collection Services for Speech and Audio

Spirelight scopes custom speech and audio data collection for AI teams. The project brief defines target languages, speaker profiles, recording conditions, consent and rights, annotation, QA, delivery artifacts, and acceptance criteria. Collection pages show planning rates where available; the written proposal confirms feasibility and final price.

What we collect

From scripted prompts to real-world, in-the-wild audio.

A custom collection can cover scripted, spontaneous, wake-word, far-field, in-car, and other defined scenarios. The written scope confirms speaker criteria, recording conditions, feasibility, rights, QA, and acceptance requirements.

Scripted speech

Read prompts, sentences, and command sets for controlled coverage of phonemes, vocabulary, and domain terms, recorded to a consistent studio or device spec.

Spontaneous and conversational

Natural monologue, dialogue, and multi-party conversation for models that need real, unscripted speech with disfluencies, overlaps, and turn-taking.

Wake-word

Targeted wake-word and short-command capture with speaker, distance, device, positive-example, and negative-example requirements defined in the brief.

Far-field

Distant-microphone and room-scale recording for smart speakers and voice interfaces that must work across a room, not just into a handset.

In-car

Cabin and in-vehicle audio can be scoped with road, engine, HVAC, device, and passenger-speech conditions relevant to the intended evaluation.

Custom scenarios

Have a scenario of your own? The brief can define custom prompts, environments, devices, staffing, security, and acceptance criteria, subject to feasibility.

By model type

Collection for ASR, TTS, and voice AI.

The right recording design depends on what you are training. Each project defines its own speaker, prompt, device, environment, rights, QA, and delivery requirements.

ASR

For speech recognition, a brief can prioritize speaker and accent diversity, realistic noise, vocabulary coverage, and evaluation slices relevant to the intended deployment.

TTS

For text to speech, a collection can be scoped for studio conditions, selected voice profiles, tone and pacing requirements, and an agreed high-sample-rate format.

Voice AI and assistants

For assistants and voice interfaces, a scope can include wake-word, command, and conversational data across agreed devices and environments.

Collection and audio and speech annotation services can be combined in one project when the scope defines both recording and labeling requirements.

Network and consent

Recruitment, consent, and rights defined before recording.

A collection brief should name the target speakers, permitted uses, consent and provenance records, retention, transfers, and acceptance criteria. Spirelight confirms those requirements in the written project scope.

Targeted recruitment

Recruitment criteria can include language, dialect, age, gender, and location. The agreed speaker matrix and feasibility are documented before launch.

Explicit consent

The project scope defines the consent language and the provenance records required to connect recordings with the applicable permission.

Clear licensing

Permitted uses, exclusivity, retention, model-related rights, delivery records, and other license terms are stated in the project contract.

Languages and dialects

Language and dialect recruitment built around your target market.

Where native-speaker recruitment is required, the written brief defines language, dialect, regional variety, speaker profile, and verification method. Feasibility is confirmed before commitment.

Browse the language-specific collection configurations for planning rates and capacity, then confirm feasibility, sample status, rights, schedule, and final price against your brief.

Formats and delivery

Delivered as structured training data, not a pile of files.

The written scope defines the schema and the audio, transcript, metadata, manifest, and handoff artifacts included in delivery.

Audio specs

Audio format, sample rate, bit depth, channel layout, and any conversion rules are agreed for the capture method and target pipeline.

Metadata and transcripts

Speaker, device, condition, transcript, and timestamp fields can be delivered in JSON, CSV, or another agreed schema.

Handoff

The written handoff plan defines delivery channel, batch cadence, checksums, identifiers, and the provenance records required for the project.

Why Spirelight

A speech specialist, not a generic data collection company.

Spirelight focuses this service on speech and voice. Buyers should compare the named recruitment method, recording protocol, QA evidence, rights, and acceptance test for their own brief.

Speech-only focus

Voice data collection is our whole business, so our prompts, rigs, and QA are built for audio quality rather than adapted from a generic labeling tool.

Consent requirements in scope

The project contract defines required consent language, permitted uses, license terms, retention, and provenance records. Buyers should verify those artifacts and review legal risk for their intended use.

In-production QA

Where in-production review is included, the QA plan defines which audio and metadata batches are sampled, what is checked, and what triggers escalation or rework.

See how collection fits alongside transcription, annotation, and QA on our services overview. Still scoping a project? Our guide to custom speech data collection covers scripted vs. spontaneous recording, a brief template, speaker recruitment, and how in-production QA can be scoped.

Request a quote

Tell us what you need to record and we will design it.

Send us your languages, speaker profiles, audio volume, recording conditions, and intended use. We will assess feasibility and respond with the information needed to scope a written proposal.

Data collection FAQ

What is speech and audio data collection for AI?

Speech and audio data collection is the process of recording real human voices under controlled conditions to build training data for AI models. It covers recruiting the right speakers, capturing scripted or spontaneous audio to a defined spec, obtaining consent, and delivering the recordings with metadata and transcripts. The result is a dataset your team can train ASR, TTS, or voice AI models on.

How is a custom collection scoped and priced?

A collection is scoped around speakers, languages, audio volume, recording conditions, annotation, QA, rights, and delivery artifacts. Published page figures are planning inputs; the written proposal confirms the pricing unit, assumptions, project minimum, schedule, and final price.

How do you scope language coverage?

We scope recruitment to the language, dialect, regional accent, and speaker profile in your project brief. We confirm recruitment feasibility for the requested locale before the collection is agreed.

How do you handle consent and licensing?

The written project scope defines the required consent language, permitted uses, provenance records, retention, transfers, and license terms before collection begins. The applicable legal basis, contracting roles, and GDPR requirements are assessed for the specific project.

What is the typical timeline?

Timing depends on recruitment, language, recording setup, volume, QA, and acceptance requirements. The written proposal states any pilot, batch schedule, review points, and delivery dates before launch.

What does the dataset catalogue contain?

The catalogue presents language-specific custom collection configurations and planning inputs. It is not a promise of finished inventory. Sample status, recruitment feasibility, recording conditions, rights, schedule, and final price are confirmed against the buyer's brief.

What does an AI data collection service actually do?

It produces training data that does not exist yet, rather than reselling data that does. The work is a chain: defining the spec, recruiting speakers who match it, capturing audio under controlled conditions, obtaining and recording consent, transcribing and annotating, then running quality control before delivery. Vendors differ in how much of that chain they own. Ask which steps are done in-house and which are subcontracted, because handoffs can create gaps in consent documentation and QA accountability.

How much do AI data collection services cost?

Collection cost depends on language and speaker recruitment, recording setup, volume, annotation, QA, rights, and schedule. Spirelight publishes planning rates and project minimums where available; the written proposal confirms what is included and the final price.

How do I choose an AI data collection company?

Compare named recruitment methods, the consent and rights required for your use, recording and annotation protocols, QA evidence, acceptance criteria, delivery artifacts, and applicable case studies. Ask whether a sample matched to your brief is currently available, and compare quotes against the same specification.

Can you collect speech data for a small project?

A small pilot can be considered when it can answer a defined evaluation question. Feasibility, the project minimum, sample status, deliverables, rights, and price are confirmed for the requested configuration before work begins.

How should buyers assess GDPR and EU AI Act requirements for speech data?

Voice recordings can be personal data, and biometric processing may trigger additional requirements. The buyer and provider should document their roles, lawful basis, purpose, consent where applicable, retention, transfers, security, provenance, and required AI Act records for the specific project. Denmark-based operations do not create blanket compliance, so the buyer should review the written scope with counsel.

Ready to start? Request a quote or add audio and speech annotation to your project.