Speech data services for AI, built to your model requirements

We scope custom speech collection and annotation for AI training. Each agreed project defines the target languages, speaker profiles, dialects, recording format, metadata, background conditions, QA, rights, and acceptance criteria before production.

Overview

The training data behind ASR, TTS, and voice AI.

Spirelight is a speech data company. A project can cover speaker recruitment, custom recording, transcription, labeling, QA, rights, and delivery under one written scope. Buyers use custom work when an existing dataset does not fit a required language, accent, recording condition, label schema, or acceptance test.

Our work splits into two core services. Data collection captures new audio to your spec, from scripted prompts to spontaneous, far-field, and in-car speech in the languages and regional varieties required by the project. Annotation turns raw recordings into structured, labeled training data with transcripts, timestamps, speaker tags, and events. Recruitment, consent, recording, annotation, QA, rights, and delivery artifacts are defined for each project. Use the catalog to review language-specific collection configurations and planning inputs.

New to sourcing speech data? Start with our guides on what speech data is, how much you need, what audio annotation involves, speaker diarization and speaker labeling, data license agreements, and how to buy AI training data. Comparing vendors or scoping a project? See how to choose an audio annotation company and how a custom speech data collection project is scoped and run.

Our services

Choose the service that fits your project.

Commission a custom collection, have your audio annotated, or start from a language-specific collection configuration.

Collection offers

Start with the model job and operating environment.

These custom-scoped offers combine collection, consent, annotation, and evaluation requirements for specific speech AI applications.

Voice agent evaluation data

Private held-out test sets scoped to your callers, channels, languages, accents, noise conditions, and failure modes, with human-reviewed references and documented holdout rules.

Explore the evaluation offer

Call-center speech data

Purpose-recorded, consented customer-service simulations with telephony delivery, separate-channel options, human-reviewed transcripts, and custom domain scenarios.

Explore the call-data offer

Full-duplex conversational speech data

Separate-channel conversations that preserve natural overlap, interruptions, backchannels, repairs, and turn timing for speech-to-speech and low-latency voice systems.

Explore the conversation offer

Audiovisual speech data collection

Synchronized audio and video for visual speech, lip alignment, audiovisual agents, and digital-human research, with the intended model use stated in participant releases.

Explore the audiovisual offer

Warehouse voice data collection

Short commands, numbers, item codes, confirmations, and corrections recorded for the headsets, languages, and operating noise used on the floor.

Explore the industrial offer

Simulated emergency-call speech data

Consented simulations—not real victim calls—covering multilingual scenarios, urgency, interruptions, degraded phone audio, and human-reviewed reference transcripts.

Explore the simulated-call offer

Emotion annotation

Per-turn emotion labels on speech: discrete classes, optional valence and arousal ratings, intent kept separate, and per-class annotator agreement reported.

Explore the emotion offer

Call-center emotion annotation

Turn-level emotion and sentiment labels on call audio for quality monitoring, escalation prediction, and agent assist, with a documented consent chain.

Explore the call-center emotion offer

Voice agent emotion data

Turn-level affect labels on spontaneous conversational speech, so a voice agent can detect frustration and be evaluated on how it responds.

Explore the voice agent emotion offer

TTS evaluation data

Recruited listener panels scoring mean opinion score, naturalness, and emotional appropriateness for expressive speech synthesis, with per-rater agreement.

Explore the TTS evaluation offer

Video data collection

Custom video datasets recorded to spec: talking-head and dialogue footage, expression and gesture sets, liveness sequences, and in-cabin captures, with per-participant consent.

Explore the video offer

Talking-head video data

On-camera monologue and dialogue for avatars, lip sync, and digital humans: fixed framing, frame-accurate sync, and likeness releases that name synthetic-media use.

Explore the talking-head offer

Voice anti-spoofing data

Genuine and spoofed speech pairs for deepfake detector training: TTS, voice conversion, and replay attacks generated only from speakers who consented to the spoof.

Explore the anti-spoofing offer
Transcription

Transcription staffed by language and dialect.

Native contributor teams matched to the language and regional variant in the recording, with the written convention agreed before production starts. Start from the full transcription service overview, or see how the global contributor workforce behind every project is recruited and managed.

Spanish transcription

Annotators staffed by the recorded variant, from Castilian to Rioplatense, with the output convention fixed before the first file.

Explore Spanish transcription

German transcription

Coverage from Bavarian to Low German, with an orthographic convention agreed before production starts.

Explore German transcription

Italian transcription

Neapolitan, Sicilian, and Venetian material handled as separate regional languages, not as accented Italian.

Explore Italian transcription

Korean transcription

Seoul, Gyeongsang, and Jeolla audio, with the speaker's actual speech level preserved throughout.

Explore Korean transcription

Chinese transcription

Mandarin transcription with the character set, segmentation, and number conventions fixed upfront.

Explore Chinese transcription

Cantonese transcription

Hong Kong audio delivered as written Cantonese or Standard Written Chinese, in traditional characters, with code-mixed English kept.

Explore Cantonese transcription

Japanese transcription

A written script policy across kanji, hiragana, and katakana, agreed before the first file.

Explore Japanese transcription

Arabic transcription

Dialect coverage across Egypt, the Gulf, the Levant, and Morocco, staffed by market.

Explore Arabic transcription

French transcription

France and francophone Africa staffed separately, with code-switching handled by convention.

Explore French transcription

Portuguese transcription

Brazil and Portugal staffed separately, because the two markets do not share one annotator pool.

Explore Portuguese transcription

Russian transcription

Sourced across Russian-speaking markets, with a stated convention for mixed contact speech.

Explore Russian transcription

Hindi transcription

Mixed English and Hindi speech transcribed to an agreed script policy, not treated as noise.

Explore Hindi transcription

Turkish transcription

Loanword conventions and dialect assignment settled before the first file is processed.

Explore Turkish transcription
Get started

Tell us the speech data you need

Send over your speaker profiles, language needs, and background noise conditions. Our team will assess feasibility and identify what is needed to scope a written project proposal.

Targeted
contributor recruitment matched to each project brief
Scoped
language and dialect feasibility confirmed per project
Defined
QA workflow and acceptance criteria agreed in scope
Written
rights, artifacts, schedule, and price documented per project