Audio Annotation Services for Speech AI

Spirelight provides audio and speech annotation services for AI teams building ASR, TTS, and voice models. We scope transcription, labels, timestamps, reviewer qualifications, QA metrics, acceptance criteria, and delivery format around the buyer's audio and model requirements.

What's included

What do audio annotation services include?

Short answer: Audio annotation services turn raw speech recordings into labeled training data: verbatim transcription, timestamps and segmentation, speaker turns, non-speech event tags, and intent and emotion labels applied under a written guideline. Quality assurance is defined before production, including reviewer qualifications, sampling, and measured inter-annotator agreement where the use case requires it. Delivery follows the buyer's label schema in an agreed file format such as JSON, JSONL, CSV, SRT, or VTT.

Annotation layerWhat it coversTypical output formats
Verbatim transcriptionHuman or ASR-assisted transcripts with agreed rules for casing, numbers, disfluencies, and domain terminology.JSON, JSONL, CSV
Timestamps and segmentationWord- or segment-level timestamps with agreed alignment units, boundary rules, and tolerances.JSON, SRT, VTT
Speaker turns and diarization labelsAnonymous speaker labels with agreed rules for overlap, turns, and unknown speakers.JSON, JSONL, CSV
Acoustic and non-speech eventsBackground noise, music, laughter, silence, and channel conditions labeled in place.JSON, JSONL
Emotion and affect classesUtterance-level intent, sentiment, and emotion labels applied against a defined tag set.JSON, CSV
Custom schema fieldsLabels, metadata fields, and edge cases mapped to the buyer's own labeling guideline.Custom manifest schema

Formats are options, not defaults: the written scope confirms the file format for each layer, and a custom manifest schema can replace any of them.

Buyer answer

What a managed annotation scope confirms.

Spirelight assesses managed audio-annotation projects for buyer-provided or newly collected speech audio. Before production, the written scope confirms the supported labels and languages, reviewer qualifications, QA method and acceptance threshold, security workflow, output schema, schedule, and price for the specific brief.

Scope fieldWhat the written proposal should confirm
Representative inputLanguage, locale, channel, codec, recording conditions, source status, and a reviewed sample where one may be shared.
Labels and edge casesDefinitions, boundary rules, overlap, unknowns, timestamp units, and the agreed machine-readable output schema.
People and calibrationAnnotator and reviewer qualifications plus any pilot, guideline-calibration, or adjudication step required for the use case.
QA and acceptanceMetric, sampling method, threshold, evidence, escalation, and the condition that triggers rework or replacement.
Delivery and controlsManifest and file formats, security and access requirements, batch cadence, schedule assumptions, and pricing basis.
What we annotate

Speech labels scoped from transcript to emotion.

The annotation brief defines the schema, target languages, reviewer qualifications, QA method, and acceptance threshold before production. Feasibility is confirmed for the requested files and locales.

Verbatim transcription

Human or ASR-assisted transcripts can follow agreed rules for casing, numbers, disfluencies, domain terminology, review, and acceptance.

Speaker labels and diarization

Anonymous speaker labels can be scoped for dialogue and multi-party audio. Agreed role or identity fields are included only when the project and applicable rights permit them; overlap, turns, unknown speakers, and acceptance rules belong in the guideline.

Event and noise tags

Non-speech events labeled in place: background noise, music, laughter, silence, and channel conditions your model needs to recognize or ignore.

Word-level timestamps

Word- or segment-level timestamps can be included, with alignment units, boundary rules, tolerances, and the acceptance test agreed before production.

Intent and emotion

Utterance-level intent, sentiment, and emotion labels for voice assistants, conversational AI, and affect models, applied against a defined tag set. For discrete classes, valence and arousal, and agreement reporting on affect specifically, see emotion annotation.

Custom label schemas

Have a labeling guideline of your own? The annotation scope can map labels, metadata fields, edge cases, and QA rules to that schema.

By model type

Annotation for ASR, TTS, and voice AI.

Different models need different labels. Each annotation project defines the transcript rules, label schema, reviewer requirements, QA metrics, and acceptance thresholds for its intended use.

ASR

For speech recognition, a scope can include verbatim text, rules for accents, overlaps and domain terms, timestamps, and speaker turns. The guideline and accuracy threshold are agreed against the target model.

TTS

For text to speech, a scope can include text normalization, phrase and pause marking, and word-level alignment on audio that meets agreed recording requirements.

Voice AI and assistants

For wake-word, intent, and conversational systems, a scope can add intent, slot, and event labels to the transcript under an agreed tag set and review plan. Emotion labels are scoped as a separate track, with their own tag set and reviewer requirements.

Need the recordings collected first? Pair annotation with our speech and audio data collection services to run collection and labeling as one project.

Process and quality

Define quality controls before annotation starts.

Audio data labeling is only useful if it is consistent. The project plan defines measurable quality controls, sampling, escalation, and acceptance criteria for the agreed batches. Buyers can start with the speech dataset specification worksheet and carry its labels, acceptance tests, and rights requirements into the quote.

Inter-annotator agreement

An agreement protocol can assign duplicated or sampled items to multiple annotators, define the metric and threshold, and specify how disagreements update the guideline or trigger rework.

Review passes

The written QA plan can assign single-pass, sampled second-pass, or full second-pass review by item type, risk, and schedule, with escalation and acceptance rules agreed per project.

QA sampling

The written QA plan states which batches are sampled, how sample sizes are chosen, which metrics are checked, and what triggers rework before acceptance.

Languages and dialects

Language-matched reviewer requirements.

Projects that require language-matched review can specify language, region, dialect, domain, slang, and code-switching qualifications. Reviewer feasibility and evidence are confirmed in the staffing plan.

Working across many locales? Review our language-specific collection configurations, then confirm reviewer feasibility, sample status, schema, rights, schedule, and final price against your brief.

Formats and delivery

Formats agreed for the target pipeline.

The written handoff plan defines the label schema and the audio, metadata, manifest, and delivery artifacts included in the project. The specimen data card illustrates documentation fields to request; it is not a real annotation sample or a promise that every field is included.

Label formats

JSON, JSONL, CSV, SRT, VTT, or a custom manifest schema can be considered and confirmed in the written scope.

Audio specs

Audio format, sample rate, bit depth, channel layout, and any conversion rules are agreed against the source files and target pipeline.

Handoff

Delivery channel, batch cadence, checksums, identifiers, and required provenance fields are specified in the written handoff plan.

Why Spirelight

A speech specialist, not a generic labeling vendor.

Spirelight focuses its annotation offer on speech and voice. Buyers should evaluate the named language reviewers, guideline, QA evidence, and acceptance test for their own project.

Project-specific provenance

When collection is in scope, the project defines the required consent, permitted uses, provenance identifiers, retention, and delivery records before recording.

Language-matched reviewers

Reviewer language, locale, dialect, and domain qualifications are matched to the agreed annotation brief and documented in the staffing plan.

Speech specialists

The staffing plan names the language, dialect, domain, and ASR or TTS knowledge required of reviewers, and the acceptance test checks the resulting labels.

See the full picture of how we work across collection, transcription, and QA on our services overview. Still comparing vendors? Our guide to choosing an audio annotation company covers the criteria, types, and cost drivers worth checking before you sign.

Request a quote

Tell us about your audio and we will scope it.

Send us your label requirements, languages, audio volume, security constraints, and target schedule. We will review the brief and identify what is needed to scope a written proposal and acceptance test.

Audio annotation FAQ

Can Spirelight annotate buyer-provided audio?

Spirelight can assess buyer-provided or newly collected speech audio against a written brief. The proposal confirms source and rights requirements, supported labels and languages, reviewer coverage, security, QA, output schema, schedule, and price before production.

Can speaker diarization be included?

A project can scope anonymous speaker-turn labels, timestamps, overlap and unknown-speaker rules, human review, scoring configuration, and acceptance criteria. Role or identity fields require a separate project and rights assessment.

Can you annotate audio to our existing schema and guidelines?

Yes. The annotation scope can map labels, metadata fields, edge cases, and QA rules to a buyer's existing labeling guideline. The written proposal confirms the schema mapping, reviewer qualifications, and acceptance criteria before production.

How is pricing structured for audio annotation services?

Annotation can be priced per audio hour, item, speaker, or batch. Language, domain complexity, label schema, QA, security, and schedule affect the rate. The written proposal confirms the pricing unit, assumptions, acceptance criteria, and final price.

What do audio annotation services cost?

Cost is scoped per project. Annotation can be priced per audio hour, item, speaker, or batch, and language, domain complexity, label schema, QA, security, and schedule affect the rate. Request a quote with your audio volume, languages, and label requirements, and the written proposal will state the pricing unit, assumptions, and final price.

What is the typical turnaround for an annotation project?

Turnaround depends on audio volume, language, label complexity, reviewer requirements, security, and acceptance testing. The written proposal states any pilot, throughput, batch cadence, review points, and delivery dates.

How do you ensure annotation accuracy and QA?

The project scope defines the labeling guideline, reviewer qualifications, sampling method, quality metrics, thresholds, escalation path, and acceptance test. Inter-annotator agreement or multi-pass review can be included when the use case requires it.

Do audio annotation services include quality assurance?

Quality assurance is defined in the project plan before annotation starts. The written QA plan can include inter-annotator agreement with a stated metric and threshold, single-pass or multi-pass review, and batch sampling with rework triggers agreed before acceptance.

How do you handle consent and GDPR?

For collected audio, the project scope defines the required consent, permitted uses, and provenance records. For buyer-provided audio, the contract defines controller and processor roles, instructions, security, retention, transfers, and any required processing agreement. Storage and access requirements are confirmed before work begins.

How do annotation services relate to the dataset catalogue?

Annotation services label buyer-provided recordings or audio collected to an agreed specification. The catalogue presents language-specific custom collection configurations and planning inputs, not a promise of finished inventory. Feasibility, sample status, schema, QA, rights, schedule, and final price are confirmed for each brief.

Ready to start? Request a quote or explore speech and audio data collection.