Spirelight provides audio and speech annotation services for AI teams building ASR, TTS, and voice models. We scope transcription, labels, timestamps, reviewer qualifications, QA metrics, acceptance criteria, and delivery format around the buyer's audio and model requirements.
Short answer: Audio annotation services turn raw speech recordings into labeled training data: verbatim transcription, timestamps and segmentation, speaker turns, non-speech event tags, and intent and emotion labels applied under a written guideline. Quality assurance is defined before production, including reviewer qualifications, sampling, and measured inter-annotator agreement where the use case requires it. Delivery follows the buyer's label schema in an agreed file format such as JSON, JSONL, CSV, SRT, or VTT.
| Annotation layer | What it covers | Typical output formats |
|---|---|---|
| Verbatim transcription | Human or ASR-assisted transcripts with agreed rules for casing, numbers, disfluencies, and domain terminology. | JSON, JSONL, CSV |
| Timestamps and segmentation | Word- or segment-level timestamps with agreed alignment units, boundary rules, and tolerances. | JSON, SRT, VTT |
| Speaker turns and diarization labels | Anonymous speaker labels with agreed rules for overlap, turns, and unknown speakers. | JSON, JSONL, CSV |
| Acoustic and non-speech events | Background noise, music, laughter, silence, and channel conditions labeled in place. | JSON, JSONL |
| Emotion and affect classes | Utterance-level intent, sentiment, and emotion labels applied against a defined tag set. | JSON, CSV |
| Custom schema fields | Labels, metadata fields, and edge cases mapped to the buyer's own labeling guideline. | Custom manifest schema |
Formats are options, not defaults: the written scope confirms the file format for each layer, and a custom manifest schema can replace any of them.
Spirelight assesses managed audio-annotation projects for buyer-provided or newly collected speech audio. Before production, the written scope confirms the supported labels and languages, reviewer qualifications, QA method and acceptance threshold, security workflow, output schema, schedule, and price for the specific brief.
| Scope field | What the written proposal should confirm |
|---|---|
| Representative input | Language, locale, channel, codec, recording conditions, source status, and a reviewed sample where one may be shared. |
| Labels and edge cases | Definitions, boundary rules, overlap, unknowns, timestamp units, and the agreed machine-readable output schema. |
| People and calibration | Annotator and reviewer qualifications plus any pilot, guideline-calibration, or adjudication step required for the use case. |
| QA and acceptance | Metric, sampling method, threshold, evidence, escalation, and the condition that triggers rework or replacement. |
| Delivery and controls | Manifest and file formats, security and access requirements, batch cadence, schedule assumptions, and pricing basis. |
The annotation brief defines the schema, target languages, reviewer qualifications, QA method, and acceptance threshold before production. Feasibility is confirmed for the requested files and locales.
Human or ASR-assisted transcripts can follow agreed rules for casing, numbers, disfluencies, domain terminology, review, and acceptance.
Anonymous speaker labels can be scoped for dialogue and multi-party audio. Agreed role or identity fields are included only when the project and applicable rights permit them; overlap, turns, unknown speakers, and acceptance rules belong in the guideline.
Non-speech events labeled in place: background noise, music, laughter, silence, and channel conditions your model needs to recognize or ignore.
Word- or segment-level timestamps can be included, with alignment units, boundary rules, tolerances, and the acceptance test agreed before production.
Utterance-level intent, sentiment, and emotion labels for voice assistants, conversational AI, and affect models, applied against a defined tag set. For discrete classes, valence and arousal, and agreement reporting on affect specifically, see emotion annotation.
Have a labeling guideline of your own? The annotation scope can map labels, metadata fields, edge cases, and QA rules to that schema.
Different models need different labels. Each annotation project defines the transcript rules, label schema, reviewer requirements, QA metrics, and acceptance thresholds for its intended use.
For speech recognition, a scope can include verbatim text, rules for accents, overlaps and domain terms, timestamps, and speaker turns. The guideline and accuracy threshold are agreed against the target model.
For text to speech, a scope can include text normalization, phrase and pause marking, and word-level alignment on audio that meets agreed recording requirements.
For wake-word, intent, and conversational systems, a scope can add intent, slot, and event labels to the transcript under an agreed tag set and review plan. Emotion labels are scoped as a separate track, with their own tag set and reviewer requirements.
Need the recordings collected first? Pair annotation with our speech and audio data collection services to run collection and labeling as one project.
Audio data labeling is only useful if it is consistent. The project plan defines measurable quality controls, sampling, escalation, and acceptance criteria for the agreed batches. Buyers can start with the speech dataset specification worksheet and carry its labels, acceptance tests, and rights requirements into the quote.
An agreement protocol can assign duplicated or sampled items to multiple annotators, define the metric and threshold, and specify how disagreements update the guideline or trigger rework.
The written QA plan can assign single-pass, sampled second-pass, or full second-pass review by item type, risk, and schedule, with escalation and acceptance rules agreed per project.
The written QA plan states which batches are sampled, how sample sizes are chosen, which metrics are checked, and what triggers rework before acceptance.
Projects that require language-matched review can specify language, region, dialect, domain, slang, and code-switching qualifications. Reviewer feasibility and evidence are confirmed in the staffing plan.
Working across many locales? Review our language-specific collection configurations, then confirm reviewer feasibility, sample status, schema, rights, schedule, and final price against your brief.
The written handoff plan defines the label schema and the audio, metadata, manifest, and delivery artifacts included in the project. The specimen data card illustrates documentation fields to request; it is not a real annotation sample or a promise that every field is included.
JSON, JSONL, CSV, SRT, VTT, or a custom manifest schema can be considered and confirmed in the written scope.
Audio format, sample rate, bit depth, channel layout, and any conversion rules are agreed against the source files and target pipeline.
Delivery channel, batch cadence, checksums, identifiers, and required provenance fields are specified in the written handoff plan.
Spirelight focuses its annotation offer on speech and voice. Buyers should evaluate the named language reviewers, guideline, QA evidence, and acceptance test for their own project.
When collection is in scope, the project defines the required consent, permitted uses, provenance identifiers, retention, and delivery records before recording.
Reviewer language, locale, dialect, and domain qualifications are matched to the agreed annotation brief and documented in the staffing plan.
The staffing plan names the language, dialect, domain, and ASR or TTS knowledge required of reviewers, and the acceptance test checks the resulting labels.
See the full picture of how we work across collection, transcription, and QA on our services overview. Still comparing vendors? Our guide to choosing an audio annotation company covers the criteria, types, and cost drivers worth checking before you sign.
Send us your label requirements, languages, audio volume, security constraints, and target schedule. We will review the brief and identify what is needed to scope a written proposal and acceptance test.
Spirelight can assess buyer-provided or newly collected speech audio against a written brief. The proposal confirms source and rights requirements, supported labels and languages, reviewer coverage, security, QA, output schema, schedule, and price before production.
A project can scope anonymous speaker-turn labels, timestamps, overlap and unknown-speaker rules, human review, scoring configuration, and acceptance criteria. Role or identity fields require a separate project and rights assessment.
Yes. The annotation scope can map labels, metadata fields, edge cases, and QA rules to a buyer's existing labeling guideline. The written proposal confirms the schema mapping, reviewer qualifications, and acceptance criteria before production.
Annotation can be priced per audio hour, item, speaker, or batch. Language, domain complexity, label schema, QA, security, and schedule affect the rate. The written proposal confirms the pricing unit, assumptions, acceptance criteria, and final price.
Cost is scoped per project. Annotation can be priced per audio hour, item, speaker, or batch, and language, domain complexity, label schema, QA, security, and schedule affect the rate. Request a quote with your audio volume, languages, and label requirements, and the written proposal will state the pricing unit, assumptions, and final price.
Turnaround depends on audio volume, language, label complexity, reviewer requirements, security, and acceptance testing. The written proposal states any pilot, throughput, batch cadence, review points, and delivery dates.
The project scope defines the labeling guideline, reviewer qualifications, sampling method, quality metrics, thresholds, escalation path, and acceptance test. Inter-annotator agreement or multi-pass review can be included when the use case requires it.
Quality assurance is defined in the project plan before annotation starts. The written QA plan can include inter-annotator agreement with a stated metric and threshold, single-pass or multi-pass review, and batch sampling with rework triggers agreed before acceptance.
For collected audio, the project scope defines the required consent, permitted uses, and provenance records. For buyer-provided audio, the contract defines controller and processor roles, instructions, security, retention, transfers, and any required processing agreement. Storage and access requirements are confirmed before work begins.
Annotation services label buyer-provided recordings or audio collected to an agreed specification. The catalogue presents language-specific custom collection configurations and planning inputs, not a promise of finished inventory. Feasibility, sample status, schema, QA, rights, schedule, and final price are confirmed for each brief.
Ready to start? Request a quote or explore speech and audio data collection.