Custom-scoped evaluation set

Voice agent evaluation data for accents, noise, and real call conditions

Build a private test set around the callers, channels, languages, and failure modes your voice agent actually faces. Spirelight scopes the collection, reference transcripts, slices, and holdout rules with your team.

Free sample: Tell us your language and channel requirements. The team selects a matching conversational or telephony example and sends it manually within 48 hours, free of charge.

Collection specification

Audio
Wideband, narrowband, mono, or separate channels by project
Coverage
Accents, languages, SNR bands, interruptions, overlap, and domain scenarios
References
Human-reviewed transcripts and agreed scoring fields
Holdout
Collection split and non-reuse terms documented in the project scope
Delivery
Audio, transcripts, metadata, split manifest, and evaluation brief
Pricing
Custom, scoped to your conditions

Test the conditions hidden by a single average score

A useful voice-agent evaluation set separates accent, language, channel, noise, interruption, and scenario effects. The result is not just one word error rate. It shows which caller groups and conditions cause the system to fail.

The collection is scoped around your deployment. If a hosted model cannot be fine-tuned, the same held-out set can still compare providers and detect regressions after each release.

Keep evaluation audio out of the training delivery

Training leakage makes a benchmark look better without making the production system better. We define the evaluation split, file manifest, access rules, and any non-reuse commitment in the statement of work before recording starts.

This offer is not the right fit when

  • You need a public benchmark with unrestricted redistribution.
  • You cannot describe the deployment channel, language, or main failure mode.
  • You need an immediately downloadable fixed pack without a scoping step.

Get a free sample

  1. Your details
  2. Your project
  3. Verify

Three short steps. The team then follows up manually within 48 hours, and confirms volume, rights, QA, and delivery if you want the scope priced.

Would rather talk it through first? Price an evaluation set

Prepare the brief

Bring a specification your vendors can price consistently

Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.

Build a dataset specification

Buyer documentation

Specimen data card and provenance structure, evaluation brief template, and acceptance and held-out questions.

Guides for this decision

Need a scoped collection plan? Send the language, channel, volume, annotation, and timing you know. The quote form opens directly in project-brief mode.