Voice agent evaluationCustom-scoped evaluation set

Voice agent evaluation data for accents, noise, and real call conditions

Build a private test set around the callers, channels, languages, and failure modes your voice agent actually faces. Spirelight scopes the collection, reference transcripts, slices, and holdout rules with your team.

Sample fulfilment: Request representative conversational and telephony samples. The team selects and sends a relevant slice manually after reviewing your language and channel requirements.

Collection specification

Audio
Wideband, narrowband, mono, or separate channels by project
Coverage
Accents, languages, SNR bands, interruptions, overlap, and domain scenarios
References
Human-reviewed transcripts and agreed scoring fields
Holdout
Collection split and non-reuse terms documented in the project scope
Delivery
Audio, transcripts, metadata, split manifest, and evaluation brief

Test the conditions hidden by a single average score

A useful voice-agent evaluation set separates accent, language, channel, noise, interruption, and scenario effects. The result is not just one word error rate. It shows which caller groups and conditions cause the system to fail.

The collection is scoped around your deployment. If a hosted model cannot be fine-tuned, the same held-out set can still compare providers and detect regressions after each release.

Keep evaluation audio out of the training delivery

Training leakage makes a benchmark look better without making the production system better. We define the evaluation split, file manifest, access rules, and any non-reuse commitment in the statement of work before recording starts.

This offer is not the right fit when

  • You need a public benchmark with unrestricted redistribution.
  • You cannot describe the deployment channel, language, or main failure mode.
  • You need a ready-made fixed pack without a scoping step.
Prepare the brief

Bring a specification your vendors can price consistently

Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.

Build a dataset specification
Relevant samples

Receive evaluation samples

Request representative conversational and telephony samples. The team selects and sends a relevant slice manually after reviewing your language and channel requirements.

Enter your work email and verify it with a 6-digit code. An admin then reviews the request and sends a relevant sample manually within two business days. Use the optional field to name the language, channel, or condition you need to inspect.

Need a scoped collection plan?

Send the language, channel, volume, annotation, and timing you know. The quote form opens directly in project-brief mode.