Speech synthesis evaluationCustom evaluation panels

Listening-test evaluation data for expressive speech synthesis

Automatic metrics rank synthesis systems that native listeners rank differently. Recruited listener panels give you mean opinion scores, naturalness ratings, and a read on whether the emotion in a synthesized line lands at all.

Free sample: Send a handful of synthesized clips. The team runs a small scored panel and sends the rated results back manually within 48 hours, free of charge.

Collection specification

Mean opinion score
Five-point MOS collected from recruited native listeners, reported with confidence intervals and per-rater variance
Naturalness
Separate naturalness and intelligibility scales, so a system is not penalized twice for one defect
Emotional appropriateness
Whether the intended emotion was perceived, scored against the same seven-class taxonomy used across our annotation work
Comparative tests
A/B and MUSHRA-style preference panels when ranking systems matters more than absolute scores
Listeners
Native speakers recruited per language and screened, never crowd workers labeling outside their language
Agreement
Per-rater agreement published with every panel, not just the aggregate score

A MOS number is only as good as the panel behind it

Mean opinion score is easy to report and easy to inflate. Panel size, listener screening, whether raters are native speakers, and how systems are interleaved all move the number more than most model changes do.

Panels are run with screened native listeners and published with per-rater variance, so a half-point difference between two systems can be read as real or as noise rather than assumed to be either.

Expressive synthesis needs an emotion question, not just a quality question

A line can be rated highly natural and still convey the wrong feeling. For expressive and character voices that is the failure that matters, and a naturalness scale will never surface it.

Listeners are asked separately what emotion they perceived, scored against the same seven-class taxonomy used in our annotation work, so evaluation results line up with training labels instead of living in a different vocabulary.

This offer is not the right fit when

  • You want an automatic objective metric rather than human listeners.
  • You need evaluation of speech recognition accuracy rather than synthesis quality.
  • You want a single unscreened crowd panel at the lowest possible price.
Prepare the brief

Bring a specification your vendors can price consistently

Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.

Build a dataset specification
Free sample

Get a free scored sample

Send a handful of synthesized clips. The team runs a small scored panel and sends the rated results back manually within 48 hours, free of charge.

Enter your work email and verify it with a 6-digit code. The team then selects a representative sample and sends it manually within 48 hours. Use the optional field to name the language, channel, or condition you need to inspect.

Add project details (optional)

Want to talk through your taxonomy first?

Category sets, dimensional scales, and agreement thresholds are usually a conversation, not a form field. Book a short call to work through the taxonomy and conditions the labels need to hold up under.