Custom evaluation panels
Listening-test evaluation data for expressive speech synthesis
Automatic metrics rank synthesis systems that native listeners rank differently. Recruited listener panels give you mean opinion scores, naturalness ratings, and a read on whether the emotion in a synthesized line lands at all.
Free sample: Tell us the languages and voice styles you are evaluating. The team replies within 48 hours to collect a handful of your synthesized clips, then runs a small scored panel and sends the rated results back manually, free of charge.
Collection specification
- Mean opinion score
- Five-point MOS collected from recruited native listeners, reported with confidence intervals and per-rater variance
- Naturalness
- Separate naturalness and intelligibility scales, so a system is not penalized twice for one defect
- Emotional appropriateness
- Whether the intended emotion was perceived, scored against the same seven-class taxonomy used across our annotation work
- Comparative tests
- A/B and MUSHRA-style preference panels when ranking systems matters more than absolute scores
- Listeners
- Native speakers recruited per language and screened, never crowd workers labeling outside their language
- Agreement
- Per-rater agreement published with every panel, not just the aggregate score
- Pricing
- Custom, scoped to your conditions
A MOS number is only as good as the panel behind it
Mean opinion score is easy to report and easy to inflate. Panel size, listener screening, whether raters are native speakers, and how systems are interleaved all move the number more than most model changes do.
Panels are run with screened native listeners and published with per-rater variance, so a half-point difference between two systems can be read as real or as noise rather than assumed to be either.
Expressive synthesis needs an emotion question, not just a quality question
A line can be rated highly natural and still convey the wrong feeling. For expressive and character voices that is the failure that matters, and a naturalness scale will never surface it.
Listeners are asked separately what emotion they perceived, scored against the same seven-class taxonomy used in our annotation work, so evaluation results line up with training labels instead of living in a different vocabulary.
This offer is not the right fit when
- You want an automatic objective metric rather than human listeners.
- You need evaluation of speech recognition accuracy rather than synthesis quality.
- You want a single unscreened crowd panel at the lowest possible price.
Get a free scored sample
- Your details
- Your project
- Verify
Three short steps. The team then follows up manually within 48 hours, and confirms volume, rights, QA, and delivery if you want the scope priced.
Would rather talk it through first? Talk through your label taxonomy
Prepare the brief
Bring a specification your vendors can price consistently
Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.
Buyer documentation
Specimen data card and provenance structure, evaluation brief template, and acceptance and held-out questions.
Related services
Guides for this decision
Want to talk through your taxonomy first? Category sets, dimensional scales, and agreement thresholds are usually a conversation, not a form field. Book a short call to work through the taxonomy and conditions the labels need to hold up under.