Mean opinion score is easy to report and easy to inflate. Panel size, listener screening, whether raters are native speakers, and how systems are interleaved all move the number more than most model changes do.
Panels are run with screened native listeners and published with per-rater variance, so a half-point difference between two systems can be read as real or as noise rather than assumed to be either.