A useful voice-agent evaluation set separates accent, language, channel, noise, interruption, and scenario effects. The result is not just one word error rate. It shows which caller groups and conditions cause the system to fail.
The collection is scoped around your deployment. If a hosted model cannot be fine-tuned, the same held-out set can still compare providers and detect regressions after each release.