Specify and evaluate speech data before you ask for a price
Use these ungated templates to compare vendors on the same requirements, inspect the documentation fields worth asking for, and keep held-out evaluation separate from training delivery.
Five practical starting points
These are scoping aids. The specimen is clearly marked and does not imply that every field is included in every delivery.
Speech data provider scorecard
Compare vendors on rights, provenance, coverage, quality, security, evaluation integrity, and delivery with 12 weighted criteria and an editable CSV.
Speech dataset specification
Define languages, speakers, hours, capture conditions, annotations, rights, evaluation splits, and acceptance tests. Copy or download the generated brief.
Speech data card and provenance structure
Inspect a sanitized example covering provenance, consent references, capture, coverage, annotations, splits, QA, rights, and a manifest row.
Voice-agent evaluation brief
Scope conditions, held-out rules, reference data, slices, metrics, acceptance criteria, rights, and release decisions before collecting test audio.
Speaker diarization evaluation
Define reference labels, overlap and collar rules, scoring configuration, held-out slices, acceptance thresholds, rework, rights, and review.
Get a free sample for the buying job, not just the language
These offers are custom-scoped. Tell us the language and channel, then the team selects a matching example and sends it manually within 48 hours.
Call-center and telephony
A project can specify simulated calls, phone-channel delivery, reference transcripts, consent requirements, and agreed scenarios.
Voice-agent evaluation
Describe the accent, noise, channel, and failure conditions you need to inspect; example availability is confirmed after review.
Full-duplex conversation
A project can specify separate-channel conversation, overlap, interruptions, backchannels, and turn timing.