Compare speech data providers on evidence, not sales claims
Use this 12-criterion scorecard to test rights, provenance, coverage, quality, security, evaluation integrity, and delivery before committing to a speech dataset vendor.
How should buyers compare speech data providers?
Define the model job and non-negotiable risks first. Give every vendor the same brief, score only claims supported by inspectable evidence, and verify the most important claims in a representative paid pilot. Use the weighted total to compare viable vendors, not to override a rights, privacy, security, or leakage deal breaker.
Make the comparison fair and falsifiable
A vendor score is useful only when each bidder is solving the same problem. State the target model behavior, deployment channel, users, languages, failure modes, intended data uses, and acceptance tests. Separate required evidence from optional differentiators.
Freeze one brief
Send identical coverage, channel, rights, format, quality, and delivery requirements to every provider.
Set gates first
Name the conditions that cannot be traded for a higher score, such as missing consent evidence or evaluation leakage.
Ask for artifacts
Request samples, manifests, consent wording, protocols, reports, policies, license terms, and named assumptions.
Score separately
Have product, ML, legal/privacy, security, and procurement owners score their fields before discussing the result.
Test the hard slices
Use a representative pilot that includes the noisiest channel, hardest locale, and most important acceptance tests.
Contract the evidence
Carry accepted definitions, metrics, rights, replacement rules, and reporting into the statement of work.
Use the same 0–5 evidence scale for every criterion
Score what can be inspected and enforced today. A polished promise without an artifact is not equivalent to a documented, sample-backed, contractable process.
Example: a vendor scores 4 on a criterion weighted at 12%. The criterion contributes (4 ÷ 5) × 12 = 9.6 points. Record the artifact or test behind the score in the evidence column.
The 12 weighted criteria
The weights total 100. They are a practical default for speech-data procurement, not a universal risk model. Adjust them before receiving bids if a project has materially different priorities, then keep them fixed for all vendors.
| # | Criterion | Weight | What a score of 5 requires | Evidence to request | Example red flag |
|---|---|---|---|---|---|
| 01 | Purpose and use-case fit | 12% | The proposed data demonstrably matches the model task, target users, error modes, and deployment conditions. | Collection or dataset specification, representative sample, and explicit assumptions. | Generic inventory is offered without mapping it to the stated model job. |
| 02 | Rights, consent, and permitted use | 12% | Notice, consent, license, and downstream terms are traceable and cover the intended training, evaluation, retention, and derivative-model uses. | Consent wording, participant notice, license, rights register, and withdrawal/deletion process. | The provider cannot show the permission chain or relies on implied permission for voice or biometric uses. |
| 03 | Coverage and representativeness | 10% | Speaker, locale, demographic, device, and environment quotas are defined, measurable, and reported after quality acceptance. | Quota plan, recruitment method, coverage report, limitations, and post-QA counts. | Headline counts cannot be reconciled with accepted unique speakers or usable hours. |
| 04 | Recording and channel realism | 8% | Capture reproduces the production channel and failure modes with controlled variation and documented equipment. | Capture protocol, device/channel metadata, acoustic measurements, and representative audio. | Clean microphone audio is used to represent telephony, far-field, in-car, or noisy deployment. |
| 05 | Provenance and documentation | 10% | Each asset traces to its session, rights record, protocol, transformations, and dataset version. | Data card, manifest, stable IDs, lineage fields, transformation log, and version history. | Files cannot be linked back to session-level provenance and permission records. |
| 06 | Annotation and transcription quality | 10% | The label policy is documented and quality is measured on a representative blinded sample using an agreed method. | Annotation guide, annotator qualification, adjudication rules, gold set, and per-slice report. | One accuracy claim is given without a method, sample definition, or slice-level result. |
| 07 | Quality and acceptance testing | 10% | Machine and human checks, rejection rules, pilot gates, and replacement terms are agreed before production delivery. | QA plan, acceptance script, defect taxonomy, pilot report, and replacement terms. | Acceptance is subjective or assessed only by the provider. |
| 08 | Privacy, security, and governance | 8% | Data flows, access, retention, incidents, subprocessors, and sensitive-data controls are documented for the project. | Data-flow diagram, security controls, retention schedule, subprocessor list, and incident process. | Raw voice data is moved or retained without an agreed purpose, access model, or deletion path. |
| 09 | Evaluation split and leakage control | 8% | Speaker, session, prompt, and source separation protect held-out evaluation data and are verifiable in the manifest. | Split protocol, deduplication method, separation report, and non-reuse terms. | The provider cannot establish whether evaluation speakers or utterances also occur in training data. |
| 10 | Pilot and delivery process | 5% | A representative pilot tests the hardest conditions and produces an auditable production and release plan. | Pilot plan, milestones, escalation path, sample manifest, and change control. | The pilot excludes the conditions most likely to fail in production. |
| 11 | Commercial and licensing clarity | 4% | Price units, minimums, reuse, exclusivity, territories, term, restrictions, and change costs are explicit. | Priced statement of work, license schedule, assumptions, and change-order terms. | A low headline price excludes required rights, QA, metadata, or replacements. |
| 12 | Continuity, corrections, and support | 3% | Named ownership, versioning, issue response, corrections, and end-of-project handover are defined. | Correction SLA, release notes, version policy, and handover checklist. | No process exists for tracing and correcting defects found after delivery. |
Do not let the total hide a deal breaker
A weighted score helps rank acceptable options. It does not make an unacceptable permission chain, security model, or evaluation design acceptable. Decide which red flags stop procurement, trigger specialist review, or require a corrected pilot.
Common stop-and-investigate signals
- Consent wording or the chain of permitted uses cannot be inspected.
- The license does not clearly cover the intended model and retention use.
- Delivered files cannot be traced to session and rights records.
- Training and evaluation speaker or utterance overlap cannot be tested.
- Security, retention, subprocessor, or deletion responsibilities are undefined.
- A pilot sample does not represent the hardest deployment conditions.
- Quality claims lack a metric, sample definition, or buyer acceptance test.
- Dataset size or speaker counts cannot be reconciled after rejection and deduplication.
Move from longlist to contracted delivery
- Define the job: write the use case, intended data uses, target slices, channel, success metrics, and risk gates.
- Issue one evidence request: send every provider the same specification and scorecard fields.
- Score the response: use 0–5 only when there is a named artifact, inspectable sample, test result, or contractual commitment.
- Run a representative pilot: include difficult slices and execute the buyer's acceptance tests on raw outputs and manifests.
- Rescore verified evidence: update provisional scores after the pilot; document disagreements and remaining assumptions.
- Contract the controls: carry the accepted rights, definitions, tests, reporting, replacements, corrections, and handover into the order.
Keep the completed scorecard with the procurement record. If the dataset, model purpose, vendor process, or applicable risk changes materially, repeat the relevant checks rather than treating the score as permanent.
What informed this framework
This is Spirelight's procurement framework, not an official compliance checklist. Its emphasis on documented context, risk controls, governance, and traceability is informed by the following primary sources. Applicability depends on the project, jurisdiction, role, and intended use.
NIST AI Risk Management Framework
NIST's voluntary framework organizes AI risk work around Govern, Map, Measure, and Manage. The scorecard turns several vendor-facing evidence questions into a repeatable procurement record.
Regulation (EU) 2016/679, the GDPR
The official regulation sets rules for processing personal data, including lawful bases and special categories. The scorecard structures evidence questions for buyer review; it is not a legal determination for a particular dataset or use.
Regulation (EU) 2024/1689, the EU AI Act
The official regulation includes provisions on data and data governance for certain AI systems, with application dates for several high-risk obligations later adjusted by Regulation (EU) 2026/1744. This scorecard flags evidence areas for review; it does not determine whether a system or provider falls within a particular obligation.
Datasheets for Datasets
The paper proposes documenting dataset motivation, composition, collection, preprocessing, uses, distribution, and maintenance. Those questions inform the provenance and documentation criterion here.
Frequently asked questions
What is a good speech data vendor score?
There is no universal passing score. Define deal breakers first, compare the total among otherwise viable options, and verify the highest-risk claims in a representative pilot. Two vendors with the same total can have very different risk profiles.
Can a high total compensate for missing consent evidence?
No. If missing rights, consent, provenance, security, or evaluation-separation evidence creates unacceptable risk for the intended use, treat it as a gate rather than a low-scoring tradeoff.
Should the weights change for a custom collection?
They can, but change them before bids arrive and use the same weights for every provider. A sensitive or regulated use may give more weight to rights, governance, and security; an evaluation set may give more weight to leakage controls and slice coverage.
Can the vendor complete the scorecard?
A vendor can populate the evidence and notes columns, but the buyer should own the score. Product, ML, privacy/legal, security, and procurement reviewers should score the fields relevant to them.
How should a pilot affect the score?
Treat proposal-stage scores as provisional. After testing representative audio, manifests, labels, quotas, and acceptance results, rescore the relevant criteria using the verified evidence.
Does this replace legal or security review?
No. It structures vendor comparison and evidence requests. Qualified legal, privacy, security, procurement, and AI-governance reviewers should assess the rules and risks for the specific project.
Send us the brief and the evidence standard
Share the languages, target speakers, channel, hours, annotations, intended use, and acceptance tests. Spirelight will review the requirements and respond with a scoped collection or annotation plan.