Free buyer framework · 100 points

Compare speech data providers on evidence, not sales claims

Use this 12-criterion scorecard to test rights, provenance, coverage, quality, security, evaluation integrity, and delivery before committing to a speech dataset vendor.

Last reviewed August 2, 2026Method 0–5 evidence scaleFormat Ungated CSV
The short answer

How should buyers compare speech data providers?

Define the model job and non-negotiable risks first. Give every vendor the same brief, score only claims supported by inspectable evidence, and verify the most important claims in a representative paid pilot. Use the weighted total to compare viable vendors, not to override a rights, privacy, security, or leakage deal breaker.

Weighted formula criterion points = (score ÷ 5) × weight Add the 12 criterion points for a total out of 100. A score of 0 means no usable evidence; 5 means complete, traceable, and contractable evidence.
Before scoring

Make the comparison fair and falsifiable

A vendor score is useful only when each bidder is solving the same problem. State the target model behavior, deployment channel, users, languages, failure modes, intended data uses, and acceptance tests. Separate required evidence from optional differentiators.

Freeze one brief

Send identical coverage, channel, rights, format, quality, and delivery requirements to every provider.

Set gates first

Name the conditions that cannot be traded for a higher score, such as missing consent evidence or evaluation leakage.

Ask for artifacts

Request samples, manifests, consent wording, protocols, reports, policies, license terms, and named assumptions.

Score separately

Have product, ML, legal/privacy, security, and procurement owners score their fields before discussing the result.

Test the hard slices

Use a representative pilot that includes the noisiest channel, hardest locale, and most important acceptance tests.

Contract the evidence

Carry accepted definitions, metrics, rights, replacement rules, and reporting into the statement of work.

Scoring rule

Use the same 0–5 evidence scale for every criterion

Score what can be inspected and enforced today. A polished promise without an artifact is not equivalent to a documented, sample-backed, contractable process.

0No usable evidence or the requirement is not addressed.
1A general assertion with no project-specific artifact.
2Partial documentation or a sample that leaves material gaps.
3Adequate, project-relevant evidence with manageable open points.
4Strong evidence verified in samples, reports, or a representative pilot.
5Complete, traceable evidence with measurable contractual commitments.

Example: a vendor scores 4 on a criterion weighted at 12%. The criterion contributes (4 ÷ 5) × 12 = 9.6 points. Record the artifact or test behind the score in the evidence column.

Evaluation matrix

The 12 weighted criteria

The weights total 100. They are a practical default for speech-data procurement, not a universal risk model. Adjust them before receiving bids if a project has materially different priorities, then keep them fixed for all vendors.

Default weights and evidence prompts. The downloadable CSV includes blank score, weighted-point, red-flag, and notes fields.
#CriterionWeightWhat a score of 5 requiresEvidence to requestExample red flag
01Purpose and use-case fit12%The proposed data demonstrably matches the model task, target users, error modes, and deployment conditions.Collection or dataset specification, representative sample, and explicit assumptions.Generic inventory is offered without mapping it to the stated model job.
02Rights, consent, and permitted use12%Notice, consent, license, and downstream terms are traceable and cover the intended training, evaluation, retention, and derivative-model uses.Consent wording, participant notice, license, rights register, and withdrawal/deletion process.The provider cannot show the permission chain or relies on implied permission for voice or biometric uses.
03Coverage and representativeness10%Speaker, locale, demographic, device, and environment quotas are defined, measurable, and reported after quality acceptance.Quota plan, recruitment method, coverage report, limitations, and post-QA counts.Headline counts cannot be reconciled with accepted unique speakers or usable hours.
04Recording and channel realism8%Capture reproduces the production channel and failure modes with controlled variation and documented equipment.Capture protocol, device/channel metadata, acoustic measurements, and representative audio.Clean microphone audio is used to represent telephony, far-field, in-car, or noisy deployment.
05Provenance and documentation10%Each asset traces to its session, rights record, protocol, transformations, and dataset version.Data card, manifest, stable IDs, lineage fields, transformation log, and version history.Files cannot be linked back to session-level provenance and permission records.
06Annotation and transcription quality10%The label policy is documented and quality is measured on a representative blinded sample using an agreed method.Annotation guide, annotator qualification, adjudication rules, gold set, and per-slice report.One accuracy claim is given without a method, sample definition, or slice-level result.
07Quality and acceptance testing10%Machine and human checks, rejection rules, pilot gates, and replacement terms are agreed before production delivery.QA plan, acceptance script, defect taxonomy, pilot report, and replacement terms.Acceptance is subjective or assessed only by the provider.
08Privacy, security, and governance8%Data flows, access, retention, incidents, subprocessors, and sensitive-data controls are documented for the project.Data-flow diagram, security controls, retention schedule, subprocessor list, and incident process.Raw voice data is moved or retained without an agreed purpose, access model, or deletion path.
09Evaluation split and leakage control8%Speaker, session, prompt, and source separation protect held-out evaluation data and are verifiable in the manifest.Split protocol, deduplication method, separation report, and non-reuse terms.The provider cannot establish whether evaluation speakers or utterances also occur in training data.
10Pilot and delivery process5%A representative pilot tests the hardest conditions and produces an auditable production and release plan.Pilot plan, milestones, escalation path, sample manifest, and change control.The pilot excludes the conditions most likely to fail in production.
11Commercial and licensing clarity4%Price units, minimums, reuse, exclusivity, territories, term, restrictions, and change costs are explicit.Priced statement of work, license schedule, assumptions, and change-order terms.A low headline price excludes required rights, QA, metadata, or replacements.
12Continuity, corrections, and support3%Named ownership, versioning, issue response, corrections, and end-of-project handover are defined.Correction SLA, release notes, version policy, and handover checklist.No process exists for tracing and correcting defects found after delivery.
Risk gates

Do not let the total hide a deal breaker

A weighted score helps rank acceptable options. It does not make an unacceptable permission chain, security model, or evaluation design acceptable. Decide which red flags stop procurement, trigger specialist review, or require a corrected pilot.

Common stop-and-investigate signals

  • Consent wording or the chain of permitted uses cannot be inspected.
  • The license does not clearly cover the intended model and retention use.
  • Delivered files cannot be traced to session and rights records.
  • Training and evaluation speaker or utterance overlap cannot be tested.
  • Security, retention, subprocessor, or deletion responsibilities are undefined.
  • A pilot sample does not represent the hardest deployment conditions.
  • Quality claims lack a metric, sample definition, or buyer acceptance test.
  • Dataset size or speaker counts cannot be reconciled after rejection and deduplication.
Decision workflow

Move from longlist to contracted delivery

  1. Define the job: write the use case, intended data uses, target slices, channel, success metrics, and risk gates.
  2. Issue one evidence request: send every provider the same specification and scorecard fields.
  3. Score the response: use 0–5 only when there is a named artifact, inspectable sample, test result, or contractual commitment.
  4. Run a representative pilot: include difficult slices and execute the buyer's acceptance tests on raw outputs and manifests.
  5. Rescore verified evidence: update provisional scores after the pilot; document disagreements and remaining assumptions.
  6. Contract the controls: carry the accepted rights, definitions, tests, reporting, replacements, corrections, and handover into the order.

Keep the completed scorecard with the procurement record. If the dataset, model purpose, vendor process, or applicable risk changes materially, repeat the relevant checks rather than treating the score as permanent.

Primary references

What informed this framework

This is Spirelight's procurement framework, not an official compliance checklist. Its emphasis on documented context, risk controls, governance, and traceability is informed by the following primary sources. Applicability depends on the project, jurisdiction, role, and intended use.

NIST AI Risk Management Framework

NIST's voluntary framework organizes AI risk work around Govern, Map, Measure, and Manage. The scorecard turns several vendor-facing evidence questions into a repeatable procurement record.

Regulation (EU) 2016/679, the GDPR

The official regulation sets rules for processing personal data, including lawful bases and special categories. The scorecard structures evidence questions for buyer review; it is not a legal determination for a particular dataset or use.

Regulation (EU) 2024/1689, the EU AI Act

The official regulation includes provisions on data and data governance for certain AI systems, with application dates for several high-risk obligations later adjusted by Regulation (EU) 2026/1744. This scorecard flags evidence areas for review; it does not determine whether a system or provider falls within a particular obligation.

Datasheets for Datasets

The paper proposes documenting dataset motivation, composition, collection, preprocessing, uses, distribution, and maintenance. Those questions inform the provenance and documentation criterion here.

Buyer questions

Frequently asked questions

What is a good speech data vendor score?

There is no universal passing score. Define deal breakers first, compare the total among otherwise viable options, and verify the highest-risk claims in a representative pilot. Two vendors with the same total can have very different risk profiles.

Can a high total compensate for missing consent evidence?

No. If missing rights, consent, provenance, security, or evaluation-separation evidence creates unacceptable risk for the intended use, treat it as a gate rather than a low-scoring tradeoff.

Should the weights change for a custom collection?

They can, but change them before bids arrive and use the same weights for every provider. A sensitive or regulated use may give more weight to rights, governance, and security; an evaluation set may give more weight to leakage controls and slice coverage.

Can the vendor complete the scorecard?

A vendor can populate the evidence and notes columns, but the buyer should own the score. Product, ML, privacy/legal, security, and procurement reviewers should score the fields relevant to them.

How should a pilot affect the score?

Treat proposal-stage scores as provisional. After testing representative audio, manifests, labels, quotas, and acceptance results, rescore the relevant criteria using the verified evidence.

Does this replace legal or security review?

No. It structures vendor comparison and evidence requests. Qualified legal, privacy, security, procurement, and AI-governance reviewers should assess the rules and risks for the specific project.

Need a scoped collection?

Send us the brief and the evidence standard

Share the languages, target speakers, channel, hours, annotations, intended use, and acceptance tests. Spirelight will review the requirements and respond with a scoped collection or annotation plan.