Custom anti-spoofing collection

Genuine and spoofed speech pairs from speakers who consented to the spoof

Commission voice anti-spoofing data as one specification: genuine speech and its spoofed counterpart from the same consented speakers, matched across prompts and channels, so the detector learns the artifact instead of the recording setup.

Free sample: Name the languages, channels, and attack types you are defending against. The team prepares a matched genuine and spoofed excerpt and sends it manually within 48 hours, free of charge.

Collection specification

Genuine capture
Telephony, VoIP, and close-mic recordings from the same speakers, matched to your deployment channel
Spoof families
TTS and voice-conversion synthesis, replay re-recording through real devices and rooms, and channel injection, scoped per project
Pairing
Every spoofed utterance keyed to a genuine utterance from the same speaker, with matched prompts and channel conditions
Languages
Speakers recruited per language and accent to your deployment matrix, not inherited from English benchmarks
Metadata
Per-file labels for speaker, attack type, generation system, device chain, and channel, delivered as JSON or CSV
Rights
Use-specific cloning and biometric consent per speaker, naming spoof generation for detector training as the purpose
Pricing
Custom, scoped to your conditions

One specification for both sides of the pair

A detector trained on genuine speech from one corpus and spoofed speech from another learns to separate corpora, not attacks. Public anti-spoofing benchmarks repeat the problem at scale: they lean heavily on English, studio-grade recordings, and a fixed list of known attack systems, so models that top a leaderboard stumble on the first unfamiliar generator. Production fraud does not announce its generator, and it arrives over telephony.

Here the pair is the unit of collection. Every spoofed utterance is keyed to a genuine utterance from the same speaker, reading the same prompt over the same channel, so the only difference left for the detector to model is the spoof itself. That matching is what makes error analysis say something useful: when the detector fails, you know it failed on the artifact, not on a microphone.

An attack surface scoped to your threat model

The spoof side is scoped per project rather than fixed. TTS and voice-conversion spoofs are generated from the consented genuine recordings, covering the synthesis families your synthetic-voice detector has to catch. Replay attacks are captured physically: genuine recordings played back through real loudspeakers and handsets and re-recorded in real rooms, which is the presentation attack a microphone actually sees.

Channels get the same treatment, because a fraud call reaches the bank through a codec chain, not a studio. Genuine and spoofed material can both be passed through live telephony and VoIP paths, or injected directly at the channel, so the detector meets spoofs the way production traffic delivers them. Languages are recruited to the deployment, not inherited from an English benchmark.

Consent that names the spoof, splits that keep evaluation honest

Every spoofed voice in the corpus belongs to a speaker who agreed to exactly that. Participants sign use-specific cloning and biometric consent that names spoof generation for detector training and evaluation as the purpose, in plain language, and the signed chain is delivered with the data. We do not clone a voice whose owner has not signed, for any customer or any reason.

The same discipline applies to evaluation. Speakers are partitioned into held-out splits before any spoof is generated, so no test speaker's genuine or spoofed audio ever appears in training, and the reported numbers describe generalization rather than memorization. That is the difference between a benchmark score and a detector you can put in front of fraud.

This offer is not the right fit when

  • You want a specific real person's voice cloned without their signed, use-specific consent, whatever the stated purpose.
  • You need spoofed audio for impersonation, social engineering exercises, or anything other than training and evaluating detectors.
  • You want a corpus to download today rather than a collection scoped to your channels, languages, and attack surface.

Get a free paired sample

  1. Your details
  2. Your project
  3. Verify

Three short steps. The team then follows up manually within 48 hours, and confirms volume, rights, QA, and delivery if you want the scope priced.

Would rather talk it through first? Scope an anti-spoofing collection

Prepare the brief

Bring a specification your vendors can price consistently

Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.

Build a dataset specification

Buyer documentation

Specimen data card and provenance structure, evaluation brief template, and acceptance and held-out questions.

Related services

Guides for this decision

Need a scoped collection plan? Send the language, channel, volume, annotation, and timing you know. The quote form opens directly in project-brief mode.