Voice cloning is the use of AI to reproduce a specific person's voice: a text-to-speech model that does not just speak, but speaks as someone. Modern systems can approximate a voice from seconds of audio and produce broadcast-quality results from a few hours of studio recording. The technology behind brand voices, film dubbing, and voice banking for people losing their speech is the same technology behind voice deepfakes, which is why every serious voice cloning project is as much a consent and licensing exercise as a technical one.

This guide explains how voice cloning works, what separates a demo-grade clone from a production voice, the recorded data a good clone actually needs, and the consent chain a legitimate project has to be able to show.

How voice cloning works

Modern text-to-speech models learn the general mechanics of human speech from thousands of hours of many voices, then condition on a sample of the target voice to take on its identity. There are two practical modes. Zero-shot cloning feeds the model a short reference clip, seconds to a minute, and gets an immediate approximation; quality varies and the voice can drift on long passages. Custom voice building fine-tunes a model on a purpose-recorded dataset of the target speaker, which is how production brand voices and synthetic voice actors are made. The difference between the two is mostly the data.

What voice cloning is used for

Legitimate uses are established and growing: brand and product voices that stay consistent across markets, dubbing and localization that keep the original actor's voice in translated releases, voice banking for people facing conditions like ALS who record their voice before losing it, synthetic voice-over for games and media, and accessibility applications that read in a familiar voice. Every one of these depends on the speaker's documented, informed agreement, which is what separates the industry from the misuse cases that make headlines.

The data a good clone needs

Demo-grade and production-grade clones have very different appetites:

  • Zero-shot: seconds to minutes of clean speech produce a recognizable approximation, suitable for prototyping and personal use.
  • Production custom voice: typically thirty minutes to several hours of professionally recorded speech: quiet room, consistent microphone and distance, consistent tone, verbatim-accurate transcripts, and coverage of the styles the voice will need (neutral narration, questions, excitement, whispers if the product calls for them).

Recording quality is the ceiling. Noise, room reverb, level changes, and inconsistent takes all imprint on the clone. The requirements are the same as any high-grade TTS corpus, covered in our TTS training data guide: it is a small dataset, so every flaw is a large fraction of it. Timestamped, verbatim transcripts matter too; our time-coded transcripts guide covers the format side.

Consent and legality

A person's voice is legally protected in most major markets: through publicity and personality rights, through a growing set of laws aimed specifically at synthetic media and deepfakes, and through transparency obligations such as those in the EU AI Act. For a company, the practical requirement is a documented chain: the speaker knew the recordings would train a synthetic voice, agreed to the commercial uses in scope, was compensated under a clear license, and the agreement covers the territories and duration of use. Cloning a voice without that chain, whatever a tool technically permits, is a legal and reputational exposure no product team should accept. Our speech data licensing guide covers what consent for AI training has to contain, and it applies doubly here because the output is identifiable as a person.

Sourcing voices the defensible way

Teams building voice products have three clean paths: record the target speaker under a proper voice-licensing agreement, commission a purpose-built voice dataset with professional speakers who consented to synthesis, or license an existing consented voice corpus. Spirelight runs custom speech collection with exactly this paperwork: vetted speakers, studio-grade recording specs, verbatim transcripts, and consent that names synthetic voice creation as a permitted use. If you are weighing vendors, our guide to buying AI training data lists the questions that expose weak provenance quickly.