Voice cloning uses AI to reproduce characteristics of a specific person's voice. Systems can condition on short reference audio or use a purpose-recorded dataset for a custom voice. The same technical capability can support brand voices, dubbing, voice banking, accessibility, or harmful impersonation, so authorization, rights, privacy, compensation, security, and transparency require project-specific review.

This guide explains the technical approaches, recording inputs, and diligence questions. It is not legal advice; laws and rights vary by person, jurisdiction, context, and intended use.

How voice cloning works

Modern text-to-speech models learn the general mechanics of human speech from thousands of hours of many voices, then condition on a sample of the target voice to take on its identity. There are two practical modes. Zero-shot cloning feeds the model a short reference clip, seconds to a minute, and gets an immediate approximation; quality varies and the voice can drift on long passages. Custom voice building fine-tunes a model on a purpose-recorded dataset of the target speaker, which is how production brand voices and synthetic voice actors are made. The difference between the two is mostly the data.

What voice cloning is used for

Legitimate uses are established and growing: brand and product voices that stay consistent across markets, dubbing and localization that keep the original actor's voice in translated releases, voice banking for people facing conditions like ALS who record their voice before losing it, synthetic voice-over for games and media, and accessibility applications that read in a familiar voice. For these uses, obtain project-specific legal review and document the speaker authorization, rights, notices, or lawful basis the facts require.

The data a good clone needs

Demo-grade and production-grade clones have very different appetites:

  • Zero-shot: seconds to minutes of clean speech produce a recognizable approximation, suitable for prototyping and personal use.
  • Production custom voice: typically thirty minutes to several hours of professionally recorded speech: quiet room, consistent microphone and distance, consistent tone, verbatim-accurate transcripts, and coverage of the styles the voice will need (neutral narration, questions, excitement, whispers if the product calls for them).

Recording quality is the ceiling. Noise, room reverb, level changes, and inconsistent takes all imprint on the clone. The requirements are the same as any high-grade TTS corpus, covered in our TTS training data guide: it is a small dataset, so every flaw is a large fraction of it. Timestamped, verbatim transcripts matter too; our time-coded transcripts guide covers the format side.

Authorization, rights, and legal review

Voice cloning can implicate contract, copyright, performer, publicity or personality, privacy, biometric, consumer-protection, synthetic-media, and transparency rules depending on the facts and jurisdiction. Do not assume one consent form resolves every issue. Have counsel identify the people and rights involved, intended outputs and uses, territories, term, compensation, withdrawal or termination, model and recording retention, sublicensing, security, disclosure, and prohibited uses.

If consent is relied on for any processing, verify that it is valid and sufficiently specific for that processing. A recording agreement does not automatically grant synthetic-voice rights, and later paperwork does not by itself validate earlier processing. The speech data licensing guide provides contract and records questions.

Sourcing a voice

Possible routes include an agreement with the target speaker, a commissioned purpose-recorded dataset, or a licensed corpus. None is automatically legally sufficient. Verify speaker authorization, source chain, permissions, license scope, compensation, privacy basis, security, permitted and prohibited uses, territories, term, retention, sublicensing, disclosure, and termination for the exact project.

Spirelight can assess a synthetic-voice collection brief, subject to speaker agreement, permissions, compensation, privacy and security requirements, and operational feasibility. Recording specification, transcripts, rights, artifacts, schedule, and price are confirmed in the written scope. The buying AI training data guide provides a wider diligence checklist.