Recruiting voice donors for a cloning project looks superficially like recruiting participants for a speech dataset, and the two are almost opposite exercises. A speech corpus wants many speakers, each contributing a little, and diversity across them is the deliverable. A cloned voice wants one speaker, or a small handful, each contributing a lot, and consistency within them is the deliverable. One bad speaker in a 500-person corpus is noise you can drop. One bad speaker in a cloning project is the project.

This guide covers where voice donors actually come from, how to screen them in a way that predicts how the clone will sound, what a recording day realistically yields, the terms that keep a speaker willing to return, and the mistakes that cost projects a re-record. Our participant recruitment for speech data collection guide covers the many-speaker case.

Voice donor, voice actor, data contributor: three different deals

The terms get used interchangeably and the commercial arrangements behind them are not interchangeable at all.

What they provideTypical termsBest for
Voice donorTheir voice as an identity, knowingly, for a model that will speak as themNamed synthetic-voice grant, defined uses and exclusions, often ongoing compensationBrand voices, product voices, anything customer-facing
Professional voice actorTrained, directable delivery and stamina across long sessionsSession fee plus a separately negotiated replica grant; union rules may applyExpressive and multi-style voices, audiobook and character work
Data contributorOne voice among many in a corpusPer-session payment, broad dataset licence, no individual identity attachedASR and base TTS corpora, not cloning

The distinction that causes real trouble is the second row. Hiring a voice actor through a normal booking process gets you a recording licence, and the synthetic-voice grant is a separate negotiation that has to happen before the session, not after. Raising it afterwards means either paying a premium under time pressure or discarding the audio.

Where voice donors come from

No single channel is right for every brief, and the choice is mostly driven by how specific your requirements are.

Voice-over marketplaces and agencies

Fast, deep, and full of people who can actually sustain a four-hour session. The catch is that a demo reel is a produced artefact: compressed, EQ’d, edited, sometimes years old. It tells you what someone sounds like at their best in a genre, not what they sound like reading your product copy for three hours. Always run your own screen. Expect to negotiate AI terms explicitly; many performers and agents have firm positions, and finding out early is the point.

An existing contributor crowd

If you already run speech data collection, you have people who have recorded before, have been paid reliably, and whose audio you can listen to today rather than requesting a screen. The screening shortcut is substantial. The limit is coverage: a crowd built for corpus collection will be thin in exactly the specific registers a brand voice needs.

Targeted outreach for a specific profile

When the brief is a particular regional accent in a particular age band, general channels produce volume and almost no matches. Community organisations, regional radio, local theatre and university language departments reach people who are not on marketplaces. Slow to start, and the only reliable route for genuinely specific requirements.

Inside your own company

Occasionally right, and it needs care. The person is available, invested and cheap. But consent between an employer and an employee deserves scrutiny, the arrangement has to survive them leaving, and an untrained speaker may not sustain the sessions. If you do it, use the same paperwork and the same fee you would offer an external donor.

Screening: listen, do not read

The screen is the highest-leverage step in the whole process, and it is cheap. Three stages.

1. The recorded screen, using your copy

Send candidates 90 seconds of your actual product copy, not a generic passage. Include the awkward parts deliberately: your product names, a number-heavy sentence, a question, and one line that has to sound apologetic. Ask for a raw file with no processing, which is both a sample and a test of whether they can follow a technical instruction.

Judge it on: does the timbre suit the product; is the read consistent across the 90 seconds; is the delivery natural or announcer-ish; did they follow the spec; and does the raw audio sound clean. An unprocessed file from a good room sounds unglamorous, and that is what you want.

2. The instant-clone check

This step is underused and it is nearly free. Take the screen audio, make a zero-shot clone, and generate your ten most common product sentences. It will not sound like the final voice. It will reveal whether the timbre survives synthesis, because some voices that sound wonderful in the room clone poorly: very breathy voices, very deep voices, and voices with strong vocal fry often lose their character or introduce artefacts. Finding that out from a 90-second screen rather than after a recording day is the entire value.

3. A paid trial session

Before committing to a multi-day booking, pay for one short session in the real room with the real chain. You learn whether they hold consistency over an hour, whether they take direction, whether the room suits the voice, and whether they turn up on time. Pay properly for it regardless of outcome.

Things that disqualify a voice for cloning specifically, as opposed to for voice-over generally: audible sibilance that survives de-essing, inconsistent pacing across takes, a strong tendency to perform rather than speak, and any condition that changes the voice week to week.

The consent conversation, held early

Raise it in the first substantive conversation, before the screen if possible. Some people will not agree to synthetic-voice terms at any price, and that is a legitimate position you want surfaced at the top of the funnel rather than after a trial session.

What to be plain about: a model will be trained to generate new speech in their voice; it will say things they never recorded; here is where it will be used and here is what it will never be used for; here is the term and what happens at the end; here is what they can withdraw and what cannot be reversed. Speakers who understand the deal are the ones who stay willing to come back. Speakers who feel they were finessed generate disputes later, and the full clause list is in the voice cloning laws and consent guide.

The prohibited-uses clause is worth real attention here. It costs you very little to exclude political content, adult content and endorsements, and it converts reluctant candidates more effectively than a higher fee.

What a recording day actually yields

Plan on yield, not booth hours. A 6-hour booked day typically produces 4–5 hours in front of the microphone and 2–3 hours of usable audio after retakes, breaks and QA rejection, and the ratio degrades as the day runs long because the voice tires and consistency slips.

So a 1–3 hour production voice is generally one to two well-run days, not an afternoon. Some practical constraints worth building into the schedule:

  • Split long collections across days, not into long days. Hour five is where drift and fatigue show up in the audio.
  • Record the reference take first. A short calibration passage at the start of every session, used to match level, tone and energy on later days.
  • Do not move anything. Same room, same microphone, same preamp, same distance, same chair. Photograph the setup on day one and reproduce it.
  • Hold back the hardest material. Emotional range and long-form passages belong early in a session, not at hour four.
  • QA during the session. An engineer listening live catches the drift that a next-day review can only fix by rebooking.

Pay and terms that keep a speaker

Rates vary far too much by market, experience and union status for a single published figure to be useful, so the structure matters more than the number. Three models are common, and they behave differently.

ModelHow it worksTrade-off
Session fee plus buyoutPaid for recording days, plus a one-time fee for the synthetic-voice grantSimple and predictable. Speakers increasingly resist unlimited perpetual buyouts.
Session fee plus usageRecording fee, plus ongoing payment tied to use, term or renewalEasier to agree and much better for retention. More administration.
RetainerOngoing relationship including re-records and refreshesBest when the voice is strategic and you know you will need updates.

Whichever model, the operational point is the same: you will need this person again. New product lines need new vocabulary, models get retrained, coverage gaps surface in production. A speaker who felt well treated returns to the same room for a top-up session; one who felt exploited does not, and then you are choosing between a visibly mismatched voice and rebuilding from scratch. Pay on time, name them as a collaborator if they want that, tell them where the voice is being used, and honour the exclusions.

Common mistakes

  • Casting from demo reels. Produced audio in a flattering genre predicts very little about three hours of your product copy.
  • Leaving AI terms until after the session. The leverage is gone and so is the goodwill.
  • Booking one long day to save money. The last two hours are usually unusable, so it costs more.
  • Not testing the clone before committing. Some excellent voices synthesise poorly, and 90 seconds of screen audio tells you.
  • Recording without a script review. Missing product names and number formats are the failures users hear, and they are free to prevent beforehand.
  • Treating the speaker as a supplier rather than a collaborator. Retention is an engineering requirement disguised as a relationship.