There is no sticker price for a custom voice, and any single published figure is describing one specific scope while implying it is the market. The useful thing is not a number but a model: the four lines a voice budget actually contains, what drives each one up or down, and the scope questions that explain why two quotes for the same stated deliverable can differ by a factor of five.
This guide gives that model, three scope scenarios to plan against, the build-versus-license comparison, and a worksheet for making quotes comparable. Spirelight prices custom collection from a confirmed specification rather than a shelf rate, so the figures below are planning inputs, not a quote.
Why there is no sticker price
"Three hours of studio audio" sounds like a specification and is not one. It omits every variable that actually determines cost: whether the speaker grants synthetic-voice rights or only a recording licence, whether the language is one with a deep talent pool or one where a suitable speaker has to be found and flown, whether transcripts are machine-generated or human-verified, whether re-records are included, and whether anyone has budgeted for the pronunciation work that comes after the first model.
Ask two suppliers for three hours and you may get a competent voice-over session with an AI clause bolted on, or a properly specified collection with consent, QA and a re-record provision. Both are honest quotes. They are not the same product.
The four budget lines
| Line | What it buys | What moves it |
|---|---|---|
| Talent and rights | The speaker’s time, and the synthetic-voice grant separate from the recording licence | Language and accent availability, professional vs non-professional, union status, buyout vs usage-based terms, exclusivity, territory and term |
| Recording | Studio or controlled room, engineer, direction, equipment, and the days themselves | Studio grade, location, number of days, whether the speaker travels, remote vs on-site, session supervision |
| Data preparation | Verbatim transcripts, alignment, segmentation, normalisation, QA and rejection handling | Human-verified vs machine transcripts, language, transcript accuracy threshold, how much audio gets rejected and re-recorded |
| Model work | Fine-tuning, evaluation, pronunciation fixes, deployment and hosting | In-house vs vendor, number of styles, pronunciation lexicon work, evaluation depth, iteration rounds |
Which line dominates tells you what kind of project you have. On a single-voice, single-language build, talent and recording usually lead. On a multi-market build, rights and legal review overtake them, because the work multiplies per jurisdiction rather than per hour. On an expressive, multi-style voice, data preparation leads, because each style needs its own coverage, its own QA pass and its own rejection budget.
What each line is actually driven by
Talent and rights
The largest swing factor is availability in the language and accent you need. A neutral voice in a widely spoken language has a deep pool and competitive rates. A specific regional accent in a specific age band may require weeks of targeted outreach before anyone is screened, and the recruitment effort is a cost line whether or not it appears as one. The rights structure is the second factor: a perpetual unlimited buyout is priced very differently from a three-year term with named uses, and speakers increasingly decline the former at any price.
Recording
Studio grade matters less than consistency, but consistency has a cost: the same room, engineer and chain across multiple days means booking a block rather than taking availability. Remote recording is cheaper and carries a materially higher rejection rate, which moves cost into the data preparation line rather than removing it. On-site and travel decisions are covered in our on-site speech data recording guide.
Data preparation
This is the line most often underestimated, because it is invisible in the deliverable. Human-verified verbatim transcripts cost meaningfully more than machine transcripts with a skim, and on a small dataset the difference shows directly in the model, since a two percent transcript error rate on three hours is not absorbed the way it would be on three hundred. Rejection handling belongs here too: whatever fraction of recorded audio fails QA has to be re-recorded, and somebody is paying for it.
Model work
Fine-tuning itself is usually the cheapest part, which surprises people. The cost is in what surrounds it: evaluating the voice properly, building a pronunciation lexicon for your product vocabulary, and the iteration rounds after real text hits it. Budget for at least one round of pronunciation work after launch, because it is not optional, it is just sometimes unbudgeted.
Three scope scenarios
These are scopes to plan against, not prices. Use them to work out which project you are actually running before asking for quotes.
Prototype voice
30–60 minutes, one speaker, one style, remote or light studio, machine transcripts with a human skim, no exclusivity, short term. Enough to decide whether a custom voice is worth building and to demonstrate it internally. Not enough to ship. The main risk is treating the result as representative of the production voice, which it is not.
Production brand voice
1–3 hours, one speaker, neutral plus two or three registers, controlled studio across two days, human-verified transcripts, a named synthetic-voice grant with defined uses and exclusions, a re-record provision, and a pronunciation lexicon for your product vocabulary. This is what most companies mean when they say custom voice. Budget for the re-record round as part of the project rather than as an overrun.
Multi-market voice family
Several speakers across languages, each with local rights review, a consistent recording specification reproduced in several countries, human-verified transcripts per language, and per-market disclosure obligations. Costs do not scale linearly from the single-voice case: recruitment difficulty and legal review vary sharply by market, so a per-language blended rate will be wrong in both directions. Price country by country.
Build, license, or use an API voice
The comparison is usually decided on control rather than unit cost.
| Route | Cost shape | Choose it when |
|---|---|---|
| Stock API voice | Per-character or per-minute, no build cost | The voice is not part of your identity and nobody would notice a change |
| Vendor custom voice | Build fee plus usage, hosted by the vendor | You want a distinct voice without running the pipeline, and vendor lock-in is acceptable |
| Own the data and the model | Higher build cost, lower marginal cost, you hold the rights | The voice is strategic, exclusivity matters, or your volume makes per-character pricing the dominant line |
The arguments that usually settle it are not about unit price. A voice you own cannot be deprecated by a vendor, repriced mid-contract, or offered to a competitor. Against that, owning it means owning the re-records, the pronunciation work and the consent relationship for as long as the voice is in service.
The like-for-like quote worksheet
Send the same specification to every supplier and require each of these to be answered explicitly. A quote that leaves any of them blank is not comparable with one that does not.
- Rights. Does it include a synthetic-voice grant, or only a recording licence? Territory, term, exclusivity, sublicensing, what happens at expiry.
- Usable hours. Is the quoted duration recorded audio or accepted audio after QA? These differ by a third or more.
- Transcripts. Human-verified verbatim or machine-generated? What accuracy threshold, and who checks it?
- Re-records. Is a top-up session included, and at what notice? Is the same speaker, room and chain guaranteed available, and for how long?
- Rejection. Who pays for audio that fails QA, and what is the expected rejection rate for this recording method?
- Styles. How many registers are in scope, and is each one separately specified and QA’d?
- Pronunciation. Is lexicon work for your product vocabulary included, or billed later?
- Consent evidence. Will you receive the speaker-facing consent document and the source chain?
- Delivery. Formats, sample rate, metadata, and what artefacts you keep if the relationship ends.
The costs that arrive later
- Pronunciation work after launch. Real text always contains words the script did not. Assume a round.
- Top-up sessions for new vocabulary. New products and features need new coverage, and the cost depends entirely on whether your speaker will return.
- Retraining when the base model changes. If your vendor updates the base, the voice may need refitting.
- Ongoing compensation. Usage-based terms are an operating cost, not a capital one, and the administration is real.
- Disclosure and watermarking. EU AI Act Article 50 duties have applied since August 2026 and carry an engineering cost if retrofitted.
- Re-recording from scratch. The expensive failure. It happens when the speaker is unreachable, the room is gone, or consent did not cover the use you now need.
The general speech-data cost model, covering corpus collection rather than single-voice builds, is in our what speech data actually costs guide.