A custom speech data collection project starts when the catalogue runs out of answers: the language you need is not licensed anywhere, the acoustic environment your model will actually face is not represented in any off-the-shelf set, or the data has to be exclusive to your project rather than resold to whoever asks next. Scoping, running, and delivering that kind of project is a different exercise from buying a ready-made dataset, and it rewards a buyer who shows up with a clear brief.
This guide walks through when custom or bespoke collection is the right call, the difference between scripted, prompted, and spontaneous elicitation, the spec fields a collection brief needs, why recruitment is usually the hardest part, and why quality control has to happen during the recording session rather than after it. A working brief template is included so a reader can bring specific answers to a scoping call instead of starting from a blank page.
When off-the-shelf data is not enough
Every speech-data project starts with the same question: does something usable already exist. For a lot of common cases, the answer is yes. Standard read speech in a major language, generic command sets, and common accents in wide use can often be covered by a licensed catalogue or a marketplace dataset, faster and at lower cost than commissioning new recordings. What counts as speech data varies by use case, and it is worth checking the catalogue before assuming a project needs a bespoke run.
Custom speech datasets earn their cost when the gap is specific. A model trained on generic call-center audio will not generalize to a noisy warehouse floor, a car cabin at highway speed, or a clinical intake room, because the acoustic conditions and vocabulary do not transfer. The same is true for domain language: a voice assistant for logistics dispatch needs the jargon, hesitations, and code-switching that real dispatch calls contain, not a proxy dataset borrowed from a different industry. A missing dialect or an under-represented language is another common trigger, since catalogues concentrate on the languages with the most existing supply, which is rarely the language a new market needs. Some buyers also need exclusivity: a dataset a competitor could also license undermines the reason for collecting it in the first place. When domain, acoustic environment, language coverage, or exclusivity is the real constraint, custom collection is the tool built for it, not a nice-to-have upgrade on a catalogue purchase.
Scripted, prompted, and spontaneous collection
Custom projects are not one method; they are a choice among three elicitation styles, and the choice determines both what the data looks like and what it costs to produce.
Scripted speech data asks speakers to read fixed text aloud. It is the fastest and most predictable style to run because the linguistic content is controlled in advance: every phoneme, digit string, or command phrase a project needs can be guaranteed to appear a known number of times. It suits wake-word detection, pronunciation coverage, and any ASR model that needs balanced phonetic coverage more than it needs natural conversation.
Prompted collection gives speakers a scenario or a goal rather than a script: describe your morning routine, place this food order, resolve this billing complaint. Speakers use their own words inside a controlled situation, which produces more natural prosody and disfluency than scripted reading while keeping the topic and turn structure predictable enough to plan around.
Spontaneous speech data captures unscripted, often multi-party speech: real support calls, meetings, or open conversation. It is the closest proxy to production conditions and the only style that reliably captures interruption, overlap, and topic drift, which is why it is the right fit for conversational speech data destined for a dialogue or voice-agent model. It is also the hardest style to plan and QA, because there is no script to check a recording against.
Most collection briefs need more than one style. A voice assistant might need scripted data for command coverage and spontaneous data for the fallback conversation flow the assistant has to survive. Deciding the mix before recruitment starts is what keeps a project from having to re-record later.
How a collection project is scoped
A collection brief is really a small set of decisions written down clearly enough that a provider can plan recruitment, recording, and QA against it without a lengthy back-and-forth. The table below is the spec: fill it in before the first scoping call and the call becomes a conversation about the plan, not a requirements-gathering session.
| Spec field | What a good answer looks like | Why it matters |
|---|---|---|
| Languages and dialects | The named language plus the specific dialect or regional variant, not just "Spanish" but "Mexican Spanish, Mexico City register" | Recruitment, script design, and transcription guidelines all change by dialect; a vague language target produces audio that does not match the deployment market |
| Speaker demographics and count | Age bands, gender split, and accent mix, each with its own target number, not just a total headcount | A model trained on one demographic slice will not generalize; the count per cell is what a recruiter actually plans against |
| Hours | Total hours of usable, QA-passed audio, stated separately from raw recording hours | Raw and usable hours differ once rejected takes are removed; quoting on usable hours avoids a mismatch at delivery |
| Recording conditions and devices | The environment (quiet room, car, call center, outdoor) and the device or microphone class the deployment will actually use | A model tuned on studio audio underperforms on the noise floor and device response of production; conditions should mirror deployment, not convenience |
| Scenario or script design | The elicitation style (scripted, prompted, spontaneous) and, if scripted, the actual script or prompt bank | This determines linguistic coverage and sets expectations for how natural or how controlled the resulting speech will be |
| Annotation depth | Transcription only, or transcription plus speaker labels, timestamps, non-speech events, or emotion tags | Deeper annotation adds review time and cost; naming it upfront avoids a scope change mid-project |
| Delivery format | File type, sample rate, channel layout, and metadata schema, matched to the training pipeline that will consume the data | A dataset that does not match the expected format adds an integration delay after delivery |
| Licence and exclusivity | Whether the data can be resold or reused by the provider, and for how long any exclusivity period holds | This determines both the price and whether a competitor could end up training on the same recordings |
None of these fields need a final answer to start a conversation, but a brief with a working answer for each one lets a provider quote from it directly instead of running a separate discovery call first.
Recruiting the right speakers
Once the spec is set, recruitment is usually the part of the timeline that takes the longest, and it is where most schedule slippage on a bespoke voice dataset actually comes from. Finding a rare combination, a specific dialect crossed with an age band and a gender split, in the numbers a brief calls for, is a different problem than finding volunteers who happen to speak a language.
The language and dialect coverage a project needs is often the hardest cell to fill, precisely because it is specific. A brief asking for fifty speakers of a widely spoken language is straightforward; a brief asking for fifty speakers of one regional dialect, split evenly across three age bands, recorded in a car cabin, is a genuine recruitment project with its own sourcing and screening steps. Consent and compliance add another layer: speakers need to understand what a recording will be used for and agree to it in writing before a session starts, and that record needs to hold up under later scrutiny, not just at the moment of recording.
The practical implication for scoping is that demographic and dialect precision should be requested at the level a model actually needs, not padded for safety. A wider net costs more to recruit and rarely improves a model if the extra breadth falls outside what will be deployed.
Quality control during collection, not after
The single most consequential decision in a collection project is when QA happens, not just how thorough it is. QA that happens after delivery can only reject a recording; it cannot fix one. Once a speaker has left the booth, hung up the call, or closed the recording session, a clipped file, a wrong script read, or a microphone that was picking up background hum cannot be recovered by review. The recording either gets used as-is or gets thrown away, and a thrown-away session usually means re-recruiting the same speaker or finding a replacement, both of which cost more time than catching the problem in the moment would have.
Live QA moves the check into the session itself. A reviewer, human or automated, listens to or scans each take against the brief while the speaker is still present: is the audio level correct, is the script being read as intended, is the background noise within the agreed range, did the speaker actually produce the target phoneme or scenario. When a take fails, the fix is immediate: the speaker re-records that one line before moving on, rather than the whole session being flagged weeks later once the speaker is gone.
A rejection loop formalizes this. Each session is checked against explicit pass criteria as it happens; failed takes are re-recorded on the spot, and only sessions that clear the bar count toward the hours a client is billed for. This is also why speech data quality is best treated as a collection-time discipline rather than a delivery-time audit: the cost of a bad recording is fixed the moment the speaker walks away, and no amount of post-processing changes that afterward. A brief that specifies pass criteria before recording starts, not after, is the one that actually protects the hours a project pays for.
Timelines, delivery, and what you own
Custom collection timelines are set by recruitment and by the specificity of the spec, not by recording or transcription, which are comparatively fast once speakers are booked. A common language with a modest speaker count can move quickly; a rare dialect, a narrow demographic cell, or unusual recording conditions extend the timeline because sourcing takes longer than recording does. Building in recruitment time honestly, rather than compressing it to hit a launch date, is what keeps a project from delivering rushed, thin coverage in the cells that were hardest to fill.
Delivery should match the format decided during scoping: audio files in the agreed codec and sample rate, transcripts and annotation in the agreed schema, and metadata that ties each file back to its speaker demographics and recording conditions without a follow-up request.
Licence and exclusivity are the other half of what a buyer is actually paying for. Options generally range from a standard licence, where a provider may reuse anonymized recordings for other work, to a fully exclusive licence, where a proprietary speech dataset is used only by the commissioning buyer and never resold or reused. Exclusivity, dialect rarity, and recording conditions are the biggest drivers of what a custom project costs, more than the raw hour count; understanding how speech data licensing works before a scoping call avoids surprises at contract stage.
If a project is ready to move from scoping to execution, Spirelight's speech and audio data collection service covers recruitment, recording, transcription, and QA end to end. Spirelight is not the right choice for every case: if the need is a common language at modest scale, a catalogue purchase will usually beat a bespoke run on both cost and turnaround, and the fields above are still the right ones to check first, on this site or with anyone else running the project.