Imagine receiving exactly the 1,000 hours of speech you asked for, only to discover that your model still struggles with real users.
One provider could satisfy the request by asking young speakers in Copenhagen to read prepared sentences into new smartphones in quiet rooms. Another could collect spontaneous conversations from speakers across Denmark using call-centre headsets, older phones, and noisy offices.
Both deliveries might contain 1,000 hours of Danish. But only one may resemble the conditions your product will encounter.
That is the problem with defining a speech data project by volume alone. Before a provider can give you a meaningful timeline, price, or collection plan, they need to understand what the model must do, who it must work for, and the conditions in which it will be used.
A good project brief makes those requirements clear.
Start with the outcome, then the hour count
A useful brief should describe what the dataset is meant to help the product achieve.
“We need 1,000 hours of Danish and Swedish speech” gives a provider very little to work with.
Compare it with:
“We need conversational mobile-phone and headset recordings from adult Danish and Swedish speakers to improve intent recognition for customer support across the Nordic region, including interruptions, numbers, names, and office noise.”
The second request tells us much more. It identifies the use case, the markets, the type of speech, the capture channels, and several situations the model needs to handle.
The final volume still matters. But now it can be estimated against an actual product requirement rather than treated as the requirement itself.
Turn your users into recruitment requirements
A brief should make it possible to answer a practical recruitment question: Can the provider recruit the people your model needs to understand, in proportions you can justify?
For each market, define the regions, accents, and dialects that matter. Decide whether the product needs to understand native speakers, second-language speakers, or both. Include age ranges where age may affect voice characteristics or vocabulary.
If users regularly switch languages within the same conversation, include code-switching in the requirement.
These details eventually become collection quotas.
Without them, a recruitment pipeline may naturally fill the easiest profiles first. The final dataset can meet its speaker count and remain weak in the groups the product most needs to understand.
A useful brief does not need to prescribe every quota before speaking with a provider. It does need to describe the user population clearly enough for the provider to propose one.
Describe the conditions the product must handle
Recording requirements should follow the environment in which the product will operate.
A dictation product used at a desk may benefit from close-microphone recordings in quiet rooms.
Vehicle assistants typically need some bit of road noise, cabin reverberation, different seating positions, and the microphone configuration used inside the vehicle.
A customer-service model may need narrowband telephony, wireless headsets, home offices, and shared workspaces.
Specify the devices, channels, microphone distances, and environments needed for the product to function well.
If you do not yet know the exact distribution, identify the important conditions and ask the provider to recommend an appropriate mix.
The goal is to make sure the dataset covers the acoustic conditions the model is expected to handle.
Specify the behaviour you need to capture
“Collect natural speech” is almost as vague as “collect good data.”
Different products need different kinds of natural behaviour.
A wake-word model may need positive examples, hard negatives, and variations in distance and delivery.
A voice agent may need spontaneous turns, interruptions, hesitation, clarification, and recovery.
A banking assistant may need customers saying account numbers, dates, currencies, and names, including the ways people pause, restart, or correct themselves.
Natural speech still needs structure.
The collection task should give contributors a situation, goal, or reason to speak while leaving enough freedom for the relevant behaviour to occur naturally.
Besides what the contributors will talk about, your brief should describe what the recordings need to reveal.
Set transcript and annotation rules early
Audio collection and annotation are part of the same system.
If transcription rules change halfway through a project, earlier batches may need to be reviewed again. The larger the collection has become, the more expensive that change is.
Decide whether transcripts should be verbatim or cleaned. Define how to handle filler words, false starts, abbreviations, numbers, dates, currencies, code-switching, and unintelligible speech.
Then identify any additional labels the model requires.
Do you need speaker labels? Timestamps? Overlapping speech? Non-speech events? Intent? Sentiment? Emotion?
The answer depends on the product.
A meeting assistant may need diarization and overlap information. A command model may depend more heavily on intent labels and normalization rules. A TTS project may require pronunciation, prosody, and recording consistency that would matter differently in an ASR project.
If a transcript rule or label will influence training or evaluation, define it in the brief.
Define what a usable delivery looks like
A project brief should describe both what will be delivered and how your team will determine whether it is acceptable.
Start with the technical requirements: audio format, sample rate, channel layout, transcript format, and required metadata.
Then define how information such as speaker profiles, devices, environments, consent records, and quality results should appear in the delivery manifest.
Include any file-naming conventions, storage requirements, or delivery methods your engineering workflow depends on.
Acceptance criteria matter just as much.
Depending on the project, these may include transcription accuracy, label consistency, clipping and silence checks, quota completion, signal-to-noise requirements for selected conditions, or a review process for ambiguous recordings.
The important point is that the provider should understand what counts as an acceptable delivery before production begins.
Use the pilot as a decision gate
A pilot should not be a polished sample chosen simply to show that the provider can produce good recordings.
It should test the project design.
That means including difficult speaker profiles, realistic devices and environments, the intended contributor instructions, the full annotation rules, and the same metadata structure planned for the final delivery.
Your team should be able to put the pilot through the workflow that will eventually consume the full dataset.
The pilot should help answer questions such as:
- Do contributors understand the task?
- Does the speech sound appropriate for the intended use case?
- Are the transcription and annotation rules clear?
- Does the metadata explain meaningful differences between recordings?
- Can the required speaker profiles and recording conditions actually be recruited?
- What needs to change before collection scales?
Finding a weak instruction after a pilot is useful.
Finding it after hundreds of hours have already been collected is rework.
What your provider should return
The client should not have to design the entire collection workflow. Once the product requirement is understood, the provider's job is to translate the brief into a workable production plan.
That plan should explain:
- Speaker recruitment targets and how difficult profiles will be sourced.
- Recording workflows, devices, environments, and contributor instructions.
- Transcription and annotation guidelines.
- Metadata fields and delivery structure.
- Pilot design, quality gates, and review process.
- Consent and provenance records that will accompany the data.
- Key assumptions, risks, timeline, and budget options.
These are the details that turn a brief into an executable collection.
If important decisions cannot be explained before production, they are likely to surface during production, when changing them is slower and more expensive.
A good brief gives the project something to test.
At Spirelight, we work with AI teams to turn product requirements into practical collection specifications: speaker targets, recording conditions, contributor tasks, annotation rules, metadata, delivery requirements, and quality gates. See how Spirelight builds custom speech datasets.
The aim is to make the product need clear enough that the collection plan can be designed, challenged, piloted, and improved before it scales.
If you already have a rough speech data brief, send it to us. We can help turn it into a collection plan your product and engineering teams can test.
Contact Spirelight: https://www.spirelight.ai/contact
