On-site speech data recording is moderated, in-person audio capture: speakers come to a controlled room, a trained moderator runs the session, and the rig is matched to the channel your production system will actually hear. It is how TTS voices, far-field assistant data, and anything that needs verified speakers and consistent acoustics gets made.
This page is the physical and operational how-to: what the room needs, which rig fits which use case, why the moderator is the real quality system, and the unglamorous work of getting speakers to show up. We run this operation daily in our own studios. Publishing the method costs us nothing and tells you exactly what you are paying for, or exactly what you are signing up to build.
What the room actually needs
The instinct is to chase perfection: the deadest room, the most expensive microphone, the cleanest waveform. That instinct is wrong for training data. What a model needs from a recording operation is repeatability: session forty sounding like session four, so the acoustic variation in your dataset comes from the speakers and the script, not from whoever set up the room that morning.
Three physical properties do most of the work. First, reverberation: aim for a dry room, and get there with heavy soft treatment. Thick absorption panels, carpet, and curtains beat glass and bare plaster every time, and a small room with hard surfaces is worse than a larger one with soft ones. Second, the noise floor, which should be low and above all stable. HVAC hum, refrigeration, and street bleed ruin more sessions than any microphone choice ever will, because a moderator can catch a mispronunciation but cannot un-bake a compressor cycling underneath an hour of audio. Third, microphone positioning that is fixed and documented, so distance and angle are the same for every speaker in every session.
Match the space to the product. A treated ordinary room covers most training-data use cases, ASR included. A vocal booth is for TTS-grade capture, where the model will clone every acoustic artifact you leave in. An anechoic chamber is a measurement instrument for testing devices, not a place to record conversation; speech captured in one sounds wrong, because no listener lives in a room with no reflections.
The rig, matched to the deployment channel
The second principle: record through a channel as close as possible to the one your production system will hear. A rig that flatters the audio can quietly mismatch the deployment target, and the mismatch only surfaces later as an error rate nobody can explain.
| Use case | Capture rig | Positioning and channel |
|---|---|---|
| TTS and read speech | Near-field condenser | Close, fixed distance and angle, in the quietest space available; a vocal booth for TTS-grade work |
| Voice assistant and far-field ASR | Far-field microphone array | Device-realistic positions: on the counter, across the room, wherever the product will actually sit |
| Telephony and call center ASR | Telephony loop | Play or route the audio through a real phone channel, so the recording carries the 8 kHz path production will hear |
Whatever the rig, capture two extras per speaker per session: a room-tone slate, meaning a stretch of the room with nobody speaking, and a fixed reference passage read by every speaker. The slate gives QA a per-session noise baseline, so hum, bleed, and gain drift show up before anyone listens to speech. The reference passage puts the same words in every voice, which makes drift in mic position, level, or room state obvious across sessions and speakers. Both cost a minute and settle arguments later.
The moderator is the quality system
An in-person session without a moderator produces the same junk as an unmoderated remote one, just slower and at greater expense. The room and the rig set a ceiling on quality; the moderator is what gets you anywhere near it.
The job spans the whole session. At the door: brief the speaker, walk through consent, check the setup. During capture: run the script, listen live on headphones for clipping, mispronunciations, dropped lines, and flagging energy, and re-record failed items on the spot, because a retake costs seconds in the room and a rejection costs a whole recruited session after the fact. For spontaneous segments the job inverts: stop directing, keep the speaker talking naturally, and resist the urge to correct anything that is merely informal, because informal is the point.
The moderator also writes things down. Every deviation gets logged: skipped prompts, accent notes, the truck that passed during item 214, a mic bumped and re-set. The log is what lets QA tell a bad session from a bad room, and it is the difference between a dataset with known provenance and a pile of audio.
Getting people to actually show up
Nothing in the acoustics literature prepares you for the real failure mode of on-site collection: the speaker who does not come. A booth with a moderator and no speaker burns money at exactly the same rate as a full one, so attendance is the operations problem that decides whether the whole thing pays.
The mechanics that work are unglamorous:
- Recruit more people than the plan needs and overbook each slot with a buffer, because some fraction always evaporates.
- Confirm twice: once at booking, once the day before.
- Pay per completed session, not per hour of presence, so the incentive points at finishing the script rather than occupying the chair.
- Keep sessions under two hours. Voices tire, attention tires faster, and long sessions get cancelled more.
- Run walk-in windows so an empty slot can be recovered by whoever is available nearby.
- Track show-rate by recruiting channel and move spend toward the channels that deliver people, not signups.
This is where in-house recording operations quietly die. The studio gets built, the first sessions run, and then the pipeline of willing, vetted, demographically right speakers dries up while the room sits idle. A standing, vetted contributor crowd removes the entire problem, which is a large part of what you are buying when you buy collection as a service rather than as a facility.
Consent and data handling at the door
In-person collection gives you something remote collection struggles to match: certainty about who is in the recording. Use it. The person who sits down should be the person on the booking, verified by an ID check, and the first minutes of the session belong to consent, not capture.
Consent should be written, and it should name AI training explicitly. A release that covers research or product improvement is not the same as one that covers training commercial models, and the difference matters when the dataset is audited years later. State the retention terms in the document the speaker signs, and record only what the spec needs: data minimization is easier to enforce at a moderated door than anywhere downstream. For the regulatory backdrop, in particular what the EU expects of training-data provenance, see our guide to speech data under the EU AI Act.
Plan capacity on usable hours, not booked hours
The calendar lies. A booked booth day contains briefings, consent, setup, breaks, re-records, QA review, and the gap where the afternoon speaker never appeared. Moderated collection trades raw throughput for usable yield on purpose: every retake in the room is a rejection that never happens downstream. The consequence is that a moderated booth day yields far less usable audio than the booked hours suggest.
So plan the schedule, the recruiting funnel, and the budget on usable hours out, not hours booked, and manage the ratio between the two as a first-class metric. Capacity plans built on calendar time come in late and over budget, and the shortfall gets discovered at delivery, which is the worst possible time. What this does to the price of an audio hour is covered in our breakdown of what speech data collection costs.
Build it, or use one that already exists
Everything above is buildable. A treated room is weeks of work, the rigs are known quantities, and moderation is a trainable role. If speech collection is a permanent part of your product, in one location and one language, building can be the right call, and this page is most of the checklist. Deciding what to record in the first place is a separate exercise, covered in our guide to custom speech data collection.
What is much harder to build is the part that is not a facility: moderators who have run hundreds of sessions, and a recruiting pipeline that reliably puts the right speakers in the chair, in more than one language, for the months you actually need it and not a day longer. That is the case for using an operation that already exists. Spirelight runs owned studios with moderators on staff and a standing multilingual contributor crowd, so on-site collection arrives as a speech collection service with none of the buildout. How to evaluate any vendor for this, us included, is in our guide to the data collection provider checklist. If you already have a spec, send it over and we will scope it.