On-site speech data collection is moderated, in-person audio capture: speakers come to a controlled environment, a trained moderator runs the session, and the rig is matched to the channel the production system will actually hear. It is how TTS voices, in-car assistants, far-field devices, and anything that needs verified speakers and consistent acoustics gets built.

This page is the operating playbook, and it is deliberately specific: who should sit in the chair and how the demographic plan is built, where the speakers come from, what the room and the rig need for each industry, why the moderator is the real quality system, what goes wrong, and what all of it costs. Location, staffing, equipment, recruitment, and schedule still have to be confirmed in writing for each project.

Who sits in the chair: the demographic plan

Before a single slot is booked, an on-site collection is specified as a speaker matrix: age bands, gender balance, dialect and region, native versus second-language speakers, and whatever attributes the deployment adds, such as driving experience or headset habits. The matrix is not paperwork. Model bias traces back to cells nobody filled, and evaluation integrity depends on knowing exactly who spoke, so the matrix is the actual product spec and the recording days are just its execution.

Size the matrix against the population the model will hear, not against a census. An in-car assistant needs licensed drivers across the regional accents of the markets the platform ships in, recorded at realistic seat positions; our automotive voice data guide covers that vertical in detail. A contact-center model needs both agent-like voices and the caller side, including the older callers who dominate many support queues. A children's education product needs tight age bands, because a nine-year-old's vocal tract is not a twelve-year-old's, and every session with a minor runs under the guardian permissions applicable law requires. Clinical documentation tools need the people who actually dictate, clinicians with their own vocabulary and pace, not general readers handed medical text.

Two mechanics separate plans that deliver from plans that slip. First, reserve evaluation speakers at recruitment time: train and test splits need disjoint speakers, and carving the held-out set out afterwards always shortchanges one side. Second, track fill per cell, never per hour. Total hours climb reassuringly while the minority cells, the specific dialect in the specific age band, sit empty, and those last cells are what set both the timeline and the price. Recruit the hard cells first and let the easy ones backfill around them.

Where the speakers come from

Every on-site plan eventually collides with the same question: who supplies the people. A standing contributor crowd is the fastest route when its footprint matches the matrix, and it comes pre-vetted from earlier projects. Fieldwork and market-research agencies add local, in-person reach for demographics a crowd does not hold, at a price per completed session. Community organizations, campuses, and clubs unlock specific dialects and age bands through trusted introductions, slowly. Targeted ads deliver raw volume with a heavy screening burden on the other end. Referrals from already-booked speakers reach inside dialect communities but narrow diversity if left uncapped.

Whatever the channel, nobody goes into the calendar unheard. A short recorded screener, one read passage and one spontaneous answer, reviewed by a native-speaker reviewer against the dialect spec, is the filter that keeps demographically perfect but linguistically wrong candidates out of the booth. The channel-by-channel mechanics, screening design, incentive structure, and the no-show arithmetic have their own page: our guide to participant recruitment for speech data collection.

What the room actually needs

The instinct is to chase perfection: the deadest room, the most expensive microphone, the cleanest waveform. That instinct is wrong for training data. What a model needs from a recording operation is repeatability: session forty sounding like session four, so the acoustic variation in your dataset comes from the speakers and the script, not from whoever set up the room that morning.

Three physical properties do most of the work. First, reverberation: aim for a dry room, and get there with heavy soft treatment. Thick absorption panels, carpet, and curtains beat glass and bare plaster every time, and a small room with hard surfaces is worse than a larger one with soft ones. Second, the noise floor, which should be low and above all stable. HVAC hum, refrigeration, and street bleed ruin more sessions than any microphone choice ever will, because a moderator can catch a mispronunciation but cannot un-bake a compressor cycling underneath an hour of audio. Third, microphone positioning that is fixed and documented, so distance and angle are the same for every speaker in every session.

Match the space to the product. A treated ordinary room covers most training-data use cases, ASR included. A vocal booth is for TTS-grade capture, where the model will clone every acoustic artifact you leave in. An anechoic chamber is a measurement instrument for testing devices, not a place to record conversation; speech captured in one sounds wrong, because no listener lives in a room with no reflections.

The rig, matched to the deployment channel

The second principle: record through a channel as close as possible to the one your production system will hear. A rig that flatters the audio can quietly mismatch the deployment target, and the mismatch only surfaces later as an error rate nobody can explain.

Use caseCapture rigPositioning and channel
TTS and read speechNear-field condenserClose, fixed distance and angle, in the quietest space available; a vocal booth for TTS-grade work
Voice assistant and far-field ASRFar-field microphone arrayDevice-realistic positions: on the counter, across the room, wherever the product will actually sit
Telephony and call center ASRTelephony loopPlay or route the audio through a real phone channel, so the recording carries the 8 kHz path production will hear
In-car voiceCabin-mounted mics or arrayMic positions fixed to the production locations, rearview or A-pillar or headliner, recorded parked and on defined drive profiles with HVAC and speed logged
Meetings and multi-partyTable-top array plus close-talk per speakerArray geometry matched to the meeting device, with one clean reference channel per participant for ground truth

Whatever the rig, capture two extras per speaker per session: a room-tone slate, meaning a stretch of the room with nobody speaking, and a fixed reference passage read by every speaker. The slate gives QA a per-session noise baseline, so hum, bleed, and gain drift show up before anyone listens to speech. The reference passage puts the same words in every voice, which makes drift in mic position, level, or room state obvious across sessions and speakers. Both cost a minute and settle arguments later.

The rig rows are where industries stop being interchangeable. Automotive sessions are scheduled around vehicles and routes, and the conditions log, speed, HVAC state, windows, is as much a deliverable as the audio; the specifics live in our in-car speech data collection guide. Far-field rooms for smart-home and assistant work need furniture, soft surfaces, and controlled noise sources, because an empty silent room is exactly what the device will never hear. Meeting-style sessions add task design that elicits natural overlap and interruption, covered in our meeting speech datasets guide. Telephony work routes audio through a real phone loop rather than simulating one. TTS studio work books the same voice across weeks, so consistency of room, rig, and speaker state becomes the whole game.

The moderator is the quality system

An in-person session without a moderator produces the same junk as an unmoderated remote one, just slower and at greater expense. The room and the rig set a ceiling on quality; the moderator is what gets you anywhere near it.

The job spans the whole session. At the door: brief the speaker, walk through consent, check the setup. During capture: run the script, listen live on headphones for clipping, mispronunciations, dropped lines, and flagging energy, and re-record failed items on the spot, because a retake costs seconds in the room and a rejection costs a whole recruited session after the fact. For spontaneous segments the job inverts: stop directing, keep the speaker talking naturally, and resist the urge to correct anything that is merely informal, because informal is the point.

The moderator also writes things down. Every deviation gets logged: skipped prompts, accent notes, the truck that passed during item 214, a mic bumped and re-set. The log is what lets QA tell a bad session from a bad room, and it is the difference between a dataset with known provenance and a pile of audio.

Getting people to actually show up

Nothing in the acoustics literature prepares you for the real failure mode of on-site collection: the speaker who does not come. A booth with a moderator and no speaker burns money at exactly the same rate as a full one, so attendance is the operations problem that decides whether the whole thing pays.

The mechanics that work are unglamorous:

  • Recruit more people than the plan needs and overbook each slot with a buffer, because some fraction always evaporates.
  • Confirm twice: once at booking, once the day before.
  • Pay per completed session, not per hour of presence, so the incentive points at finishing the script rather than occupying the chair.
  • Keep sessions under two hours. Voices tire, attention tires faster, and long sessions get cancelled more.
  • Run walk-in windows so an empty slot can be recovered by whoever is available nearby.
  • Track show-rate by recruiting channel and move spend toward the channels that deliver people, not signups.

This is where in-house recording operations quietly die. The studio gets built, the first sessions run, and then the pipeline of willing, vetted, demographically right speakers dries up while the room sits idle. A standing, vetted contributor crowd removes the entire problem, which is a large part of what you are buying when you buy collection as a service rather than as a facility.

Consent and data handling at the door

In-person collection gives you something remote collection struggles to match: certainty about who is in the recording. Use it. The person who sits down should be the person on the booking, verified by an ID check, and the first minutes of the session belong to consent, not capture.

Define the applicable lawful basis, notices, and permissions before capture; if consent is relied on, make it demonstrable and sufficiently specific for the intended AI use. A release that covers research or product improvement is not the same as one that covers training commercial models, and the difference matters when the dataset is audited years later. State the retention terms in the document the speaker signs, and record only what the spec needs: data minimization is easier to enforce at a moderated door than anywhere downstream. For the regulatory backdrop, in particular what the EU expects of training-data provenance, see our guide to speech data under the EU AI Act.

Plan capacity on usable hours, not booked hours

The calendar lies. A booked booth day contains briefings, consent, setup, breaks, re-records, QA review, and the gap where the afternoon speaker never appeared. Moderated collection trades raw throughput for usable yield on purpose: every retake in the room is a rejection that never happens downstream. The consequence is that a moderated booth day yields far less usable audio than the booked hours suggest.

So plan the schedule, the recruiting funnel, and the budget on usable hours out, not hours booked, and manage the ratio between the two as a first-class metric. Capacity plans built on calendar time come in late and over budget, and the shortfall gets discovered at delivery, which is the worst possible time. What this does to the price of an audio hour is covered in our breakdown of what speech data collection costs.

What an on-site collection costs

On-site budgets are built from line items that scale differently, which is why two quotes for the same hour count can sit far apart. The facility is either a build-out or a rental. Moderator time scales with booked days, not usable hours. Participant payments scale with completed sessions and, on-site, often include travel. Recruitment and screening overhead scales with how hard the matrix is, not with how much audio comes out. Equipment is mostly fixed once the rig is chosen. QA review and project management run through the whole schedule.

The drivers that actually move a quote are the shape of the brief: how rare the language, dialect, and demographic cells are; how many recording locations the plan needs; how long a session can run before voices tire; whether transcription and annotation are bundled; and the usable-hour ratio the operation achieves, because every line above is paid against booked time while the deliverable is measured in usable audio. Freeze the brief with the dataset specification worksheet before comparing vendors, and read what speech data collection costs for how these lines turn into a per-hour price. Published figures anywhere, ours included, are planning inputs; the written proposal confirms the pricing unit, the project minimum, and the final number.

Build it, or use one that already exists

Everything above is buildable. A treated room is weeks of work, the rigs are known quantities, and moderation is a trainable role. If speech collection is a permanent part of your product, in one location and one language, building can be the right call, and this page is most of the checklist. Deciding what to record in the first place is a separate exercise, covered in our guide to custom speech data collection.

What is much harder to build is the part that is not a facility: moderators who have run hundreds of sessions, and a recruiting pipeline that reliably puts the right speakers in the chair, in more than one language, for the months you actually need it and not a day longer. That is the case for using an operation that already exists. Spirelight can scope moderated on-site collection, recording-site requirements, and project-specific recruitment, so on-site collection arrives as a speech collection service without requiring the buyer to build the operation. How to evaluate any vendor for this, us included, is in our guide to the data collection provider checklist. If you already have a spec, submit it for a feasibility assessment.