In-vehicle speech data collection means recording people speaking inside a car, under the acoustic conditions the deployed model will actually face: road and wind noise, HVAC, overlapping passengers, and the raised, strained voice speakers naturally produce at speed. There are three ways to run it, and they differ sharply in realism, cost, and legal overhead.

Most automotive data collection services treat the method as a black box: you describe your assistant, they return a quote. This page opens the box. What follows is how we run in-car speech data collection at Spirelight: the three setups and when each is defensible, the microphone rig, the condition matrix that actually drives cost, consent inside a moving vehicle, and what a finished delivery should contain.

Three ways to collect in-cabin speech

Every in-vehicle project uses one of three setups, or a mix of them. They trade realism against cost and repeatability, and choosing wrong is the expensive mistake in this category, because the flaws only surface after you have trained on the data.

SetupRealismCostWhat you loseRight when
Real fleet on real roadsHighest: real road texture, vibration, wind, genuine Lombard speechHighest: instrumented vehicles, safety protocol, consent overheadRepeatability; a failed session costs a day, not a retakeProduction training and evaluation sets, multi-zone systems
Stationary vehicle plus noise playbackReal cabin acoustics, recorded road noise injected through calibrated speakersMid: one parked vehicle, no safety overheadStructure-borne vibration and most of the real Lombard effectVolume collection once the hard cells are covered by fleet hours
Studio booth plus cabin impulse responsesSimulated: clean speech convolved with measured cabin IRs, noise mixed in afterwardsLowestEverything acoustic is syntheticEarly prototyping only

A real fleet is the only setup that captures what a cabin actually does to speech: vibration through the seats and mic mounts, road texture changing under the tires, and drivers whose attention is genuinely on the road. It also carries the full weight of safety protocol, insurance, and per-occupant consent, which is why it costs what it costs.

A stationary vehicle with calibrated noise playback keeps the real cabin geometry and mic positions, and it is repeatable: the same noise bed for every speaker, and a failed prompt rerun in minutes. What it loses is structure-borne vibration and most of the genuine Lombard effect, because people do not strain their voices against playback the way they do against real road noise.

A studio booth with measured cabin impulse responses and noise mixing is the cheapest option, but everything acoustic about it is synthetic. It is defensible for early prototyping and little after that. Most production projects mix setups: fleet hours for evaluation sets and the hardest conditions, stationary hours for volume. The general mechanics of scoping any custom project, recruiting, pilot sessions, and QA gates, are covered in our guide to custom speech data collection; everything below is what changes when the recording room has wheels.

The microphone rig

The standard rig has two layers. Each occupant wears a near-field reference microphone, a headset or lapel mic, that captures their speech cleanly regardless of what the cabin is doing. A far-field array is then mounted at positions that match production hardware.

Position matters more than microphone quality. The AISHELL-5 corpus, recorded in a real electric vehicle, mounts its far-field microphones above the door handles of all four doors, per the AISHELL-5 paper. Production voice assistants more commonly place microphones at the A-pillar, in the rearview mirror housing, or in the headliner. Data recorded at the mirror does not transfer for free to a headliner array: the distances, reflections, and noise pickup all change. Ask where your target vehicle puts its mics, and put the array there.

Three engineering requirements are non-negotiable. All channels record at matched sample rates. Capture is sample-synchronized across channels through a single clock, because beamforming and source separation both assume aligned channels. And gain is staged per channel, so the driver channel does not clip when a truck passes and the quiet rear passenger is not sitting at the noise floor. The near-field channels double as the transcription reference and as the clean signal for computing per-channel SNR later.

Scoping the driving-condition matrix

The condition matrix is where in-vehicle collection budgets are won or lost. The usual dimensions: speed tier (parked, urban, highway), road surface, windows open or closed, HVAC level, rain or dry, EV or combustion drivetrain, and passenger configuration from driver-only to a full cabin.

Each dimension multiplies the others. Three speed tiers, two window states, three HVAC levels, two weather states, and three occupancy patterns is already 108 cells before you touch road surface or drivetrain, and every cell costs recording hours. Nobody should record the full matrix.

Scope from the deployment instead. If the assistant ships in an EV, the cabin is much quieter at low speed and wind noise dominates earlier, so many combustion cells may not matter at all. If field telemetry says your users rarely open windows, one token cell covers it. The right approach is to rank cells by how often the model will see them and how badly it currently fails in them, then weight recording hours toward that ranking rather than spreading them evenly. A matrix scoped this way routinely comes in at a fraction of the naive version with no measurable loss in model quality.

Session design: scripted, prompted, spontaneous

Within each condition cell, sessions layer three kinds of speech. Scripted commands cover the wake word plus a command inventory spanning your intents: navigation, media, climate, telephony. Each speaker reads controlled variants, so the model sees the same intent phrased many ways in many voices.

Prompted tasks give the speaker a goal rather than a sentence: find a charging station, make it warmer in the back, call the last missed number. This elicits the hesitations, restarts, and phrasing that scripted reading suppresses, which is closer to what the deployed model will hear.

Spontaneous conversation between occupants fills the third layer. It is what the assistant must ignore, reject, or attribute to the right person, and it is the raw material for barge-in and diarization robustness.

One rule about the Lombard effect: at highway speed, speakers raise their voices, shift pitch upward, and flatten their spectral tilt without being told to. Do not ask people to talk loudly in a quiet cabin to simulate this. Faked Lombard speech has the wrong acoustics, and a model trained on it learns the wrong compensation. Record at real speed and the effect arrives on its own.

Per-seat capture for multi-zone systems

Four-zone recognition systems need to know not just what was said but which seat said it: a rear passenger asking for more heat should not change the driver's settings. Collections for these systems need per-seat reference microphones, deliberate overlap in the session design, and transcripts attributed to seats, not just to speakers.

Deliberate overlap means prompting two occupants to speak at once, and capturing rear-cabin conversation running underneath driver commands. Overlap is expensive to transcribe accurately, which is why it is often quietly missing from cheaper collections, and why single-speaker in-car data does so little for a multi-zone product. Seat attribution is part of our standard automotive delivery; the wider collection and annotation offering is on our services page.

Consent and GDPR inside the cabin

Voice recordings are personal data under the GDPR, and when they are processed to identify a person they become biometric special-category data under Article 9, which requires explicit consent. A car full of occupants is not a gray area: every person in the vehicle gives explicit consent before recording starts. A sticker on the dashboard or a line in a rental agreement is not consent.

Consent has to state the purpose (training and evaluating speech recognition systems), the retention period, and who the data can be licensed to. A passenger who joins mid-session consents before the session resumes; recording cannot cover them retroactively. In practice this means consent capture built into the session tooling, with each record tied to a session ID and a seat.

This paper trail is not just compliance hygiene. Buyers deploying in the EU need to document training data provenance for AI Act technical documentation, and per-occupant consent records are exactly the evidence that requirement calls for; our guide to the EU AI Act and speech data covers what buyers are asked to show. A vendor that cannot produce this trail is selling a liability with a dataset attached.

What a delivery should contain

A finished in-vehicle dataset is more than audio files. The delivery should include:

  • Multi-channel WAV with a channel map documenting which channel is which microphone, at which position, in which vehicle.
  • Time-aligned transcripts produced against the near-field reference channels, so far-field error rates can be measured honestly.
  • Speaker and seat labels on every utterance, including overlap regions.
  • Condition metadata per utterance: speed tier, surface, windows, HVAC, weather, drivetrain, occupancy.
  • SNR statistics per session and per channel, computed against the near-field reference.

The last two are the ones buyers forget to demand and then miss most. Per-utterance condition metadata is what lets you slice error rates by cell and find out where the model actually fails. SNR statistics let you verify that the vendor delivered the noise conditions you ordered rather than sixty hours of a parked car with the engine off.

When not to commission an in-cabin collection

An honest scoping conversation starts with whether you need this at all. If the goal is a robustness fine-tune of an existing model, you can often get most of the way with clean conversational speech, recorded cabin noise mixed in at controlled SNRs, and measured impulse responses, at a fraction of the cost of a fleet. Train, measure against a small real in-cabin evaluation set, and only commission a full collection if the remaining gap justifies it.

Commissioning is the right call when augmentation cannot manufacture what you need: seat-attributed multi-zone data, audio at your exact production mic geometry, genuine Lombard speech, or an evaluation set that must reflect reality rather than a simulation of it. Before scoping anything, check what already exists: our guide to automotive speech datasets covers the open and commercial corpora, and the broader automotive voice data guide maps the whole domain. If you are not sure which side of the line your project is on, talk to our team: describing the deployment takes ten minutes and sometimes ends with us recommending the cheaper route.