In-vehicle speech data collection means recording people speaking inside a car, under the acoustic conditions the deployed model will actually face: road and wind noise, HVAC, overlapping passengers, and the raised, strained voice speakers naturally produce at speed. Three common setups make different tradeoffs in acoustic realism, repeatability, safety, privacy, and cost.

This page provides a procurement framework: compare the recording setup, microphone rig, condition matrix, occupant and privacy controls, and required delivery evidence. Spirelight can assess an in-car brief; site, fleet, equipment, recruitment, data-protection requirements, deliverables, schedule, and price require written confirmation.

Three ways to collect in-cabin speech

Every in-vehicle project uses one of three setups, or a mix of them. They trade realism against cost and repeatability, and choosing wrong is the expensive mistake in this category, because the flaws only surface after you have trained on the data.

SetupRealismCostWhat you loseRight when
Real fleet on real roadsHighest: real road texture, vibration, wind, genuine Lombard speechHighest: instrumented vehicles, safety protocol, consent overheadRepeatability; a failed session costs a day, not a retakeProduction training and evaluation sets, multi-zone systems
Stationary vehicle plus noise playbackReal cabin acoustics, recorded road noise injected through calibrated speakersMid: one parked vehicle, no safety overheadStructure-borne vibration and most of the real Lombard effectVolume collection once the hard cells are covered by fleet hours
Studio booth plus cabin impulse responsesSimulated: clean speech convolved with measured cabin IRs, noise mixed in afterwardsLowestEverything acoustic is syntheticEarly prototyping only

A real-road fleet can capture vibration, changing road texture, and naturally task-focused driver speech that a stationary or simulated setup may not reproduce. It can also add safety, insurance, permission, privacy, and scheduling requirements that must be costed in the project plan.

A stationary vehicle with calibrated noise playback keeps the real cabin geometry and mic positions, and it is repeatable: the same noise bed for every speaker, and a failed prompt rerun in minutes. What it loses is structure-borne vibration and most of the genuine Lombard effect, because people do not strain their voices against playback the way they do against real road noise.

A studio booth with measured cabin impulse responses and noise mixing is a simulated option whose suitability must be tested. A hybrid may combine real-road evaluation cells with stationary or simulated volume, but the mix should follow a held-out benchmark rather than a category rule. The general mechanics of scoping any custom project, recruiting, pilot sessions, and QA gates, are covered in our guide to custom speech data collection; everything below is what changes when the recording room has wheels.

The microphone rig

One possible rig has two layers. Each occupant wears a near-field reference microphone, a headset or lapel mic, that captures their speech cleanly regardless of what the cabin is doing. A far-field array is then mounted at positions that match production hardware.

Position matters more than microphone quality. The AISHELL-5 corpus, recorded in a real electric vehicle, mounts its far-field microphones above the door handles of all four doors, per the AISHELL-5 paper. Production voice assistants more commonly place microphones at the A-pillar, in the rearview mirror housing, or in the headliner. Data recorded at the mirror does not transfer for free to a headliner array: the distances, reflections, and noise pickup all change. Ask where your target vehicle puts its mics, and put the array there.

Write three engineering requirements into the acceptance specification when beamforming, separation, and channel comparison depend on them. All channels record at matched sample rates. Capture is sample-synchronized across channels through a single clock, because beamforming and source separation both assume aligned channels. And gain is staged per channel, so the driver channel does not clip when a truck passes and the quiet rear passenger is not sitting at the noise floor. The near-field channels double as the transcription reference and as the clean signal for computing per-channel SNR later.

Scoping the driving-condition matrix

The condition matrix is where in-vehicle collection budgets are won or lost. The usual dimensions: speed tier (parked, urban, highway), road surface, windows open or closed, HVAC level, rain or dry, EV or combustion drivetrain, and passenger configuration from driver-only to a full cabin.

Each dimension multiplies the others. Three speed tiers, two window states, three HVAC levels, two weather states, and three occupancy patterns is already 108 cells before you touch road surface or drivetrain, and every cell costs recording hours. Recording every combination is rarely an efficient default; justify each cell against deployment frequency, measured risk, and the evaluation plan.

Scope from the deployment instead. If the assistant ships in an EV, the cabin is much quieter at low speed and wind noise dominates earlier, so many combustion cells may not matter at all. If field telemetry says your users rarely open windows, one token cell covers it. The right approach is to rank cells by how often the model will see them and how badly it currently fails in them, then weight recording hours toward that ranking rather than spreading them evenly. A weighted matrix can remove low-priority cells; validate the reduced matrix on a held-out set before scaling.

Session design: scripted, prompted, spontaneous

Within each condition cell, sessions layer three kinds of speech. Scripted commands cover the wake word plus a command inventory spanning your intents: navigation, media, climate, telephony. Each speaker reads controlled variants, so the model sees the same intent phrased many ways in many voices.

Prompted tasks give the speaker a goal rather than a sentence: find a charging station, make it warmer in the back, call the last missed number. This elicits the hesitations, restarts, and phrasing that scripted reading suppresses, which is closer to what the deployed model will hear.

Spontaneous conversation between occupants fills the third layer. It is what the assistant must ignore, reject, or attribute to the right person, and it is the raw material for barge-in and diarization robustness.

One rule about the Lombard effect: at highway speed, speakers raise their voices, shift pitch upward, and flatten their spectral tilt without being told to. Do not ask people to talk loudly in a quiet cabin to simulate this. Prompted loud speech and naturally noise-induced speech are not equivalent. If Lombard conditions matter, include a representative real-noise evaluation cell and validate whether the proposed capture method is adequate.

Per-seat capture for multi-zone systems

Four-zone recognition systems need to know not just what was said but which seat said it: a rear passenger asking for more heat should not change the driver's settings. Collections for these systems need per-seat reference microphones, deliberate overlap in the session design, and transcripts attributed to seats, not just to speakers.

Deliberate overlap means prompting two occupants to speak at once, and capturing rear-cabin conversation running underneath driver commands. Overlap adds transcription and adjudication work. Confirm the required overlap share, seat labels, reviewer method, and acceptance threshold rather than inferring them from price or total hours. Seat attribution can be included when agreed in the collection specification; the wider collection and annotation offering is on our services page.

Occupant, recording-law, and GDPR diligence

Recorded voices can be personal data under the GDPR. They are biometric special-category data under Article 9 when processed through specific technical means to uniquely identify a person; that classification and the applicable Article 9 condition are project-specific. Before recording, assess controller and processor roles, lawful basis, notices, local recording law, occupant permissions, purpose, retention, transfers, security, and withdrawal or objection handling.

If consent is relied on, it must meet the applicable requirements and cover the relevant processing; signage alone does not establish valid consent. Define how late-joining occupants are handled, when recording pauses, and how each applicable record links to a session without collecting unnecessary metadata.

EU AI Act documentation duties depend on the system, role, and classification. Ask for the source, permission, provenance, and data-governance evidence applicable to the project. Without those records, the buyer cannot complete its diligence; see the EU AI Act and speech data for the scoped assessment.

What a delivery should contain

A finished in-vehicle dataset is more than audio files. The delivery should include:

  • Multi-channel WAV with a channel map documenting which channel is which microphone, at which position, in which vehicle.
  • Time-aligned transcripts produced against the near-field reference channels, so far-field error rates can be measured honestly.
  • Speaker and seat labels on every utterance, including overlap regions.
  • Condition metadata per utterance: speed tier, surface, windows, HVAC, weather, drivetrain, occupancy.
  • SNR statistics per session and per channel, computed against the near-field reference.

The last two are the ones buyers forget to demand and then miss most. Per-utterance condition metadata is what lets you slice error rates by cell and find out where the model actually fails. SNR statistics and condition logs help test whether the delivered cells meet the written acceptance criteria.

When not to commission an in-cabin collection

An honest scoping conversation starts with whether you need this at all. If the goal is a robustness adaptation of an existing model, compare a documented augmentation pilot with a real-road collection. The quote and held-out benchmark determine whether the simulated route is economical and adequate. Train, measure against a small real in-cabin evaluation set, and only commission a full collection if the remaining gap justifies it.

A real-cabin collection may be appropriate when the brief requires seat attribution, target microphone geometry, naturally noise-induced speech, or a real-road evaluation set that simulation has not validated. Before scoping anything, check what already exists: our guide to automotive speech datasets covers the open and commercial corpora, and the broader automotive voice data guide maps the whole domain. If you are not sure which route fits, send the deployment brief. Spirelight can assess feasibility and compare a proposed real-cabin scope with a documented augmentation pilot; the response confirms what can be included.