A team building a meeting assistant, an overlap-aware segmentation system, or multi-speaker ASR hits the same question early: which meeting speech dataset to train and benchmark on. The public corpora are known by name, AMI, ICSI, DipCo, AISHELL-4, AliMeeting, and the CHiME dinner-party sets, and each fixes a different combination of language, room, microphone geometry, and conversational style. None of them was recorded in your acoustic conditions, and the obvious substitute, your own production meetings, is exactly the audio that privacy promises and customer contracts take off the table.

This guide compares the named meeting corpora, sizes, capture setups, and licenses included, shows where the family runs out, and lists what to specify when commissioning meeting-style collection of your own.

What meeting speech data has to contain

Meeting audio is defined by properties read-speech corpora do not have, and a corpus is only useful for meeting models if it captures all of them. Overlap comes first. In a natural meeting, a real share of speech happens while someone else is talking: backchannels, interruptions, two people starting at once. A dataset in which speakers politely alternate teaches turn-taking that does not exist in deployment, so overlap has to be present in the audio and marked in the transcript rather than smoothed out of it.

Capture setup is the second half. Meeting devices sit on a table, so the training signal is far-field: multi-channel microphone arrays with documented geometry, recorded in rooms with real reverberation and noise. The strongest corpora pair those arrays with a close-talk headset per participant, which provides a clean reference channel for transcription and speaker-attributed ground truth while the array carries the deployment condition. Finally, every utterance needs a speaker label with start and end times, including inside overlapped regions; without per-speaker timing the corpus supports transcription but not separation or turn-taking work. These are the same properties that define conversational speech generally, covered in the conversational speech data guide; meetings add more speakers and more distance.

The public meeting corpus landscape

Six named corpora account for most published meeting-speech work. Sizes, setups, and licenses below are taken from the official distribution pages and corpus papers.

CorpusSizeCapture setupLicense or access
AMI100 hours, EnglishInstrumented meeting rooms; close-talk headset and lapel mics plus 8-channel microphone arrays; about two thirds scenario-based design meetingsCC BY 4.0, free
ICSIAbout 72 hours, 75 meetings, EnglishNatural weekly research-group meetings at ICSI Berkeley, 2000 to 2002; a close-talk mic per speaker plus 6 table-top mics, about 6 participants per meetingLDC catalog (LDC2004S02), fee-based
DipCo10 sessions of 15 to 45 minutes, EnglishStaged dinner parties recorded by Amazon, 4 participants; a close-talk mic per speaker plus five 7-mic far-field array devicesCDLA-Permissive 1.0, free via Zenodo
AISHELL-4120 hours, 211 sessions, MandarinReal conference-room meetings, 4 to 8 speakers; 8-channel circular table-top arrayCC BY-SA 4.0, free via OpenSLR
AliMeeting118.75 hours, 240 sessions, MandarinReal meetings, 2 to 4 participants; 8-channel far-field array plus a headset mic per speakerCC BY-SA 4.0, free via OpenSLR
CHiME-6About 50 hours, 20 sessions, EnglishDinner parties in real homes, 4 participants; six 4-mic array devices plus binaural mics per speakerCC BY-SA 4.0, free via OpenSLR

Two clarifications. About two thirds of AMI is scenario meetings, participants role-playing a design team against a brief, with the rest naturally occurring; ICSI is the opposite, real recurring research meetings, all captured in one instrumented room at one institute. DipCo and CHiME-6 are not meetings at all but dinner parties; they earn their place in the family because they are the hardest public far-field multi-talker material and are routinely used to stress-test meeting stacks.

Where the public family runs out

Age and provenance first. The AMI corpus was collected in the mid-2000s and ICSI between 2000 and 2002, before the meeting platforms, codecs, and devices current systems listen through. The ICSI meeting corpus draws its 75 meetings from 53 unique participants in a single room, and the AMI scenario portion recycles one design brief, so vocabulary and group dynamics are narrower than the hour counts suggest.

Scale and language second. Roughly 100 hours per corpus was generous in 2005 and is small against what current multi-speaker models consume. The sets recorded at modern scale with modern arrays, AISHELL-4 and AliMeeting, are Mandarin, so teams building for English or other European languages get the newest far-field material in a language their product will never hear. Beyond English and Mandarin, the public meeting family covers essentially nothing.

The archive teams reach for next is their own production meetings, and that path is blocked for good reasons. Meeting recordings are made so participants can review them, not so vendors can train on them: privacy promises, the scope of recording consent, and customer-data restrictions in enterprise contracts generally rule out training on production meeting audio. What remains is a public family that undershoots modern needs and a private archive nobody may touch, which is the gap commissioned collection exists to fill.

What to specify when commissioning meeting-style collection

A commissioned meeting corpus is a specification exercise: the sessions are designed so the delivered data has the properties above by construction. The items that carry the value:

  • Participant counts. Sessions of 2 to 4 speakers behave differently from sessions of 6 to 8: overlap rate, turn length, and interruption patterns all shift. Match the distribution to the meetings the product will see, and recruit enough distinct speakers that held-out evaluation splits are possible.
  • Microphone geometry. Fix the array to the deployment device: channel count, spacing, and placement, recorded simultaneously with a separate close-talk channel per speaker so far-field models train against near-field references.
  • Overlap elicitation. Task design controls conversational dynamics. Debates, joint planning tasks, and time pressure raise interruption and backchannel rates; moderated turn-taking suppresses them. Specify the overlap behavior you need and verify it in delivered sessions rather than hoping it emerges.
  • Languages and accents. The public family gives you English and Mandarin. Any other language, and any deliberate accent mix within a language, has to be recruited for.
  • Consent and metadata. Every participant signs recording consent that names model training as the purpose, so the rights question that blocks production meetings never arises. Session metadata records room, device, seating, and pseudonymous speaker identifiers.
  • Deliverables. Per-channel audio, time-coded speaker-attributed transcripts with overlap marked, and agreed train, development, and test splits with disjoint speakers.

The commercial starting point for this specification is Spirelight's full-duplex conversational speech collection: a separate channel per speaker, natural overlap, interruptions and backchannels elicited by task design, and turn-timing metadata, built around two-speaker interaction. A meeting-style brief is that specification extended, and the extensions are exactly the items above: more participants per session, a defined mic geometry, and the languages the public corpora skip, agreed at scoping rather than assumed.