Human speakers acknowledge, interrupt, hesitate, restart, and speak at the same time. A model trained only on isolated utterances or neatly alternating dialogue does not learn when to listen, yield, continue, or recover.

Separate synchronized channels preserve those events without forcing diarization to guess who spoke. The annotation scope can stay lightweight or extend into precise turn-taking labels depending on the model job.