A video that happens to contain speech is not automatically an audio-visual speech dataset. Frame timing, audio alignment, mouth visibility, capture consistency, metadata, and the intended model use all decide whether the recording can be trained on responsibly.
Face video is biometric-adjacent data, and several of the model uses on this page push it into special-category territory under the GDPR. So every participant signs likeness, biometric, and AI-training language that names the intended use before a camera rolls, which is the documented consent chain scraped footage and public research corpora cannot offer. Existing recordings are never represented as suitable for avatar, lip-reading, or biometric work without a separate rights review.