A video that happens to contain speech is not automatically a visual-speech dataset. Frame timing, audio alignment, face visibility, capture consistency, metadata, and the intended model use all affect whether the recording can be trained on responsibly.

Spirelight scopes these fields before recruiting. Existing recordings are never represented as suitable for avatar, lip-reading, or biometric use without a separate rights review.