Custom talking-head collection
Talking-head video for avatar training, with likeness rights that name the use
Commission on-camera monologue and dialogue recorded for talking-head models: fixed framing, frame-accurate sync, phoneme-balanced scripts, and a signed release in which every participant approves avatar and synthetic-media use.
Free sample: Tell us the framing, resolution, and languages you need. The team prepares a matching example clip and sends it manually within 48 hours, free of charge.
Collection specification
- Video
- 1080p or 4K at 25, 30, or 60 fps, fixed front framing from mid-chest up
- Cameras and lighting
- Single front camera as standard, optional profile angles, diffuse lighting held constant per session
- Audio
- Lapel or boom microphone per speaker, 48 kHz WAV on separate channels, frame-accurate sync to video
- Content mix
- Phoneme-balanced scripted reads, spontaneous monologue, and two-person dialogue with listening segments
- Coverage
- Languages, accents, age bands, and appearance variation recruited to your demographic matrix
- Rights
- Signed release per participant naming avatar, lip-sync, and synthetic-media training use
- Pricing
- Custom, scoped to your conditions
What makes talking-head footage trainable
A talking head dataset is mostly a consistency exercise. Framing stays fixed from mid-chest up, head pose is held within an agreed range, and lighting is diffuse and constant across the session, so the model learns faces rather than shadows. For lip movement, frame rate matters more than resolution: bilabial closures that smear at 25 fps resolve cleanly at 60.
Sync is the quiet requirement, because a lip-sync model trained on footage that drifts half a frame learns the drift as truth. Audio is locked to video at capture rather than corrected afterward. Scripts are phoneme balanced so every viseme appears often enough to learn from, and each session includes expression range, silences, and listening segments instead of a single neutral read.
Monologue and dialogue capture modes
Scripted reads give you controlled viseme coverage and clean alignment between text, audio, and lip movement, which is where most lip sync training data starts. Spontaneous monologue adds the hesitations, head movement, and natural prosody that scripted speech suppresses. Two-person dialogue adds turn-taking: interruptions, backchannels, and the gaze shifts that mark a speaker handing over the floor.
Dialogue also fills the gap most digital human datasets leave open: the listening face. An avatar in conversation spends half its time not speaking, and if the corpus contains only speech, the model has nothing honest to generate between turns. Recording both sides of every conversation produces training data for AI avatars that speak and avatars that listen, from the same session.
Likeness rights that name the avatar
Face video sits next to biometric data in European privacy law, and an avatar is the most literal reuse of a likeness there is. Every participant signs a release that names the intended use in plain language: avatar generation, lip-sync training, synthetic media, whichever applies to your project. The consent chain is documented per participant and delivered with the data, so a rights question two years from now has a paper answer.
That is the line scraped footage cannot cross. Public talking face datasets built from broadcast or YouTube material are generally licensed for research only, and the people on screen never agreed to become a digital human. A custom collection costs more than a download, and it is the difference between a model you can ship and one you can only publish about.
This offer is not the right fit when
- You want to scrape public faces or celebrity footage and train an avatar on people who never agreed to it.
- You need tens of thousands of hours by next week at web-scrape prices.
- Your synthetic-media use cannot be described in plain language to the person being recorded.
Get a free sample
- Your details
- Your project
- Verify
Three short steps. The team then follows up manually within 48 hours, and confirms volume, rights, QA, and delivery if you want the scope priced.
Would rather talk it through first? Scope a talking-head collection
Prepare the brief
Bring a specification your vendors can price consistently
Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.
Buyer documentation
Specimen data card and provenance structure, evaluation brief template, and acceptance and held-out questions.
Related services
Guides for this decision
Need a scoped collection plan? Send the language, channel, volume, annotation, and timing you know. The quote form opens directly in project-brief mode.