Custom talking-head collection

Talking-head video for avatar training, with likeness rights that name the use

Commission on-camera monologue and dialogue recorded for talking-head models: fixed framing, frame-accurate sync, phoneme-balanced scripts, and a signed release in which every participant approves avatar and synthetic-media use.

Free sample: Tell us the framing, resolution, and languages you need. The team prepares a matching example clip and sends it manually within 48 hours, free of charge.

Collection specification

Video
1080p or 4K at 25, 30, or 60 fps, fixed front framing from mid-chest up
Cameras and lighting
Single front camera as standard, optional profile angles, diffuse lighting held constant per session
Audio
Lapel or boom microphone per speaker, 48 kHz WAV on separate channels, frame-accurate sync to video
Content mix
Phoneme-balanced scripted reads, spontaneous monologue, and two-person dialogue with listening segments
Coverage
Languages, accents, age bands, and appearance variation recruited to your demographic matrix
Rights
Signed release per participant naming avatar, lip-sync, and synthetic-media training use
Pricing
Custom, scoped to your conditions

What makes talking-head footage trainable

A talking head dataset is mostly a consistency exercise. Framing stays fixed from mid-chest up, head pose is held within an agreed range, and lighting is diffuse and constant across the session, so the model learns faces rather than shadows. For lip movement, frame rate matters more than resolution: bilabial closures that smear at 25 fps resolve cleanly at 60.

Sync is the quiet requirement, because a lip-sync model trained on footage that drifts half a frame learns the drift as truth. Audio is locked to video at capture rather than corrected afterward. Scripts are phoneme balanced so every viseme appears often enough to learn from, and each session includes expression range, silences, and listening segments instead of a single neutral read.

Monologue and dialogue capture modes

Scripted reads give you controlled viseme coverage and clean alignment between text, audio, and lip movement, which is where most lip sync training data starts. Spontaneous monologue adds the hesitations, head movement, and natural prosody that scripted speech suppresses. Two-person dialogue adds turn-taking: interruptions, backchannels, and the gaze shifts that mark a speaker handing over the floor.

Dialogue also fills the gap most digital human datasets leave open: the listening face. An avatar in conversation spends half its time not speaking, and if the corpus contains only speech, the model has nothing honest to generate between turns. Recording both sides of every conversation produces training data for AI avatars that speak and avatars that listen, from the same session.

Likeness rights that name the avatar

Face video sits next to biometric data in European privacy law, and an avatar is the most literal reuse of a likeness there is. Every participant signs a release that names the intended use in plain language: avatar generation, lip-sync training, synthetic media, whichever applies to your project. The consent chain is documented per participant and delivered with the data, so a rights question two years from now has a paper answer.

That is the line scraped footage cannot cross. Public talking face datasets built from broadcast or YouTube material are generally licensed for research only, and the people on screen never agreed to become a digital human. A custom collection costs more than a download, and it is the difference between a model you can ship and one you can only publish about.

This offer is not the right fit when

  • You want to scrape public faces or celebrity footage and train an avatar on people who never agreed to it.
  • You need tens of thousands of hours by next week at web-scrape prices.
  • Your synthetic-media use cannot be described in plain language to the person being recorded.

Get a free sample

  1. Your details
  2. Your project
  3. Verify

Three short steps. The team then follows up manually within 48 hours, and confirms volume, rights, QA, and delivery if you want the scope priced.

Would rather talk it through first? Scope a talking-head collection

Prepare the brief

Bring a specification your vendors can price consistently

Use the free worksheet to define coverage, rights, delivery fields, held-out rules, and acceptance tests before requesting a collection plan.

Build a dataset specification

Buyer documentation

Specimen data card and provenance structure, evaluation brief template, and acceptance and held-out questions.

Related services

Guides for this decision

Need a scoped collection plan? Send the language, channel, volume, annotation, and timing you know. The quote form opens directly in project-brief mode.