Lip sync: making a generated face speak

Lip sync is the difference between a generated person and a generated person who is talking to you. It is also unforgiving: the audience has watched human mouths their entire lives, and they notice.

You have a generated presenter and a voiceover, and right now they are two unrelated things sitting on a timeline. Lip sync is what joins them: the mouth in the video is reshaped to match the words in the audio.

It is also the least forgiving thing in this whole category. Your audience has been watching human mouths form words since infancy. They cannot tell you what is wrong with a bad sync, but they will not believe it.

The audio is most of the result

Before touching anything visual, get the audio right. A lip sync model works from the audio track — so everything wrong with the audio becomes something wrong with the mouth.

Clean speech, on its own. No music, no room echo, no background noise. Sync the voice alone and add the music underneath afterwards. A voice with a bed of music under it produces a mouth reacting to the music.

Normal pace, normal volume. Very fast speech, whispering and shouting are all failure cases. Conversational delivery syncs best.

Trim the silences first. Dead air at the start produces a face doing nothing for a beat before speaking, which reads as a delay.

Decide between a real voice and a generated one before syncing, not after. Generated voiceover is faster; your own voice is more convincing and free. Either works, but re-syncing because you changed your mind is wasted.

The source clip matters more than the model

Given decent audio, almost every remaining failure comes from the clip you are syncing onto.

The face should be reasonably large in frame. A head occupying a small part of a wide shot gives the model very few pixels of mouth to work with. Head-and-shoulders is the format this works best on — which is, conveniently, the format most talking-head content uses anyway.

Front-on beats profile. A face turned away has a mouth the model can barely see. Three-quarter is usually fine; full profile usually is not.

Minimal head movement. A subject who is largely still syncs cleanly. One who is turning, walking toward camera or gesturing energetically gives the model a moving target.

Nothing in front of the mouth. Hands near the face, a microphone, hair across the jaw — all of it degrades the result.

One face. Several people in frame is an ambiguity, and it will not necessarily pick the one you meant.

Short. Same as everything else here. Sync a sentence or two per clip, not a three-minute monologue.

If you are generating the presenter, all of this is under your control — so generate for the sync. A calm head-and-shoulders shot, framed reasonably tight, subject facing camera, minimal motion. It is a less exciting shot than the one you might have prompted for, and it will sync far better.

Where it breaks

Worth knowing so you stop early rather than burning generations:

  • Profile and extreme angles. Not much to work with.
  • Very fast or heavily accented speech. Timing gets approximate.
  • Singing. Sustained vowels and pitch do not map like speech.
  • Shouting or whispering. Mouth shapes fall outside the normal range.
  • Non-speech sounds. Laughter, sighs, breaths tend to come out wrong.
  • Long takes. Drift accumulates.

The order that works

  1. Record or generate the voiceover first.
  2. Trim it — silences and filler out.
  3. Generate or choose a clip framed for syncing: front-on, head-and-shoulders, still.
  4. Sync the clean voice onto the clip.
  5. Add music underneath, ducked, after the sync.
  6. Cut, caption, grade, export.

The single most common mistake is doing this in the opposite order — generating an exciting, dynamic shot, then discovering it will not sync and having to regenerate it as a boring one.

Where this stops being a technique and starts being a decision

Lip sync is the point in this toolkit where "can" and "should" come apart, so it is worth saying plainly.

Putting words into the mouth of a real, identifiable person who did not say them is not a creative technique. It does not matter how good it looks or how harmless the words are — it is the mechanism behind a category of fabrication that is doing real damage, and consent is the whole difference.

The uses this tool is genuinely for: your own face, a performer who has agreed, a character you generated, or a translation and dub of something a person actually said, done openly.

That is not a legal disclaimer, and there is no setting that enforces it. It is just the line, and it is worth being deliberate about which side of it you are working on.