Motion transfer: how to make a character move like a reference video

If you have ever tried to prompt your way to a specific dance, walk or gesture, you know it does not work. Motion transfer takes the movement from a video that already has it — but which half you keep is a decision most people get wrong the first time.

Try to describe a specific movement in a prompt. Not "she dances" — an actual routine, with the timing and the weight shifts and the thing the arms do on the third beat.

You cannot. Prose is a terrible format for choreography, which is why dance notation exists and why nobody uses it. Text-to-video will give you a dance, cheerfully, and it will not be the one in your head.

Motion transfer solves this by not asking you to describe anything. You supply a video that already contains the movement, and the movement gets applied somewhere else.

The two directions, which are not the same thing

This is the part people get wrong, and getting it wrong wastes generations. There are two operations, they take similar-looking inputs, and they produce completely different output.

Animate — keep your scene, borrow the motion

You supply a reference character image. The output is your character, in your image's world, performing the motion from the reference video.

The reference video is used only as a source of movement. Its background, its lighting and the person in it do not appear in the result.

Use this when the character is the point: your generated presenter doing a specific gesture, your illustrated mascot walking the way a real person walks.

Replace — keep the video, swap the subject in

You supply a subject image. The output is the reference video's world — its background, its lighting, its camera move — with your subject in place of the person who was there.

Use this when the scene is the point: you have footage you like and want somebody else in it.

The quick test: ask which half of the reference video you want to keep. If the answer is "only the movement", that is animate. If the answer is "everything except the person", that is replace. Picking the wrong one produces something that looks broken in a way that is oddly hard to diagnose, because the output is technically what you asked for.

Choosing the reference video

This matters more than the model, and it is where most failures come from.

One subject, clearly visible. Two people in frame gives the model an ambiguity it will resolve badly.

Whole body, if the motion involves the whole body. A walk cycle cannot be transferred from a clip cropped at the waist.

Clean separation from the background. A subject in dark clothes against a dark background is hard to track. Contrast helps enormously.

A stable camera. Motion transfer is trying to extract the subject's movement. A handheld camera adds motion that is not the subject's, and it gets baked in.

Nothing crossing in front. Occlusion — a passing object, an arm covering the torso — is where tracking breaks.

Short. A few seconds. Same rule as everything else in generation: quality degrades over length, so take the section you need rather than the whole clip.

Ordinary phone footage of yourself doing the movement is an excellent reference and is usually better than something found online, because you can control every item on this list.

Choosing the character image

Match the starting pose, roughly. A character standing front-on transfers into a front-on movement far more reliably than one photographed sitting in three-quarter profile.

Show what needs to move. If the reference has legs in it, your character image needs legs in it.

Keep it clean. A cluttered background gives the model more to reinterpret. Generating the character on a plain background first — or cutting it out — makes the transfer noticeably more stable.

Match the aspect ratio to the reference video where you can.

What it is good and bad at

Good: whole-body movement, walks, dances, gestures, sports motion, repeatable actions you can film yourself. Anything where the timing is the thing you cannot write down.

Bad: fine hand detail, facial performance, anything involving objects being handled, fast motion with heavy blur, and interactions between two subjects.

For faces specifically, this is the wrong tool. Making a face talk is lip sync, not motion transfer — different problem, different model.

The order that works

  1. Film or find the reference first. The movement is the constraint; build around it rather than discovering afterwards that no usable reference exists.
  2. Generate the character as a still, on a clean background, in a matching pose.
  3. Pick your direction using the "which half do I keep" test.
  4. Run it short, check it at full speed, then frame by frame at the moment it looks off.
  5. Fix rather than re-roll. Trim to the good section, upscale if it is soft, and grade it to match the shots around it.

That last point is the same advice as everywhere else in this workflow, and it is the same reason: generation is a draw from a distribution, and a re-roll does not give you back the take you liked.