Why your image-to-video clips look wrong — and how to fix them

Every one of these failures is common, and most of them have a cause you can control. The important half is the last section: what to do with a clip that is nearly right, because re-rolling it means losing the take you liked.

You supply a good photo, write a reasonable prompt, and get back four seconds of something subtly wrong. This is the normal experience of image-to-video, not a sign you are doing it badly — but most of the failures have a cause you can do something about.

It falls apart the longer it runs

The most common one. The first second is clean, the third is questionable, the fifth has stopped making sense.

Generation quality degrades over the length of a clip, because each moment is extrapolated from the last and small errors compound. There is no prompt that fixes this.

The fix is to stop asking for long clips. Generate two to four seconds. If you need eight seconds of screen time, that is two shots and a cut, which is what an editor would have done anyway — long unbroken takes are rarer in real edits than people assume.

Hands, faces and text

The three things models are worst at, and the three things viewers look at hardest. A distorted background goes unnoticed; a hand growing a finger does not.

Some practical mitigations:

  • Frame them out. If the hands are not the subject, crop or compose so they are not in shot.
  • Ask for less movement from them. Motion is where distortion happens. A person who turns their head slightly is far safer than one who gestures.
  • Cut before it happens. Watch the clip frame by frame, find the moment it breaks, and trim to before it. Half of what makes generated footage usable is trimming aggressively.
  • Do not expect readable text. Signage, labels and logos generally come out as convincing-looking nonsense. If text matters, add it as a title layer over the shot instead.

The whole frame churns

You asked for a subtle movement and everything is swimming — background warping, edges crawling, the image never settling.

This is the model over-interpreting. Two causes worth separating: you asked for too much, or you asked for nothing in particular and it filled the gap.

Say how much. "Subtle movement, camera locked off" is a real instruction and it produces far more usable clips than an open prompt. Being explicit that the camera does not move is often the single highest-value phrase.

Or constrain both ends. Several models accept a start frame and an end frame. Supplying an end frame close to the start frame is a hard limit on how far the shot can travel — a much more reliable control than adjectives.

It ignored your prompt

Usually because the prompt described the picture rather than the motion.

The model can already see the image. Describing what is in it wastes the prompt and sometimes convinces the model to reinvent what is already there — which is why you occasionally get a different person wearing the same coat.

  • Not this: "a woman in a red coat on a city street at night"
  • This: "she turns to look over her shoulder, camera holds still"

Describe the change. The picture is already handled.

The subject drifts off-model

You animate the same character three times and get three slightly different people — a face that shifts, a jacket that changes shade.

Consistency comes almost entirely from feeding in the same source image, not from prompt wording. Generate the character as a still first, lock it, and use that exact file as the start frame for every shot. Trying to hold a character together through description alone does not work reliably, and you will spend a lot of generations proving it.

The part that matters: fix, don't re-roll

This is the habit that separates people who find AI video workable from people who find it maddening.

A re-roll does not give you your take back. Generation is not deterministic. If a clip is 80% right, rolling again does not produce that clip plus improvements — it produces a different clip, quite possibly worse, and the one you liked is gone.

So before re-rolling, try:

Trim it. The most underrated fix. Most failures are at the end. Two clean seconds is worth more than four compromised ones.

Interpolate frames. If the motion is right but stuttery or too short, interpolation smooths it and extends it — from the footage you already have.

Upscale it. If it is soft rather than wrong, upscaling is cheaper than another roll and preserves the take. This also covers the case where the source photo was too small, which is worth fixing on the input next time.

Cut around the problem. Overlay a title, cut to another shot at the bad moment, or slow it down. Editing has solved "this bit is unusable" for a century.

Change the input, not the dice. If the composition is the problem, edit the source image and animate again. That is a targeted change with a predictable effect, unlike re-rolling the same prompt hoping for better luck.

When to actually re-roll

When the clip is fundamentally wrong rather than imperfect: the motion is not what you asked for, the subject is wrong, the whole thing is unusable. Then roll again — but change something first. The same prompt, the same image and the same settings will just give you another draw from the same distribution.

And if three rolls have all missed in the same way, the problem is the shot you are asking for, not the model. Go back and rewrite what the beat is meant to do.