How to turn a photo into a video with AI
There are two completely different things people mean by “turn a photo into a video”, and picking the wrong one is why the result looks off. One costs nothing and always works; the other moves the subject and sometimes melts their hands.
"Turn a photo into a video" means two completely different things, and most of the disappointment with AI animation comes from reaching for the wrong one.
A camera move on a still. The photo does not change; the frame moves across it. Slow push in, drift left, pull back to reveal. The subject stays exactly as photographed.
Generated motion. The model invents what happens next — the person turns their head, the water moves, the leaves shift. The photo becomes the first frame of something new.
The first is free, instant and always looks correct. The second is the impressive one, costs tokens, and sometimes produces a hand with six fingers. Knowing which one a shot needs is most of the skill.
When a camera move is the right answer
More often than people expect. A Ken Burns move — a slow zoom and drift across a still — is what documentary television has run on for fifty years, and viewers do not read it as a limitation. They read it as production.
Reach for it when:
- The photo is of a real thing you cannot afford to have altered — a product, a person, a place, anything where "close enough" is wrong.
- The still is on screen briefly, under narration. Nobody is studying it.
- You need many shots. Forty camera moves cost nothing; forty generations cost real money.
- The photo is detailed — a landscape, a crowd, an interior. There is enough in the frame that moving across it reveals something.
The one rule: keep it slow, and pick a direction with a reason. Zooming into a face works because it tightens. Zooming into nothing in particular just looks like a mistake in the edit.
When you need generated motion
When the subject itself has to move, and no camera trick will fake it. A person speaking. Fabric in wind. A car actually driving. That is image-to-video, and the photo becomes the start frame.
Prepare the photo before you animate it
Three things, in this order, and they matter more than the prompt.
Crop to your output aspect first. If the video is 9:16 and the photo is 3:2, something is getting cut — decide what, rather than letting a generator decide by filling the edges with invention.
Upscale it if it is small. Generators work from what you give them; a soft input produces soft motion. Upscaling the still first is cheaper than re-rolling the video three times wondering why it looks mushy.
Clean it up. Remove the thing you did not want in frame before animating. Fixing it afterwards means fixing it in every frame instead of one.
Prompt the motion, not the picture
This is the mistake almost everyone makes on their first try: describing what is in the photo. The model can already see the photo. Describing it again wastes the prompt and sometimes convinces the model to re-invent what is already there.
Describe what changes:
- Bad: "a woman in a red coat standing in a city street at night"
- Better: "she turns to look over her shoulder, camera holds still"
Two more things worth stating explicitly, because models will otherwise decide for you: whether the camera moves, and how much should happen. "Subtle movement, camera locked off" is a real instruction and it produces far more usable clips than an open-ended prompt.
Pin where the shot ends
The underused trick. Several models accept a start frame and an end frame — you supply both ends and the model fills the middle. Instead of hoping the shot lands somewhere usable, you decide where it lands.
Two uses that pay off immediately:
Landing on a specific pose or framing so the next clip cuts cleanly against it.
Perfect loops. Make the end frame the same image as the start frame. The clip returns to where it began and plays forever without a visible seam — ideal for a background, a hero banner or anything that sits behind text.
What goes wrong, and what to do
Faces and hands distort. The most common failure, and it gets worse the longer the clip. Generate shorter, cut before it degrades. A clean two seconds beats a five-second clip you have to hide.
Too much motion. The model over-interprets and the whole frame churns. Ask for less explicitly, or supply an end frame close to the start frame to constrain how far it can travel.
The clip is right but too short. Do not re-roll it — you will not get that take back. Interpolating frames smooths and extends the motion you already have.
It changed something you needed unchanged. That is the signal to switch techniques. Use a camera move on the original photo instead.
Putting several together
A sequence of animated stills is a real format — it is how a great many explainer and history videos are actually made. What makes it read as deliberate rather than cheap:
- Vary the move. Push in, then drift, then hold. Three identical zooms in a row reads as a template.
- Cut on the audio, not on the motion finishing. The narration decides the rhythm.
- Mix the two techniques. Generated motion on the two or three shots that need it, camera moves on the rest. Viewers cannot tell which is which, and your costs collapse.
What it costs
Camera moves on stills cost nothing — they are an editing effect, rendered on export along with everything else, at 4K and unwatermarked. Generated motion is priced by what the underlying model charges, per clip.
That gap is the whole argument for mixing them. A ten-minute video built from forty animated stills and six generated clips is a fraction of the cost of one built from forty generated clips, and for most formats nobody watching could tell you which shots were which.