How to make a product video from product photos
You do not need a shoot. You need three clean photos, a clear idea of the one thing the product does, and the discipline to keep the real product real while everything around it is generated.
A product video used to start with a shoot, which meant a studio, a day, and a budget that most products would never earn back. What most sellers have instead is a folder of packshots: the product on white, from three angles, at high resolution. That turns out to be enough, provided you understand what the AI is good at (everything around the product) and what it is not (the product itself).
This is the workflow: from packshots to a thirty-second video in landscape, square and vertical, with sound and titles, without ever putting the product in front of a generator that will redraw it.
Decide the one thing first
Before any generation, write one sentence: what is the one thing this video should make someone believe? "It fits in a jacket pocket." "It is the quietest one." "It looks good on a desk." A product video that tries to make four points makes none; the one that makes one point, with every shot serving it, sells.
That sentence decides the scenes. A pocket claim needs a jacket and a hand. A quiet claim needs a bedroom at night. A desk claim needs a desk that someone would want. Write three scenes, one line each, that show the claim rather than state it. That is your shot list, and it will save you from generating things you cannot use.
Cut the product out once
Take the best packshot and remove its background. You now have the product on transparency, and this is the single most valuable asset in the project, because it is the real product: the real label, the real proportions, the real colour. Everything else will be built around it.
Do the same for the second and third angles if the video will show them. Three cutouts is plenty.
Build the scenes with reference edits, not from scratch
The temptation is to prompt "a red headphone on a wooden desk, morning light" and let the generator make the product. Do not. It will make a red headphone, with a logo that is nearly yours and a shape that is nearly right, and every viewer who knows the product will see it.
Instead, use reference editing: give the editor your cutout as a reference and ask for the scene around it. "Place @Image1 on a wooden desk by a window, morning light, shallow depth of field." The model keeps the product from the reference and generates the setting. Up to fourteen references are allowed, so a second cutout can supply the angle and a mood image can supply the light.
Generate a few variants per scene. Judge them on two things only: does the product look like the product, and does the scene make the claim? A gorgeous scene with a subtly wrong product is a rejection. Check the label text, the proportions, and any hardware detail (ports, buttons, stitching) at full size, because this is where reference edits drift and where a customer will notice.
You now have three still scenes, each one a frame you would be happy to run as a static ad. That is the point at which to start thinking about motion, and not before; a still that is wrong becomes a video that is wrong at many times the price.
Animate the scene, not the product
Two ways to get motion, and the choice depends on what moves.
Ken Burns on the still is the cheap, safe one. A slow push toward the product over five seconds, or a drift across the desk that ends on it, gives the scene a camera without touching a pixel of the product. It costs nothing, it cannot draw the product wrong, and for a scene where the product is at rest it is often the better-looking choice. Alternate the direction across the sequence: push in, drift, pull out.
Image-to-video is for when something should move: steam from the cup beside the product, curtains, a hand entering the frame, light changing. Animate the still with a prompt that describes motion around the product and asks for the product itself to stay still: "camera slowly pushes in, steam rises from the mug, the headphones remain perfectly still". Keep the shot short, four to six seconds, because drift accumulates and the product is the thing that will drift. Generate two variants and inspect the product on the last frame at full size before you accept either.
The rule that keeps the video honest: the product never moves unless it is the real product moving. Turntable spins, folding mechanisms, anything that shows the product's own geometry changing, are the shots where a model will invent detail. If the claim needs them, shoot them on a phone against a plain wall and cut the phone footage in; a real ten-second turntable clip on a real product beats a generated one every time.
The hero shot is the real photo
Somewhere in the video, usually the opening or the close, there should be a shot that is simply the packshot: the real product, sharp, on a clean background, filling the frame. It is the shot that answers "what does it actually look like", and it is the one that a generated scene cannot replace. Give it a slow push and hold it for three seconds. This is also where the price, the name and the call to action belong.
Words on screen, not spoken
Product videos are watched muted more often than not, so the claim goes on screen. Three or four titles across thirty seconds, one thought each, in the brand's typeface, arriving with a short fly-in and leaving before the next shot. Not a paragraph. Not a feature list. The one-sentence claim from the top, broken into the beats that the scenes show.
Then a voiceover as well, for the people with sound on, saying the same things in a fuller form. A generated voice is fine here and consistent across re-cuts, which matters when you will be making variants.
Sound
A product at rest in a beautiful scene with no sound is a catalogue page. Under every scene, an ambience: the room, the street, the morning. Over it, a short music bed that fits the brand and ducks under the voice. If a scene has a physical event (the cup set down, the case clicking shut) put a sound on it. Thirty seconds of sound design is ten minutes of work and it is the thing that makes a viewer feel they are looking at an object rather than a render.
Three shapes
The video will run in a feed (vertical), in a grid (square) and on a page (landscape). Build it in the shape you will use most and then reframe for the others; a product that sits centre-frame reframes easily, and this is a reason to compose the scenes with the product in the middle third from the start. Check every title in every shape; a lower third that was safe in landscape is under the interface in vertical.
Export each at the resolution the platform wants, with loudness matching on.
The two ways it goes wrong
The product is wrong. A generated logo, a fourth button, a strap that attaches somewhere it does not. This is not a small defect; it is the one thing the video must not do, and it is why the cutout, the reference edit and the inspection of every last frame are non-negotiable. When in doubt, use the real photo.
The scene is generic. Wooden desk, plant, window, morning light; every AI product video has it. Go back to the claim. The scene should be the specific place where the claim matters, and a specific place is never generic.
The order that works
- Write the one-sentence claim and three scenes that show it.
- Cut the product out from the best packshot; this is the asset everything is built on.
- Build the scenes with reference edits around the cutout, and inspect the product at full size in every variant.
- Animate around the product: Ken Burns for scenes at rest, short image-to-video for scenes with movement, and phone footage for anything where the product itself moves.
- Open or close on the real packshot with the name, price and call to action.
- Titles for the muted, voice for the rest, sound under everything, then reframe to three shapes and export.