What an AI video actually costs to make, and how to spend less
Most of the money spent on AI video is spent finding out what the video should look like, using the most expensive tool available to find out. Cheaper tools answer the same question. The trick is to ask it in the right order.
Ask ten people what an AI video costs and you get ten numbers, because the honest answer is "it depends on how you make it" and the range is enormous. The same thirty-second video can cost ten times more or ten times less depending on choices that have nothing to do with how good it is.
This is a guide to those choices. It does not quote prices, because prices change and the price for your settings is shown on the button before every run, and that number is the one to trust. What does not change is the shape of the bill: what makes a generation expensive, what makes a project expensive, and where the waste is.
What sets the price of one generation
Four things, in roughly descending order of weight.
The model. The spread between the cheapest video model and the most expensive is wide, and the difference is not simply quality. Some models are priced flat per run, some per second, and a couple of the most expensive are best at one specific thing (long coherent shots, say, or native audio) that you might not need. The model cards show a starting price and the live price updates with your settings; a glance before every run is the habit that saves most.
Duration. Most models bill per second, so a fifteen-second shot is three five-second shots' worth. Given that coherence also decays over the length of a shot, longer is both dearer and less reliable, and the short end of a model's range is nearly always the better buy.
Resolution. 1080p costs more than 720p, and for a draft you will judge on a laptop the difference is invisible. The final render is where resolution pays; the exploration is not.
Audio. On models that can generate sound with the picture, the toggle usually adds to the price. Turn it on when you want native sound and off when you are exploring a look.
Then there is the multiplier: variants. Four variants of one prompt cost four times one. That is sometimes right, when the prompt is nailed and you want the best of four renders. It is usually wrong during exploration, when what you want is to change the prompt, not to roll dice on it.
Where a project's money actually goes
Watch someone make a video from scratch and most of the spend goes into a specific phase: finding out what it should look like. The scene, the character, the framing, the light, the style. They generate a video, see something wrong with the setup, change the prompt, generate another video, repeat. Each iteration costs a video generation, and it took five or ten of them to discover the setup.
That is the waste, and it is structural rather than careless. The question being answered in that phase, "is this the right picture?", is an image question. A video generation answers it too, at many times the price, and then adds a motion question on top that you cannot evaluate until the picture question is settled.
The ladder
So the method, which is the method every experienced person converges on:
Answer the picture in stills. Generate images of the shot. Images are cheap enough to iterate freely: change the prompt, try three compositions, fix the character, move the light. Four variants of an image is a reasonable spend; four variants of a video is not, yet. Keep going until you would be happy to put that frame on a poster. Now the picture question is answered and it cost a fraction of one video run.
Refine the still, not the video. If the frame is nearly right, edit it: a reference edit fixes the hand, changes the jacket, adds the prop, without regenerating from scratch. This is another stage where the video route would have cost a full run for a small correction.
Animate that still, short and low. Take the frame to image-to-video with a short duration and a modest resolution, and answer the motion question: does this move the way you want? A few runs here, possibly on a cheaper model. Motion is easier to judge than it looks, and a draft at 720p shows you the same motion the final will have.
Render the keeper once, at full. Now, and only now, the expensive run: the model you want, the resolution you want, audio on if you want it, one or two variants. You are rendering a shot you have already approved in every respect but its polish, so the odds of it being the last run are high.
Upscale at the end, not on the way. If the delivery is 4K and the model output 720p, a 2× upscale on the finished, trimmed clip beats generating at a higher resolution to begin with, because it processes only the frames you kept.
The ladder does not lower the bar. The final shot is rendered exactly as it would have been. What it lowers is the number of expensive runs it took to reach the shot worth rendering.
Smaller economies that add up
Trim before you process. Anything that runs on a clip you already have (a restyle, an upscale, a caption pass, a filler-word scan) bills on the length it is given. A trimmed timeline clip processes only its visible window, so cut first. Sending a full sixty-second source to restyle when you use eight seconds of it is the single most common avoidable charge.
Draft on a cheap model, finish on a good one. For the tasks that transform footage, the model choice can swing the price enormously for the same job. Try the idea on the cheap one; re-run the take you are keeping on the better one.
Two considered variants beat four random ones. If you are running variants, you are past the exploration stage. Two is usually enough to pick from; four is for when the prompt is final and the shot is the hero.
Reuse history. Every run is kept, prompt and result, in the History tab. A shot you generated on Tuesday for one project can be re-applied on Thursday to another for nothing. Before generating, check whether you already have it.
Audio is nearly free. Music, voiceover, effects and ambience cost very little at editor lengths. It is worth generating several music variants and keeping the best, in a way that is not worth doing with video.
What costs nothing
It is easy, in a token economy, to start seeing prices everywhere. Editing costs nothing. Cutting, trimming, transitions, titles, grades, masks, keyframes, ducking, the mix, the export at any resolution: none of it is metered. The only thing tokens pay for is AI compute, and a failed run refunds itself. So the timeline is where to spend your time, and the generator is where to spend, carefully, your tokens. A video where the generations are ordinary and the edit is excellent reads as a good video. The reverse does not.
The order that works
- Explore in stills until the frame would stand on its own.
- Fix small things with an image edit, not a regeneration.
- Animate short and low to settle the motion.
- Render the keeper once, at the model and resolution the delivery needs.
- Trim, then process: upscale, restyle and caption only what is on the timeline.
- Check history before generating, and put the saved effort into the edit.