Why your AI video sounds dead, and the three layers that fix it
You can tell an AI video with your eyes closed. The picture may be flawless; the sound is a voice, some music, and a silence under both of them that the audience feels as "nothing is really there".
Close your eyes during a film and you can still tell where you are. A kitchen has a hum, a street has traffic under everything, a forest has wind in the leaves and a bird somewhere far off. None of it is the point of the scene, and you would only notice it if it stopped, which is exactly what happens in most AI video.
Generated footage arrives silent, or with audio that is right for one shot and wrong for the next. The natural fix is to lay a music bed over the whole thing and a voiceover on top, and the result is a video that looks like a place and sounds like a studio. The audience cannot say what is wrong. They just do not believe it.
The fix is not better music. It is the two layers people forget exist.
What a soundtrack is made of
Every film soundtrack, from a feature to a thirty-second ad, is three layers stacked in a specific order of importance, and it is roughly the reverse of the order beginners add them.
Ambience comes first. Room tone, atmosphere, the bed of the place. It is continuous, low, and it never draws attention; it is what silence sounds like in a real location. Its job is to make the audience believe the picture is somewhere.
Foley and effects come second. The sounds of things happening: footsteps, a door, a cup set down, wind gusting as a coat moves. These are synchronised to the picture, and they are what makes motion feel physical rather than painted.
Music comes last. It says how to feel about what is happening. It is the least necessary layer and the most noticeable one, which is why it gets added first and why it cannot do the other two jobs.
A voiceover sits on top of all three, and it is the one thing the audience is consciously listening to, which is precisely why it cannot cover for the absence of the layers under it. The gaps between sentences are where the silence shows.
Ambience: the cheapest fix with the biggest effect
If you do one thing from this article, do this. Under every scene, put a continuous bed of the place: the street, the room, the field, the café. Set it low, low enough that a viewer would never mention it, and let it run across the cuts within a scene without a break.
Two things happen. The cuts stop feeling like cuts: a change of angle over a continuous sound bed reads as the same place from a different position, which is what a cut is meant to mean. And the pauses in the voiceover stop being holes. Dead air over a picture is a defect; a quiet room over a picture is a moment.
The sound-effects generator makes these from a description and a duration: "quiet office, air conditioning, distant keyboard" or "night, crickets, occasional car far away". Ask for longer than you need and trim; ambience that has to loop will eventually be heard looping. If a scene runs longer than the generator will give you, extend the track rather than repeating it.
Change the bed when the scene changes, and overlap the change: let the new ambience start a beat before its picture arrives. That is the oldest trick in sound editing and it makes a scene change feel intentional instead of abrupt.
Foley: motion needs a sound
A generated shot of someone walking across a room with no footsteps is a ghost. A car pulling away in silence is a still image with a motion effect on it. The picture can be perfect and the brain will still refuse it, because in a lifetime of watching things move, movement has always made a sound.
Foley was historically a person in a room with a box of props, and it is now a model that reads the clip. The video-matched foley task looks at a clip's source window and generates sound that fits the motion in it: the footsteps land on the steps, the impact lands on the impact. It is not perfect and it works best on shots with a clear physical event; a subtle gesture in a wide shot may get little from it. But for the shots that need it, it is the difference between a moving picture and a scene.
For everything else, one sound at a time. Describe the specific event ("a heavy wooden door closing", "keys dropped on a table") and place it by hand on the frame where it happens. Placement matters more than quality: a mediocre door sound on the right frame is convincing, and a perfect one three frames late is not. A synchronised effect can be a few frames early and still read; late is what the ear catches.
Set foley a little below the voice and clearly above the ambience. It should be the second-most-present thing in the mix and it should never be a surprise.
What to do with native audio
Several video models can now generate sound with the picture: Veo and Kling with a toggle, Sora 2 always. When it works it is startling: a shot arrives with its own room tone, its own footsteps, sometimes its own dialogue. It is worth turning on.
It is also worth being ready to throw away. Native audio is generated per shot, so the room tone of shot one is not the room tone of shot two, and the difference is audible at the cut even when the pictures match. The dialogue, where it appears, is rarely what you wrote. And the level varies from run to run.
The practical approach is to treat native audio as a foley source, not as the soundtrack. Keep the effects that landed, the footstep, the door, the gust; split the clip's sound onto its own lane so you can trim it to just those moments; and put your own continuous ambience under the whole scene so the room does not change at every cut. The native audio becomes one more layer in a mix you control, rather than the mix.
If a model generated speech you did not want, mute it and put a real voiceover on. Do not try to work around it.
Music last, and less than you think
With ambience and foley in place, the music has a much smaller job, and it can be smaller: lower, sparser, sometimes absent. A scene that carries its own sound does not need a bed telling the audience what to feel every second, and a music track that drops out for a moment over a strong physical sound is more dramatic than one that never stops.
That said, if the music has to be present, two rules. Duck it under the voice automatically. And let it run under the cuts too: music that restarts at every scene is as jarring as ambience that does.
Mixing the layers
Balance from the bottom up, in the order the layers were added: ambience low enough to be felt rather than heard, foley clear but not startling, music under the voice, voice on top. Check on a phone speaker, where ambience largely disappears and the mix has to hold up on foley and voice alone.
Then let the loudness matching on export bring the whole thing to streaming level, so it plays at the volume everything else in the feed does.
A test worth running
Play the finished video with the picture covered. It should still sound like somewhere. If it sounds like a voice in a room with music, the two layers are missing; if it sounds like a place where something is happening, they are there.
The order that works
- Lay a continuous ambience bed under every scene, changed at scene boundaries and overlapped across them.
- Foley the physical events, from the matched-foley task where there is a clear action and by hand for the rest, placed on the frame.
- Harvest what native audio got right onto its own lane and mute the rest.
- Add music last, ducked under the voice, running across the cuts, quieter than instinct says.
- Balance from the bottom up and check on a phone.
- Listen with the picture covered. It should sound like a place.