How to mix voice over music so both can be heard

Every beginner mix has the same two faults: the music fights the voice, and the whole thing plays quieter than the video before it. Both have simple causes, and neither is fixed by turning something up.

The complaint arrives in two forms and they have opposite causes. "I can't hear what they're saying over the music" is a balance problem: the music is too loud relative to the voice. "My video is so much quieter than everyone else's" is a loudness problem: the whole mix is too quiet relative to what platforms expect. People fix the second by turning everything up, which makes the first one worse, and then fix the first by turning the music down, which brings back the second.

They are separate jobs. Do them in order and each one stays fixed.

Balance is a ratio, not a level

The only number that matters for balance is the gap between the voice and the music, and the useful starting point is around twelve decibels: music at roughly −12 dB relative to the voice. That sounds like a lot. On the meters it looks like the music is barely there. Listen, and it is exactly where a broadcast mix sits: the music is clearly present, and nobody has to strain for a word.

The reason it seems too low on the way in is that you set it while listening for the music. The audience is listening for the voice. What reads as "nice and full" when you are checking whether the track sounds good reads as "why is the music so loud" to someone who came for the words.

So the mechanics, in order:

  1. Put the voice on its own lane and set it so that the loudest words peak comfortably without clipping. Leave it alone from now on. The voice is the reference and everything else is set against it.
  2. Bring the music in on a separate lane and pull it down until you can follow every word without effort with the music running. Then pull it down a little more; you will be over-hearing it.
  3. Check on something small. Laptop speakers and phones lose the bass, which is where a lot of a music bed's weight is, so a balance that is right on headphones is often music-heavy on a phone. If it works on the phone it works everywhere.

Sound effects and ambience are a third lane, set between the two: audible as texture, never competing with a word.

Ducking does the part you cannot do by hand

A fixed balance is a compromise. During speech you want the music low; in the gaps between sentences, and over the intro and the outro, you want it up. Riding the fader by hand through a ten-minute video is possible, and it is the reason old mixing desks had motorised faders, but it is not a good use of an afternoon.

Auto-ducking is the answer: flag the music clip as Duck, and it dips whenever the voice lanes carry speech and recovers in the pauses. Two things are worth knowing about how it behaves.

It ducks under speech, not under the presence of a clip. A voice lane that is silent for four seconds lets the music back up for those four seconds, which is what you want: the music breathes with the narration rather than sitting flat for the whole duration.

And it is a dip, not a mute. The music never disappears; it steps back. If you can hear it disappearing, the base level of the music is set too high and the duck is doing the job the balance should have done. Fix the balance and the duck becomes subtle.

When a line needs its own treatment

Ducking is uniform: every sentence gets the same dip. Occasionally one moment needs more. A quiet line that has to land, a beat where the music should fall away entirely, a name you want the audience to catch.

That is a volume keyframe job. Set a key before the line at the current level, a key at the start of the line pulled well down, and a matching pair after it. Four keys, a few seconds of work, and the music gets out of the way for that line alone while the rest of the video keeps the automatic behaviour. It stacks with ducking rather than replacing it.

The same tool handles the opposite move: a music swell into a section title, where you want the bed to come up for a few seconds while nobody is speaking.

Edges are where the amateurs show

Music that starts at full volume on frame one and stops dead at the last cut is the audio equivalent of a hard cut to black at the end of a film. It reads as "the file ended" rather than "the video finished".

Every music clip should have a fade in and a fade out. The fade in can be short, half a second or so, because a music track usually has its own way of starting. The fade out wants longer: two to three seconds over the closing shot, ending just before the last frame, so the picture and the sound arrive at their ends together.

Where two music cues meet, overlap them and let the fades cross. A butt-cut between two tracks is a lurch even when the keys happen to match; a crossfade over a second or so hides the join. If you are cutting music to fit and the join lands mid-phrase, the better fix is often to extend the track so it reaches the end in the same style instead of splicing.

Loudness is a different problem

Now the second complaint. Your video, balanced beautifully, plays noticeably quieter than the one before it in a feed.

This is not about peaks. A mix can peak right at the ceiling and still be quiet, because loudness is about how much of the time the signal is near that ceiling, not how high it gets once. Streaming platforms measure that, in a unit called LUFS, and they normalise every upload to a target, about −14 LUFS on the large ones. Crucially, they only turn things down. A video mastered louder than the target gets pulled down to it. A video quieter than the target stays quiet.

So a natural, dynamic mix that measures −20 LUFS arrives on the platform six decibels under everything around it, and it stays there. That is the whole reason the video "sounds quiet".

The fix is to master to the target before uploading, which used to be a job for a separate tool and a separate afternoon. Here it is a switch in the export dialog: Match streaming loudness, which measures the whole mix and corrects it to −14 LUFS in a second pass. It is on by default, because there is almost no situation in which you would not want it, and the one where you would not (someone else is mastering the audio) is what the switch is for.

Two things it does not do. It does not fix balance: it moves the whole mix, so a music-heavy mix arrives music-heavy at the right level. And it does not need running twice: loudness correction is not the kind of thing you can stack, and re-mastering an already-mastered file pushes it somewhere you did not intend. Export once from the timeline.

Rough recordings

A voice recorded on a laptop in a kitchen will have noise under it, and the noise gets more obvious once the mix is at a healthy level. The export-time cleanup handles the mild cases: per-clip noise reduction and normalisation that cost nothing. Anything rougher than that (a fan, a room, traffic) is a job for the denoise or voice-isolate pass on the clip itself, done before you set the balance, because a hissy voice pushes you into wrong decisions about how loud the music can be.

If the voice is generated rather than recorded, most of this section does not apply; a generated voiceover arrives clean and consistent, which is one of the less-advertised reasons people use them.

The order that works

  1. Set the voice so the loudest words sit comfortably below clipping, and then never touch it again. It is the reference.
  2. Bring the music down to roughly twelve decibels under the voice, then check on a phone.
  3. Flag the music to duck, so it steps back under speech and breathes in the pauses.
  4. Keyframe the exceptions: a line that needs silence, a swell that needs room.
  5. Fade every edge and crossfade every music join.
  6. Export with loudness matching on, once, and the video plays at the level everything else does.