How to score a video without licensing a track

Library music sounds like library music and a commercial track gets your upload claimed. Generating the bed solves both — but the part that actually changes how a video feels is the ambience underneath it, which no generated clip comes with.

Music is the last thing most people sort out and the first thing an audience feels. The usual options are all bad in their own way: a commercial track gets the upload claimed, a subscription library gives you the same eleven cues everyone else is using, and silence makes a two-minute video feel four minutes long.

Generating the bed sidesteps all three. It is also, in editor quantities, one of the cheapest things you can do — audio runs cost a fraction of what a video generation costs, which changes how you should approach it: make three and throw two away.

Prompt a bed, not a song

The music under a video has a job, and the job is to not be noticed. That means the prompt is mostly a list of things it must not do.

A useful music prompt names four things:

  • Genre and instrumentation — "warm analogue synth pad with soft electric piano", not "cinematic".
  • Tempo and energy — "slow, 80 bpm, unhurried" reads very differently from "driving".
  • The mood in plain words — "hopeful but restrained" beats a pile of adjectives.
  • The negatives — no vocals, no build, no big drop, no percussion after the first bar. This is the part people skip, and it is the part that decides whether the track fights your narration.

Anything with a vocal, a hook or a dramatic structure competes with speech. If nobody is talking — a montage, a title sequence, a product loop — the opposite applies and you want the track to actually go somewhere.

There is a separate lyrics field with verse and chorus tags if you do want a song. That is a different job from scoring, and worth knowing exists.

Generate two or three variants and keep one. At these prices, running a second option costs less than the ten minutes you would spend arguing yourself into the first.

Getting it to the right length

A generated track comes back at whatever length you asked for, which will not be your edit's length. Two ways out, and looping is the worst of them:

Extend it. Extending a track continues the same style for up to a minute more. This is what you want for "the track is 45 seconds and the video is 80" — a continuation sounds like the same piece of music, where a loop sounds like a loop, because the ear finds the seam within two repetitions.

Patch the bad bit. If a run is good except for eight seconds in the middle where the drums arrive uninvited, patching regenerates just that section — two sliders, in seconds. Cheaper than re-rolling a whole track you otherwise liked.

If you genuinely need a loop, cut on a bar line rather than on a silence, and crossfade the join by a fraction of a second.

Stems, which most people never touch

Any track can be split into its parts: vocals, drums, bass, guitar, piano — or "everything else", which is the backing without the lead vocal.

Three things this unlocks:

  • An instrumental bed from a track with a vocal. Split out everything else and the vocal is gone, which is usually all that stood between a good generation and a usable bed.
  • A drums-only section. Under a title card or a fast montage, percussion alone is stronger than the full mix and leaves far more room for a voice.
  • Dropping instruments as the narration starts. Full mix over the intro, bass and pads only once you begin talking. That is a mix move that normally requires the multitrack, and stems are the multitrack.

The part that matters more than the music

Generated video is silent. Not quiet — silent. No room tone, no footsteps, no traffic, no wind. That absence is a large part of why a beautifully generated shot still reads as synthetic, and no amount of music covers it, because music sits in a different layer of attention from ambience.

Two tools fix it:

Sound effects, generated from a short description plus a duration. A door, a whoosh, a keyboard, a distant siren.

Foley matched to the clip — the foley model watches a video clip's source window and generates sound that fits the motion in it: footsteps that land where the feet land, impacts on the impacts, ambience that matches the space. This is the single most underused thing in the audio tab, and on generated footage it does more for realism than another re-roll of the picture.

A shot of a street with faint traffic and one distant horn is a street. The same shot in silence with a synth pad over it is a screensaver.

Mixing it so it disappears

The mechanics are quick and the ducking guide covers them: put music on its own lane, start around 12 dB under the voice, and flag the music clip as Duck so it dips under speech automatically and comes back in the pauses.

Three judgement calls that ducking does not make for you:

Do not score everything. Silence is punctuation. A track that runs unbroken from the first frame to the last flattens the whole video; dropping it out for six seconds before your closing line does more than any swell.

Enter and leave on a picture change. Music that starts mid-shot sounds like a mistake. Music that starts on a cut sounds intentional, even when it is the same track at the same level.

Check it on phone speakers. A bed that sits politely under a voice on monitors routinely swallows it on a phone, and a phone is where this will be watched. Small speakers lose the bass and keep everything in the range your voice lives in.

About the licensing question

The reason people reach for generated music is usually the copyright claim, so it is worth being precise about what that solves and what it does not.

Automated content matching on the big platforms works by fingerprinting known recordings. A track generated for your video is not a known recording, so it does not match one — which is exactly the problem people are trying to escape when they stop lifting music from other videos.

What that does not settle is what you are allowed to do with the output commercially, which depends on the terms of the tool that made it. Ours are in the terms; read whichever apply to whatever you use. And a track labelled "royalty free" on some aggregator site is not automatically either of those things — that phrase has been attached to plenty of music that later generated claims.

The order that works

  1. Decide what the music is for — under speech, or carrying a montage. They are opposite prompts.
  2. Prompt genre, tempo, mood and the negatives, then run two or three variants.
  3. Extend to length rather than looping; patch a section rather than re-rolling.
  4. Split stems if the mix is too busy under the voice.
  5. Add ambience and foley to any generated footage — this is the realism step.
  6. Duck the music, start 12 dB down, and drop it out at least once.
  7. Check the mix on a phone before you export.