How to add captions to a video (and why they decide retention)

Captions are the highest-return ten minutes in most edits, and they are also the most commonly botched. Here is how to generate them, style them so they do not look amateur, and what transcription always gets wrong.

If you only do one thing to a finished video before publishing it, add captions. It takes about ten minutes, it costs nothing, and it changes the numbers more reliably than any effect you could spend an hour on.

Why they matter more than they should

A large share of your audience has the sound off. Phones in public, offices, late at night next to someone sleeping, and every autoplaying feed that starts muted by default. Without captions those viewers get a silent film — and they leave in the first two seconds, which is exactly the window that decides whether the video gets shown to anyone else.

Comprehension goes up even with sound on. Names, technical terms, numbers and accents all land better when they are also written. This is not a concession to distracted viewers; it is how people read.

Accessibility. Deaf and hard-of-hearing viewers are not an edge case, and in a lot of contexts captions are a requirement rather than a nicety.

Machines read them. Captions are one of the clearest signals a platform has about what your video is actually about.

Burned-in or platform captions?

Both, ideally, and they do different jobs.

Burned-in captions are part of the picture — rendered into the video itself on export. They cannot be switched off, they look exactly as you designed them, and they survive being re-uploaded, downloaded or reposted somewhere that has no caption support. For short-form, this is the one that matters.

Platform captions are a separate subtitle track the player displays. They can be toggled, translated and read by the platform's indexing. Uploading a subtitle file alongside the video is worth doing for anything long-form.

If you are choosing one: burn them in. A reposted clip with no captions is a clip nobody watches.

Generate them, do not type them

The words are already in your audio track. Auto-captions transcribe the audio and place the text on the timeline in sync, which turns a forty-minute typing job into a click.

The thing worth knowing is that good caption tools produce word-level timing, not just per-line timing. That is what allows each word to highlight as it is spoken — the style you see on almost every high-performing short-form video. It reads as movement, and movement holds attention on a screen that is otherwise a talking head.

While you are in the audio, strip the silences and filler words first. Captions make every "um" visible in a way the audio alone does not, and tightening the track before captioning means you are not re-syncing text afterwards.

The style rules

This is where captions go from helpful to amateur, and all of it is fixable in a minute.

Three to five words on screen at a time for short-form. Full sentences force reading instead of glancing, and glancing is all a scrolling viewer will give you. Long-form can carry more, because the viewer has committed.

Never more than two lines. A third line means you are covering the picture with text, and the picture is why they are here.

Keep them out of the platform's furniture. The bottom fifth of a vertical video is buried under usernames, captions and buttons on every app. Sit the text in the lower-middle, not the bottom edge. Check it against a real screenshot of the app you are posting to, not against your editor.

Contrast, non-negotiable. White text on a bright sky is unreadable, and it will happen at least once per video. A shadow, an outline or a semi-transparent plate behind the text fixes it permanently, and costs nothing visually.

One font, and make it yours. Consistent type across every video is one of the cheapest ways a channel starts looking like a brand instead of a collection of uploads. Pick one, set it once, stop thinking about it.

Do not animate every word. A highlight on the spoken word is good. Words that bounce, spin and scale are a distraction competing with your own content.

Proofread the four things transcription gets wrong

Automatic transcription is very good at ordinary speech and reliably bad at exactly the words that matter most to you:

  1. Proper nouns — your brand, product names, people's names. The words you least want misspelled on screen.
  2. Technical vocabulary — your subject-matter jargon is, by definition, unusual.
  3. Numbers and units — "fifteen hundred" versus "1,500", percentages, currencies.
  4. Acronyms — usually rendered as whatever they sound like.

Read the caption track once before exporting. It takes two minutes, and a misspelled product name in forty-point type on your own channel is the kind of thing viewers screenshot.

A five-minute pass

  1. Strip silences and filler from the audio.
  2. Generate captions from the track.
  3. Read them once. Fix names, numbers and jargon.
  4. Set position clear of the platform's UI, and check contrast on your brightest shot.
  5. One font, one style, saved for next time.
  6. Export.

Steps 1 through 4 are the difference between captions that help and captions that look like nobody checked. Step 5 is why the next video takes two minutes instead of ten.