How to cut a long video into short clips
Cutting a long recording into shorts is not really a cutting problem — it is a selection problem, then a framing problem. Both have rules, and neither is where people spend their time.
You have an hour of something — a podcast, a webinar, a stream, a long talking-head take — and you want eight short clips out of it.
The slicing is trivial. Two cuts and you have a clip. That's why the tools that do this automatically feel so impressive for about a day, and why their output feels so interchangeable by the end of the week: they're solving the easy half. The hard halves are which sixty seconds and what the frame looks like once you've squeezed a 16:9 conversation into a phone-shaped hole.
Read before you cut
Don't scrub. Scrubbing an hour of footage looking for good bits takes about an hour, and you'll miss things because moments that read well in text are easy to skim past in audio.
Transcribe it first and work from the words. What you're hunting for is any of five shapes:
- A claim. One sentence someone would repeat to a colleague.
- A number. Concrete, surprising, checkable.
- A disagreement. Two people not agreeing is inherently watchable.
- A story. A tiny one — a specific thing that happened to a specific person.
- A reversal. "Everyone thinks X. It's actually Y."
What is never a clip: pleasantries, setup, "as I was saying earlier", and anything requiring context from twenty minutes ago. If a moment needs a preamble to make sense, it needs the preamble inside the clip — which usually means it isn't one.
Mark the timestamps in the transcript before you touch the timeline. Twenty minutes of reading beats two hours of scrubbing, and the clips are better because you chose them by content rather than by whichever waveform happened to look energetic.
What makes it a clip rather than an excerpt
Start inside the sentence. Not at its beginning — a beat into it. Clips that open on "so, uh, I think the interesting thing here is..." are asking a stranger for four seconds of faith. Start on the interesting thing.
One idea. Two ideas is a video that ends twice, and viewers leave at the first ending.
End on the point, then stop. The most common mistake in short clips is a trailing three seconds of nodding. Cut on the last useful word.
Fifteen to sixty seconds. Long enough to say something; short enough that "I'll watch this later" never comes up.
No orphan pronouns. If the clip opens with "and that's why he did it", nobody knows who he is. Move the start earlier or pick a different moment.
Cut wide, then tighten
On the timeline, take each moment with a couple of seconds of padding on both ends first. Then trim inwards while watching, which is much easier than trying to find the exact frame in one go.
Then run the two passes that make a talking clip feel professionally edited, in this order:
- Remove silences — the pauses that are invisible in an hour-long conversation are enormous in a forty-second clip. This is free; it's just waveform analysis.
- Remove filler words — every "um" and "uh", from a checklist. Untick the ones with character; cut the rest.
Do silences first. The filler scan is priced on clip length, so a tightened clip costs less to analyse.
The result is tight, which reads as confident, which is most of what people mean by "well edited".
Reframe to vertical without decapitating anyone
A 9:16 crop of a 16:9 frame keeps about a third of the width. Something is leaving the frame; your only decision is what.
Three options, and the right one differs per clip:
Centre-crop (Fill). Fine when one person sits mid-frame. Check it — most two-camera podcast setups place speakers off centre, and the default crop finds an ear and a wall.
Blurred-background fill. The whole 16:9 frame letterboxed over a blurred, blown-up copy of itself. Nothing is lost, everything is smaller. The right call for screen shares, slides, or any shot where the whole width matters.
Stack two shots. Split the vertical frame between two PiP layers — speaker on top, slide underneath. The best option for interviews, and the one automatic tools rarely attempt.
When you do crop, put the eyes on the upper third, not dead centre. Faces centred vertically in a 9:16 frame look like passport photos, and the bottom fifth of the frame is where the platform's UI lives — captions, handle, buttons. Anything important down there is buried.
Captions, then a title card
Roughly 80% of these are watched muted, at least at first. Captions are not an accessibility nicety here; they're the audio track. Word-level captions — highlighting each word as it's spoken — hold attention measurably better than a static block, and they come from the transcript you already made.
Keep them clear of the bottom fifth. And if the clip needs context, a two-second title over the first beat is far better than the speaker explaining who they are.
Export once per platform, not once per video
Export at the vertical canvas size, 1080×1920, and stop worrying about bitrate — everything gets re-encoded on upload anyway. What you're protecting against is double compression: export at high quality so the platform's pass is the only lossy one.
One thing worth doing per clip: watch the first two seconds with sound off. That's the entire audition.
The order that works
- Transcribe the whole recording. Read it; mark timestamps.
- Pick moments by shape — claim, number, disagreement, story, reversal.
- Cut wide with padding, then trim inwards.
- Silences, then fillers.
- Set the canvas to 9:16 first, then decide crop, blur-fill or stack per clip.
- Captions — word-level, above the UI zone.
- Check the first two seconds muted. If it doesn't hold you, the clip doesn't start there.