Video upscaling: what it can and can't recover

Upscaling is the most misunderstood tool in the category, mostly because of a television trope. It is genuinely useful — but for a narrower set of problems than people expect, and running it at the wrong point in the pipeline wastes it entirely.

"Enhance." Then the blurry reflection in the window resolves into a legible face, and the detective has their man.

That scene has done more damage to expectations around video upscaling than anything else, because the real thing works in a fundamentally different way — and knowing the difference is what tells you when it will help you and when it will waste your money.

What upscaling actually does

An upscaler takes a small or soft image and produces a larger, sharper one by inventing detail that is consistent with what it can see. It has been trained on enormous numbers of image pairs, so it knows what a sharp version of a soft eyelash tends to look like, and it draws that.

The output is plausible. It is not recovered. If the original never captured the texture of the fabric, the upscaler is guessing what that fabric would look like — usually well, occasionally not at all.

For most practical purposes the distinction does not matter. The result looks better, and better is what you wanted. It matters enormously in two cases:

Faces. An upscaler will happily invent pore-level detail on a face that was twelve pixels wide. The result is a convincing human face that is subtly not the person. If the footage is of someone you know — or, worse, someone the audience knows — look closely at the eyes and mouth before accepting it.

Anything evidential. Text, licence plates, documents, signage. The upscaler will produce crisp, confident, invented characters. Never treat upscaled text as readable information.

Where it genuinely helps

Small-but-clean sources. An old scan, a 480p export, a stock clip that only exists at one size. Clean information, just not enough of it — this is the ideal case.

Soft AI-generated footage. Generated video is often a little mushy, particularly on faster or cheaper model tiers. Upscaling after generation is one of the highest-return moves available, and far cheaper than re-rolling for sharpness.

Rescuing a take you cannot get again. The clip is right, the performance is right, the resolution is not. This is the case where upscaling earns its cost outright, because the alternative is losing something you cannot regenerate.

Preparing a still before animating it. Covered below — this is the one people get backwards most often.

Where it makes things worse

Heavy compression artifacts. Blocking, banding, the mush around fast motion in a re-uploaded social clip. The upscaler treats artifacts as detail and sharpens them into crisper artifacts. Start from the least-compressed copy you can find, always.

Motion blur. Blur from movement is not missing detail, it is smeared detail. Sharpening it produces smeared-and-crunchy.

Footage that is already sharp enough for delivery. See the last section.

Run it at the right point

The order matters more than the settings, and there is one rule that saves the most money:

Upscale a still before you animate it, not the video afterwards.

Generators work from what you give them. A soft input produces mushy motion, and then you are upscaling four seconds of already-degraded video rather than one frame of clean source. Upscaling one image is dramatically cheaper than upscaling the clip it becomes — and the animation itself comes out better, which is the part that actually matters.

The reverse order is the single most common wasted spend I see in this workflow.

Two other pipeline notes:

Trim before you upscale. In Hiroska a trimmed timeline clip only processes its visible window, so cutting first means not paying to enhance footage you already discarded. Do the edit, then the enhancement.

Interpolation is a different tool. If the problem is stutter or a clip that is too short, that is frame interpolation, not upscaling. They get confused constantly because both are "make the clip better" buttons.

2× or 4×?

2× is the safe default. It roughly doubles the linear resolution and rarely introduces anything strange.

4× shines on genuinely small sources — an old photo, a thumbnail-sized still — and starts inventing more aggressively as the source gets larger. On an already-decent 1080p clip, 4× is mostly spending money to make a 4K file with imagined texture in it.

The token price for the run is shown before you start, which is worth actually reading when you are about to enhance a long clip.

When not to upscale at all

Worth saying, because it is the cheapest advice here.

If it is going to social, 1080p is almost always enough. The platforms re-encode aggressively, and the difference between your 4K master and your 1080p master is largely destroyed by their compression. Upscaling before an Instagram upload usually buys nothing but upload time.

If it is background footage. Nobody is studying the shot behind your title card.

If the clip is going to be small in frame. A picture-in-picture inset at a quarter size does not need 4K.

If the problem is not resolution. A badly composed, badly lit or badly timed shot is still all of those things at 4K. Upscaling makes things bigger and sharper; it does not make them good.

The honest summary: upscaling is a rescue tool and a preparation tool. As a rescue it is excellent, and it is the first thing to reach for before re-rolling a clip you liked. As a habit applied to everything on the way out, it mostly inflates your export times.