AI video tools, decoded: what all those names actually are

There are maybe a dozen real video models and hundreds of products built on top of them. Once you can tell which is which, the whole category stops being a wall of names and starts being a short list of decisions.

If you have spent any time searching for AI video tools, you have hit the wall of names. Veo, Kling, Sora, Seedance, Hailuo, Hunyuan, LTX, Wan — and then Pika, Higgsfield, Viggle, Vidnoz, Pollo, Digen, Crayo, and a hundred more arriving monthly.

It looks like a hundred competing products. It is mostly three layers stacked on top of each other, and once you can tell them apart the category gets much smaller.

Layer 1: the models

A handful of large organisations train the actual video models. Everything else in this space is built on their output. This layer moves slowly enough to be worth learning:

Model family Made by
Veo Google
Sora OpenAI
Kling Kuaishou
Seedance (video), Seedream (images) ByteDance
Hailuo MiniMax
Wan Alibaba
Hunyuan Tencent
LTX Lightricks
Grok Imagine xAI
Flux Black Forest Labs

When you read that a tool "uses Veo 3.1", this is the layer being referred to. The model determines what the output can look like. Nothing above it changes that.

Most families ship several tiers — a fast, cheaper variant and a slower, better one — which is why you see names like "Fast", "Pro", "Turbo" and "Standard" attached. Those are not marketing gloss; they are genuinely different price and quality points, and picking the wrong tier for the job is the most common way people overspend.

Layer 2: the products

Everything else. If a name is not in the table above, it is almost certainly an interface over one or more models in it.

That is not a criticism — a good interface is worth paying for, and nobody wants to write raw API calls to make a video. But it changes what you are actually comparing. Two products can produce identical output because they are calling the same model, and differ entirely in price, speed and what happens after the clip is generated.

So the questions worth asking a product are not about quality in the abstract:

  • Which model does it use, and which version? Many do not say. A tool quietly running a year-old model is a real thing, and you cannot tell from the output samples on the landing page — those were cherry-picked whatever the model.
  • Can I choose? Being locked to one model means being locked to its weaknesses.
  • What happens to the clip after it is generated? This is the one people discover last and regret most. See below.
  • What does a re-roll cost? You will do several. The per-clip price matters far less than the price of being wrong four times.

Layer 3: the single-purpose tools

A third group does not generate video at all — it does one job to video you already have. These names come up constantly and confuse the picture because they are not competing with the models:

Motion transfer — take the movement from a reference video and apply it to your character. If you have searched for Viggle, this is the job you are looking for.

Upscaling — increase resolution and sharpness on existing footage. Topaz is the name most people know here. Worth understanding that upscaling adds plausible detail; it cannot recover what the camera never captured.

Frame interpolation — generate the in-between frames to smooth motion or make a short clip longer. The cheapest fix in the entire category and the least known.

Lip sync — match a face's mouth to an audio track.

Background removal and object erasure — take something out of the frame.

These are all things we run alongside generation, because the single most useful thing you can do with a clip that is nearly right is fix it rather than roll the dice again.

Why "which one is best" has no answer

Three reasons, and they are all structural rather than evasive:

The ranking changes monthly. Any leaderboard you read is a snapshot. Building a workflow around one model being best is building on sand.

Different models are good at different things. The property that most changes how you work is not fidelity — it is how much of the shot you are allowed to specify. Some models take a start frame only. Some take a start and an end frame, which lets you pin where a shot lands instead of hoping. That distinction narrows your choice faster than any quality comparison.

Most products use the same models anyway. If two tools call the same endpoint, comparing their "AI quality" is comparing nothing. Compare the workflow around it.

What this means in practice

Learn the model layer, ignore most of the product layer, and pick tools on what happens around generation rather than on generated samples.

Concretely, the thing that separates a usable workflow from a frustrating one is what happens when shot three of eight is wrong. If fixing it means going back to a browser tab, regenerating, re-downloading and re-importing, that round trip is the actual cost of the tool — not the price per clip.

We take the same view internally, which is why the generator sits inside the editor rather than beside it, why you can swap between models per shot instead of being locked to one, and why upscaling and motion transfer are there next to the generate button. A bad shot should be a re-roll in place, with the rest of the cut untouched.

If you want the specifics — which model does what, which accept an end frame, what each costs — every model we run has its own page under models, and tokens explains how the pricing works.