How to write a text-to-video prompt

Most prompts describe a photograph and then wonder why the video does nothing. A video prompt has a time axis, and there are exactly five slots worth filling in it.

You typed "a woman walks through a market at sunset". You got a woman, a market, a sunset, and a shot that just sort of sits there.

The instinct at this point is to add words. More adjectives, more atmosphere, a comma-separated tail of cinematic, 8k, masterpiece, highly detailed. It almost never helps, because the problem was never length. The problem is that you described a photograph, and then asked a video model to make it move — leaving the single most important decision, what changes over the next five seconds, entirely up to the model.

A prompt is a shot description, not a picture description

Everything you have learned about image prompts transfers, except the one part that matters. An image prompt describes a state. A video prompt has to describe a change — and if you don't specify one, the model picks. That is where the drifting, wobbling, weirdly aimless clip comes from.

The fix is unglamorous: say what moves, and say what the camera does.

The five slots worth filling

Think of a prompt as five slots. You don't need all five, but the ones you leave empty are the ones the model gets to invent.

Subject — concrete, and only one

"A woman in a red raincoat, mid-thirties, wet hair" beats "a beautiful woman". Specific nouns constrain the model; evaluative adjectives don't. And keep the count low: two subjects doing two things is where identity starts smearing between them.

Action — one action, in progress

One. A single continuous action, already happening when the clip starts. "Pours coffee, then looks up, then smiles" is three shots wearing one prompt, and you'll get a mush of all three.

Prefer the present participle: pouring, turning, reaching. It reads as an ongoing state rather than an event that needs a beginning and an end crammed into five seconds.

Camera — the slot everyone skips

This is the highest-leverage word in the prompt. "Slow dolly in", "static tripod shot", "handheld follow", "slow pan left", "low angle looking up". A model given no camera instruction will usually add a lazy drift, which is the visual signature of AI video that people have learned to recognise and dislike.

If you want nothing to move, say so: "static camera, locked off". Stillness is an instruction too.

Light — time and direction

"Late afternoon sun from the left, long shadows" does more work than "cinematic lighting". Light is also your continuity tool: if this shot has to cut against another one, matching the light in the prompt saves you a grading pass later.

Look — lens and medium, briefly

"Shallow depth of field, 35mm", "shot on 16mm film, visible grain", "crisp digital, deep focus". One or two of these. This is the slot where prompt bloat lives, and past about three descriptors they start fighting each other.

Things that don't do what you think

Negatives written as prose. "No text, no distortion, not blurry" often adds the thing you named, because the words are in the prompt. If a model exposes a real negative-prompt field, use that field. If it doesn't, describe the positive instead: "clean empty wall" rather than "no signs".

Counting. "Five birds" gets you some birds. Numbers above about three are suggestions.

Text in frame. Video models render text worse than image models do, and worse still while it moves. Put your words on with a title layer, where they'll be sharp, correct, and editable.

Cuts and sequences. "She enters, then the camera cuts to her face" asks one generation to be two shots. Generate two clips and cut them on the timeline — you'll get both shots, and you'll get to choose the cut point.

Choreography. Specific movement can't be written down in prose; that's why dance notation exists. If you need a particular routine or gesture, that's a motion transfer job, not a prompting job.

Style-word stacking. Cinematic, hyperrealistic, award-winning, trending on ArtStation — inherited from image-model folklore, mostly inert here, and it dilutes the words that were actually doing something.

So how long should it be?

One or two sentences, in the order above: subject, action, camera, light, look. Roughly 20–40 words. Long enough that all five slots are filled; short enough that every word is load-bearing.

If a prompt isn't working, the productive move is almost never "add more". It's to find the slot you left empty.

When to stop prompting and start seeding

There's a hard ceiling on what a text prompt can control, and you hit it faster than you'd like. You cannot prompt your way to a specific face, a specific product, or the same room you generated yesterday.

That's the moment to switch recipes. Generate a still first — images are cheaper, faster, and far more controllable — then animate that still. The image locks composition, wardrobe, face and set; the prompt is then only responsible for the motion, which is the one job it's good at.

Anything that has to appear in more than one shot should start life as an image.

The order that works

  1. Decide the shot before you type. What's in frame, what moves, what the camera does. A prompt is a transcription of that decision, not a substitute for it.
  2. Fill the five slots in one or two sentences.
  3. Name the camera explicitly, even when the answer is "static".
  4. Run variants, not re-rolls. Four variants of one prompt teach you what the model does with it; four rewrites teach you nothing, because you changed the prompt and the seed at once.
  5. Keep it short — a few seconds. Quality degrades with length in almost every model, so generate the beat you need and build the sequence on the timeline.
  6. Fix, don't re-roll. A clip that's 80% right is a trim, an upscale and a grade away from usable. Generation is a draw from a distribution — the take you liked isn't coming back.