KROMALOCA TRANSMISSION

AI Video Prompts: How to Direct a Shot Instead of Describing a Scene

Learn how AI video prompts differ from image prompts, including camera movement, subject motion, timing, continuity, shot structure and visual restraint.

AI Video Prompts: How to Direct a Shot Instead of Describing a Scene

You can make a beautiful image with a few carefully chosen words.

Video is less forgiving.

You can describe a woman standing on a rooftop beneath a red moon and get something spectacular. Then you add: “She slowly turns around while the camera circles her.” And suddenly the moon moves. Her hair changes length. The building develops a second floor. Her face decides it belongs to somebody else. The camera starts travelling in a direction nobody requested.

Welcome to AI video.

The image was obedient. The moment you gave it time, it developed opinions.

That’s the fundamental difference between generating an image and generating a video. An image only has to get one moment right. A video has to survive the moments between the moments. And that changes how you should write your prompts.

An image describes a state. A video describes a change.

Take an image prompt:

A lone astronaut standing beside a damaged rover on Mars at sunset.

The model only has to build one visual state. Now turn it into video:

A lone astronaut stands beside a damaged rover on Mars at sunset. The astronaut slowly walks toward the rover and places one hand on its damaged panel while dust moves gently across the ground. The camera makes a slow lateral tracking movement, keeping the astronaut and rover in frame.

Now the model has a problem the image version never had. It has to decide what the astronaut does first, how quickly they move, what the dust does, how the camera moves, what stays still, how the light changes, and what happens between the beginning and the end of the shot.

Video prompting is less about describing what exists and more about describing what happens. Google’s, Adobe’s and Runway’s own prompting guides all lean the same way: they ask for camera movement, subject action, shot framing and how the scene develops over time, not just a static description with motion bolted on.

how AI video prompts differ from image prompts

Don’t write a video prompt like a poster

A common mistake:

A beautiful futuristic city, cinematic, ultra detailed, neon lights, atmospheric, cyberpunk, dramatic, 8K.

That might make an interesting still. But what is the video supposed to do? Does the camera move? Does a vehicle pass? Does rain fall? Does anyone walk? Does anything happen? The prompt describes a place. It doesn’t describe a shot.

Try this instead:

A dense futuristic city at night, viewed from street level beneath elevated transit rails. Rain falls continuously onto reflective pavement while pedestrians move through the foreground. The camera slowly pushes forward along the street, passing beneath a flickering sign and revealing a massive illuminated tower through the rain.

Now there’s time in it. Something happens at the start, something happens during the shot, something gets revealed. The camera has a job. That’s much closer to directing.

Think in shots, one job at a time

A scene might contain:

A woman enters a hotel, walks across the lobby, speaks to the receptionist, takes an elevator, exits on another floor and discovers a strange room.

That’s a treatment, not a prompt. Asking one generation to solve all of it is asking the model to solve several different shots at once. Break it apart instead:

Shot 1: A woman enters an enormous empty hotel lobby at night. The camera remains static as the automatic doors close behind her.

Shot 2: She walks slowly across the marble floor toward the reception desk. The camera tracks backward in front of her.

Shot 3: She presses the elevator button. The camera holds a close-up as the indicator changes from 7 to 6 to 5.

Shot 4: The elevator doors open onto a dark corridor. The camera slowly pushes forward as she hesitates at the entrance.

Each generation now has one job. If Shot 3 fails, you fix Shot 3. You don’t regenerate the whole sequence.

Don’t ask the machine to direct your entire film in one breath. Even machines need an edit.

breaking an AI video idea into shots

Give the subject one important action

If you hand a six-second clip five unrelated actions, you’re asking the model to solve several problems at once:

A man walks toward the car, opens the door, gets inside, starts the engine, checks his phone, looks behind him, reverses and drives away.

That’s a miniature screenplay. Simplify:

A man walks toward a parked black car and opens the driver’s door. The camera follows slowly from behind.

Now the action is clear, and you can make the next beat its own shot:

The man sits inside the car and starts the engine. The dashboard lights illuminate his face as the camera slowly moves closer through the windshield.

One action, one shot, one problem to solve. That doesn’t mean a video can only ever contain one movement — it means you should know what the primary action is before you add the rest.

Camera movement is direction, not decoration

“Cinematic camera movement” sounds impressive and tells the model almost nothing. Compare:

Cinematic camera movement around the character.

with:

Slow clockwise orbit around the character while maintaining a medium shot.

The second gives the camera an actual job. Google’s, Runway’s and most other current video-model guides converge on the same vocabulary: shot size (close-up, medium, wide), camera angle (low, high, eye-level) and camera movement (push-in, pull-back, dolly, tracking, pan, tilt, crane, orbit, handheld, locked-off). Learn a dozen of those terms and you can direct almost anything.

The catch: don’t give the camera six jobs at once.

The camera slowly pushes forward while orbiting clockwise, then cranes upward, pans left, zooms out and tracks the character from behind.

You’ve asked one camera to be a helicopter, a dolly, a drone and a handheld operator simultaneously, and the result is often unpredictable. Runway’s own Gen-4 guidance is explicit about this: pick one camera movement per generation and split anything more complex into separate shots. So instead:

Slow push-in toward the character. Keep the character centred throughout the shot.

One primary movement, one visual intention. That’s not a limitation — it’s direction. And the movement should have a reason to exist. Ask what it reveals. A push-in builds anticipation. A pull-back reveals scale. An orbit reveals shape. A static camera makes the subject’s own movement matter more.

A huge machine stands motionless inside a dark chamber. The camera begins close to one mechanical component and slowly pulls backward, revealing that the component is only a small part of an enormous structure.

Here the camera movement is telling the story, not decorating it.

camera movement and video prompts in the Video Prompt Vault

Sometimes the best camera movement is none at all

A locked-off camera can be devastatingly effective.

Static wide shot of an abandoned railway station at night. Rain falls through the broken roof while a single distant train slowly approaches through the darkness.

Nothing is wrong with the camera staying still. The stillness is what makes the train’s movement noticeable. If everything moves, nothing feels important.

Motion is cheap. Attention is not.

Say what moves and what doesn’t

This is especially useful for busy scenes. Take:

A huge stone statue standing in a forest while fog moves around it.

Strengthen it by defining the relationship directly:

The statue remains completely motionless while thin fog slowly moves around its base and tree branches sway gently in the wind.

Now the model has an anchor. The statue stays. The fog moves. The trees move. Those are separate, named behaviours, and you can do the same with a character:

The character remains still while the camera slowly moves around them.

or the reverse:

The camera remains locked while the character walks from left to right through the frame.

Naming what stays fixed is often as useful as naming what moves.

Temporal language matters

Video happens across time, so use language that tells the model how a shot progresses: slowly, gradually, suddenly, begins to, continues to, over the course of the shot, by the end of the shot.

Compare:

A door opens.

with:

The heavy metal door slowly begins to open, revealing a narrow line of intense white light that gradually expands across the dark room.

The second has a beginning, a progression and a reveal — a small timeline instead of a snapshot. Several video-model guides, including Google’s Veo documentation, structure prompts this way by default. Adobe’s Firefly video tool works a little differently: each generation starts fresh from the prompt rather than carrying context from the previous one, so “iterating” there usually means writing a new, adjusted prompt or using Firefly’s dedicated camera and shot-size controls, though its newer prompt-based editing feature now lets you nudge an existing clip (swap a background, adjust lighting) without a full regeneration.

The first frame matters more than you think

Before describing movement, establish the starting state. What does the viewer see when the video begins?

Wide shot of a deserted desert highway at dawn. A vintage car is parked on the shoulder, facing toward the horizon.

Then add the motion:

The car slowly begins moving toward the horizon as the camera tracks alongside it.

That’s easier for the model to follow than opening with “a cinematic car drives through the desert while the camera tracks it,” which leaves the starting position undefined. Establish the state, then start the action.

This matters even more with image-to-video. If you’re animating an existing image, it has already defined the composition, subject, colours, lighting and camera position — repeating all of that in the prompt is often wasted words. Focus on what changes:

The woman remains in the same position and keeps the same expression. Her hair moves gently in the wind while the camera slowly pushes toward her.

The image establishes her. The prompt establishes the change.

image-to-video prompting

Don’t animate everything

A beautiful still image, followed by “animate everything,” is a classic trap. Now the clouds move, the walls breathe, the furniture slides, the clothes shift, the lights flicker, the camera spins, the moon grows. You’ve accidentally opened the portal.

Choose what deserves motion and protect the rest:

Keep the architecture completely static. Only the hanging cables move gently in the air, while a faint glow pulses through the central machine.

A character can stay still while the world moves around them. A machine can stay still while light travels through it. Movement means something when it has something stable to move against.

Physics is part of the prompt

If you want realistic motion, describe a mechanism the movement can believe in.

A heavy metal door opens rapidly upward.

may struggle without support. Try:

A massive hydraulic door slowly lifts upward, with visible mechanical pistons extending as it opens.

Or compare “a curtain moves gently in a light breeze” with “the curtain dramatically dances through the room.” The first is filmable. The second is poetic, which is a different job.

Light can carry the same kind of motion:

The machine remains still as a narrow internal light slowly travels through its circular mechanisms, illuminating one section after another.

or:

A train passes outside the room, causing bands of moving light to sweep across the walls.

The scene feels alive without the whole environment having to move.

Continuity is the real monster

Generating one impressive clip is easy. Generating five clips that look like they belong to the same world is much harder.

Shot 1: a woman in a red coat enters a hotel. Shot 2: she’s suddenly in black. Shot 3: the hotel has different architecture. Shot 4: it’s daytime now. Shot 5: her hair is completely different. Each clip might look great on its own. Together, they look like five unrelated dreams.

If continuity matters, keep track of character appearance, clothing, environment, lighting, time of day, camera language, colour palette, important props and spatial relationships — and use reference images or your model’s consistency features when they’re available.

Editing your way to the right shot

Good AI video prompting is mostly editing, not generating from scratch each time. Look at what’s wrong, and fix one thing.

Camera’s right, but the character walks too fast:

Keep the existing composition and camera movement. Slow the character’s walking pace considerably. Everything else remains unchanged.

Character’s right, but the background keeps drifting:

Preserve the character and environment exactly. Keep the architecture static throughout the shot. Only the character and falling rain should move.

If your clip has five problems at once — wrong camera, wrong lighting, wrong movement, wrong background, wrong colour — resist the urge to rewrite everything. Fix one, generate, see what changed, fix the next. That’s slower than rewriting the whole prompt, but it’s the only way to learn what your particular model actually responds to, since camera terms like “dolly” or “orbit” carry across models but can behave a little differently in each one.

A useful shot anatomy

Not a formula, just questions worth answering before you generate: What’s visible (subject and environment)? What’s happening (the primary action)? Where’s the camera, and how does it move (pick one movement)? What else moves — rain, smoke, fabric, light, hair, crowds? What stays stable? What changes over the course of the shot? What does the viewer discover, and where does the shot end?

You don’t need to write those headings into the prompt. They’re a thinking tool.

From vague prompt to directed shot

Version 1: “A futuristic robot walking through a city. Cinematic and realistic.” Tells us almost nothing about the shot.

Version 2: “A realistic humanoid robot walks through a futuristic city at night. Neon signs illuminate wet streets. The camera follows behind the robot.” Better.

Version 3:

A tall humanoid service robot walks alone through a rain-soaked futuristic city at night. Its brushed-metal body reflects the coloured lights of the surrounding buildings. The camera tracks slowly behind it at waist height, maintaining the same distance throughout the shot. Rain falls steadily while distant pedestrians move through the background. Near the end of the shot, the robot stops and looks toward a massive illuminated structure at the end of the street.

Now there’s a subject, an environment, an action, a camera position, a camera movement, atmospheric motion, background activity and a final beat. Notice we didn’t need fifty cinematic adjectives. We needed direction.

Sound is another dimension

Some video generators can produce or incorporate sound, which is another layer to direct — but don’t write a screenplay’s worth of audio into every prompt. If sound matters, name the important event: “heavy rain, distant traffic and the low mechanical hum of the machine,” or “no music, only footsteps, room ambience and the quiet sound of the ventilation system.” Tell the model what matters. Leave the rest alone.

What actually makes a video feel cinematic

Not the word “cinematic.” That word has become almost decorative. The feeling comes from deliberate camera movement, controlled framing, foreground depth, motivated lighting, a restrained palette, purposeful subject movement, environmental motion, a clear visual hierarchy, a beginning and ending state, and something being revealed. Describe those, and you may get the feeling without saying the word at all.

If you have to keep telling the machine that something is cinematic, it may be time to give the camera something interesting to do.

The strangest trick: remove movement

If your AI video looks obviously AI-generated, try taking things away. Remove the unnecessary camera movement. Remove the spinning. Remove the extra particles, the dramatic zoom, the five simultaneous actions. Leave one beautiful thing moving.

Locked-off camera. An enormous machine remains completely still in the darkness. Only a thin line of violet light slowly travels around its inner ring.

That’s often enough. Restraint can create more tension than spectacle.

One final example

The idea: “A machine wakes up inside an impossible room.” Fun, but vague.

The shot:

An enormous circular machine rests in the centre of a black geometric chamber with no visible walls or ceiling. The machine remains completely motionless for the first few seconds. A thin line of violet light begins travelling around its inner ring, gradually accelerating until the entire structure becomes illuminated. The camera performs an extremely slow push toward the centre, maintaining perfect symmetry. As the light reaches the final segment of the ring, a small opening appears at the centre, revealing a vast field of stars beyond it.

Now there’s a beginning, a progression, an escalation, a camera movement, a reveal and an ending. That’s the difference between “make a trippy video of a machine” and directing this particular moment.

The machine doesn’t need more adjectives. It needs choreography.

The real skill is learning what not to animate

AI video hands you an enormous amount of freedom, and that’s the actual problem. Everything can move, so everything starts moving — camera, character, background, light, clouds, even the floor. The shot turns into a screensaver having an existential crisis.

Good direction usually runs the other way. Decide what deserves motion, then protect everything else. A character can stay still while the world moves around them. A single door can open while an entire room stays frozen. Movement becomes meaningful when it has something stable to move against.

The next time you write a video prompt

Don’t start with “make a cinematic video.” Start with: what does the viewer see at the beginning? What changes? What stays stable? Where’s the camera, and what does it do? Where does the shot end?

Answer those and you already have the bones of a useful prompt. The rest is refinement — and yes, the machine will still occasionally decide your perfectly normal human has three elbows. You fix it, generate again, keep the good parts, move on.

AI video generation isn’t about finding the one perfect sentence. It’s about learning to direct something that doesn’t quite understand what you mean yet. That’s a much more interesting skill.

Further reading

Video-model prompting keeps changing as the models do, so it’s worth checking the source documentation alongside your own testing:

  • Veo prompt guide, Vertex AI (docs.cloud.google.com/vertex-ai/generative-ai/docs/video/video-gen-prompt-guide) covers prompt anatomy, camera movement and shot composition.

  • Runway’s help centre (help.runwayml.com/hc/en-us/articles/39789879462419) documents camera-motion terminology and style descriptors for its Gen-4 family.

  • Adobe Firefly’s video guide (helpx.adobe.com/au/firefly/work-with-audio-and-video/work-with-video/generate-videos-using-text-prompts.html) covers shot size, camera angle, motion presets and reference-based control.

Try a guide’s advice, break it, change one variable, generate again, watch what happens. That’s where the useful stuff starts.

Official references

Explore KROMALOCA