Image prompts describe a frame. Video prompts describe a frame plus time — and time is where most prompts fall apart. The mental shift that fixes it: stop writing like a painter and start writing like a director giving instructions to a camera operator and one actor.
The formula
Camera language + Subject action + Scene + Lighting + Style- Camera language — how the shot moves: push in, orbit, static
- Subject action — one clear thing the subject does
- Scene — where it happens and what surrounds the subject
- Lighting — time of day, source, mood
- Style — cinematic, anime, documentary, stop motion…
Compare:
Weak: a woman walking in a city at night, cool vibes
Strong: slow tracking shot following a woman in a red coat as she
walks past neon storefronts, rain-slicked street reflecting
the signs, shallow depth of field, moody cinematic styleThe strong version answers the three questions every video model is silently asking: where is the camera, what moves, and what does it look like.
Camera language glossary
Models are trained on captioned footage, so real film terminology lands better than improvised descriptions. The core vocabulary:
| Term | 中文 | What it does |
|---|---|---|
| Static shot | 固定镜头 | Camera locked off; only the subject moves |
| Push in / dolly in | 推镜头 | Camera moves toward the subject; builds focus |
| Pull back / dolly out | 拉镜头 | Camera retreats; reveals context |
| Pan | 摇镜头 | Camera rotates left/right from a fixed point |
| Tilt | 俯仰镜头 | Camera rotates up/down from a fixed point |
| Tracking shot | 跟踪镜头 | Camera travels alongside a moving subject |
| Orbit / arc shot | 环绕镜头 | Camera circles the subject |
| Crane shot | 升降镜头 | Camera rises or descends vertically |
| Handheld | 手持镜头 | Slight shake; documentary energy |
| Aerial / drone shot | 航拍镜头 | High, wide, sweeping perspective |
| Slow motion | 慢动作 | Stretched time; emphasizes detail |
| Rack focus | 变焦点 | Focus shifts from one plane to another |
One camera instruction per prompt. "Orbit while pushing in and then crane up" reads like three shots; the model will usually mangle all three.
Describe motion the model can actually do
Motion is the hardest thing for video models, so be deliberately conservative:
- One subject, one action. "A chef flips a pancake" works. "A chef flips a pancake while a waiter weaves between tables and a dog chases a cat" produces limbs in the wrong places.
- Pick continuous verbs. Drifting, flowing, rotating, walking, rippling sustain well across a clip. Sharp state changes — a vase shattering, a card trick — are much harder.
- Make the action self-explanatory. "She reacts to the news" means nothing visually. "She covers her mouth and her eyes widen" is filmable.
- Let the environment move too. Drifting fog, falling leaves, flickering signage add life at low risk, because nothing has to stay anatomically coherent.
Think in short clips
Current models generate seconds, not scenes. Write each prompt as a single beat — one camera move, one action arc that completes within the clip. If your idea needs a beginning, middle, and end, split it into separate generations and cut them together afterwards, exactly as an editor would. Prompts that try to compress a story ("a man wakes up, gets dressed, misses his bus, and meets a stranger") return a blur of half-finished gestures.
A good test: can you act out the entire prompt in five seconds? If not, cut it down.
Image-to-video: describe the motion, not the picture
When you start from a still frame, the image already carries the subject, scene, lighting, and style. Re-describing them wastes your prompt and can fight the source image. Describe only what should move:
Weak: a beautiful mountain lake at sunset with pine trees and
orange clouds, cinematic
Strong: slow push in toward the lake, clouds drifting right,
gentle ripples spreading across the waterEverything in the weak prompt is already in the image. The strong prompt spends every word on motion. This is also the cheapest way to get reliable results: generate a strong still in the image generator, iterate until the frame is right, then animate it in the AI video generator.
Five complete examples
Cinematic product shot:
slow orbit around a matte black wireless earbud case on a stone
pedestal, soft studio light, dust particles drifting in a single
beam, dark background, premium commercial styleDocumentary nature:
aerial drone shot gliding over a winding river at dawn, mist
rising from the water, golden light breaking over forested hills,
naturalistic documentary styleAnime mood piece:
static shot of a girl standing on a train platform at dusk, her
hair and scarf moving in the wind, lights of a passing train
streaking behind her, anime style, soft pastel paletteStreet photography energy:
handheld tracking shot following a skateboarder weaving through
a narrow alley, late afternoon sun flaring between buildings,
gritty 35mm film lookMacro abstract:
extreme close-up of ink dropping into clear water, tendrils of
deep blue unfurling in slow motion, white background, studio
lighting, minimal and cleanEach example is one camera idea, one motion idea, one style. More finished prompts in this format live in the prompt library.
FAQ
How long should a video prompt be? Shorter than an image prompt — roughly 15 to 40 words. Video models juggle space and time at once, and every extra clause is another thing that can drift. Spend your words on camera and motion.
Why does my subject morph or deform mid-clip? Usually the prompt asks for more motion than the model can keep coherent — multiple subjects, fast actions, or conflicting camera moves. Simplify to one subject and one continuous action, or switch to image-to-video so the model has a stable frame to anchor on.
What do I need to generate video on Bno AI? Just enough credits — video is charged per clip, not locked behind a subscription. The free tier's 10 daily credits cover one budget clip (Grok Imagine 1.5, 6 seconds at 480p, 9 credits) or about five GPT Image 2 images at 2 credits each for perfecting your starting frames. Heavier models cost more per clip, so for regular video work see pricing — Pro starts at $20/month with 2,000 credits.
