Image to video
Image to video AI, tested on ten photos
Image to video AI works reliably on stills with one subject, a simple background and no readable text, and it fails on faces held past about four seconds, on any lettering in frame, and on fast human motion. Pick the photo for the model rather than fighting the model with prompts.
Every image to video AI page tells you to upload your image and press generate. None of them tell you which images are worth uploading. That is the whole difficulty. The model is not deciding whether to do a good job, it is being handed a problem that is either easy or impossible, and the photo you choose decides which. Below are ten source photos run through current image to video models, seven that worked and three that did not, with the prompt printed under each clip.
The seven that worked
Each of these is a photo type, not a one off. If your source falls into one of these categories you can expect a usable clip on the first or second attempt.
Landscapes and skies
Nothing in the frame has a fixed identity the model has to preserve. Clouds and grass can move any way at all and still look right.
Products on a plain background
One rigid object, one clean background, no limbs. Ask for an orbit or a slow push and the shape holds. Logos and printed text are the exception, see the failures below.
Empty rooms and architecture
Straight lines bend if you push the camera hard, so keep the move small. A five percent dolly reads as expensive. A fifty percent dolly reads as a warp.
Food, steam and liquid
Steam, drips and small surface movement are the easiest wins in the whole category. The subject stays put and only the vapour has to be invented.
Texture and macro
Fabric, sand, dust, ink. High detail, low structure. The model has nothing to get wrong because there is nothing it must keep consistent.
Animals, if the motion is small
A head turn works. A run does not. The rule is the same as for people, only animals are more forgiving because we are worse at spotting a wrong dog than a wrong face.
Vehicles with a moving camera
Move the camera rather than the car. Reflections travelling across paint sell motion without asking the model to animate a complicated object.
The three that failed
These are kept in deliberately. Every competing page shows only successes, which is how people end up burning credits on shots that were never going to work.
Faces held longer than about four seconds
The identity survives the first couple of seconds and then drifts. Eyes go first, then the mouth, then the jaw. If your shot needs a face, cut at three seconds or generate a wider frame where the face is small.
Any readable text in the frame
Shop signs, packaging, screens, number plates. The letters re-form every frame and the word becomes a different word. There is no prompt that fixes it. Crop the text out and add it back as an overlay.
Fast human motion continued from a still
A skateboarder mid air, a dancer mid spin. The model has to invent a whole body trajectory and limbs separate on the landing. Ask for the moment before or after the action instead.
Camera moves, and what each word does
These five words change the output more than any adjective you can add. Most people write cinematic and hope. Naming the move is the difference between directing the shot and accepting whatever comes back.
| Term | What it means | When to use it |
|---|---|---|
| Pan | The camera rotates left or right from a fixed position. | Revealing a scene wider than the frame. Cheap and stable because the camera never moves through space. |
| Orbit | The camera circles the subject at a fixed distance. | Products. It is the single most reliable move on a rigid object and the one most people forget to ask for. |
| Dolly in | The whole camera travels toward the subject. | Building tension. Keep it small on interiors and architecture, where a hard dolly bends straight lines. |
| Zoom in | The focal length changes while the camera stays put. | A flatter, more clinical look than a dolly. Worth naming precisely, because if you just write closer the model picks for you. |
| Tilt | The camera pivots up or down from a fixed position. | Revealing height. Pairs well with architecture and tall products. |
A prompt shape that works for stills
Image to video prompts are shorter than text to video prompts, because the photo already carries the look. You are only describing motion. The shape that behaves is: what moves, how the camera moves, and what must stay still.
Template
[subject] [does one small thing], camera [named move] [direction and amount], [everything else] stays still.
Worked example: the steam rises from the cup, camera pushes in five percent, the cup and the table stay still.
Two habits are worth building. Ask for one motion, not three. And say what should not move, because the models over animate by default and a still background is something you have to request.
Resolution, length and what you get back
| Setting | Typical range | What to do about it |
|---|---|---|
| Clip length | 4 to 10 seconds | Generate longer than you need and cut. The best two seconds are usually the first two. |
| Output resolution | 720p or 1080p | Generate at the lower one while you are iterating, then re run the winning prompt higher. |
| Input aspect | Resized to the model's native ratio | Crop to 16:9 or 9:16 yourself before uploading. |
| Frame rate | 24 or 30 fps | Match it to the rest of your timeline before you edit, not after. |
| Takes per usable clip | 2 to 5 on the easy categories, more on faces | Budget in takes, not in clips. |
Where this fits
Image to video is one of two input modes on any video AI generator. If you want vertical output for social, the framing rules change and the useful ones are on the TikTok format page. If you have a written script and need footage to cut against it, that is the script to video workflow instead.
How this differs from generating from text
Starting from a photo is a smaller problem than starting from a sentence, and that shows up in the hit rate. The model is not deciding what the scene looks like. It has the scene. It only has to decide what happens next.
| From a photo | From text alone | |
|---|---|---|
| What is fixed | Composition, colour, subject, light | Nothing |
| What you control with the prompt | Motion and camera only | Everything, and therefore nothing precisely |
| Takes per usable clip | Fewer, on the easy categories | More, because the look changes every run |
| Consistency across shots | Good, if you feed related photos | Poor, every run reinvents the world |
| Best use | Product, place, anything that must match something real | Scenes that do not exist yet |
The practical consequence is worth stating plainly. If you can photograph the thing, photograph it and animate the photo. You will get closer to what you meant, in fewer takes, than by describing it.
Preparing the source photo
Five minutes on the input saves a dozen takes. Every item here is something the model cannot fix for you.
Crop to the output shape first
The tool will resize your photo to the model's ratio whether you like it or not. Crop it yourself so you decide what leaves the frame.
Remove text from the frame
Signs, labels, watermarks, packaging copy. If it is legible in the photo it will become illegible in the clip, and that is more noticeable than not having it at all.
Check the resolution
A small or heavily compressed source produces a soft clip. The model works from what you give it and does not restore detail that is not there.
Simplify the background
Busy backgrounds are where warping is most visible. A plain wall behind a product is worth more than any prompt adjustment.
Pick a frame with room to move
If the subject fills the frame edge to edge, there is nowhere for a camera move to go. Leave space in the direction you want to travel.
Troubleshooting
The failures in this category are consistent enough to be tabled. Find the symptom, apply the fix, re run.
| What you see | Why | What to change |
|---|---|---|
| The subject melts after two seconds | Too much motion asked for, or the subject is complex | Ask for a smaller motion, or move the camera instead of the subject |
| Nothing moves at all | The prompt described the scene rather than the motion | Rewrite the prompt as a verb. Say what happens, not what it looks like |
| The camera drifts when you wanted a locked shot | The default behaviour | Add static camera, no camera movement, explicitly |
| Colours shift through the clip | The model is regenerating the palette as it goes | Shorten the clip, and state the light in the prompt so it has something to hold on to |
| Extra objects appear | The model filling space it thinks is empty | Add exclusions. No people, no birds, no vehicles |
| The edges of the frame warp | A camera move larger than the source photo can support | Reduce the move, or start from a wider photo |
| Faces become someone else | Drift, and it is not solvable in the prompt | Cut before it happens, or frame the face smaller in the source |
Getting several clips to look like one shoot
If you are animating a set of photos for the same project, consistency comes from the inputs rather than the prompts. Shoot or select photos with the same light direction and the same lens. Then keep the motion vocabulary consistent too, so every clip in the set uses moves from the same family.
The mistake is to vary the camera move for variety. Six clips with six different moves feel like six different videos. Two moves used across six clips feels like a shoot.
Image to video questions
What is image to video AI?
Image to video AI takes a still photograph as the first frame and generates the frames that follow it. The photo constrains the look, and a text prompt describes what should move. It is more predictable than generating from text alone because the model has less to invent.
Which photos animate best?
Photos with one clear subject, a simple background, and no readable text. Landscapes, products on plain backgrounds, food, and texture close ups work first time. Faces, crowds, and fast action need several attempts and often do not land.
How long can an image to video clip be?
Usually four to ten seconds. Quality degrades as the clip runs on, so a five second generation you cut to three seconds is often better than a ten second generation you use whole.
Can I control the camera movement?
Yes, and you should. Name the move directly using pan, orbit, dolly, tilt or zoom. Vague instructions like make it cinematic let the model choose, and it usually chooses a slow push that fights whatever else you asked for.
Why does my photo change colour or crop when I animate it?
Most tools resize the input to the model's native resolution and aspect ratio before generating. If your photo is not close to that shape you lose edges. Crop to the output ratio yourself first so you decide what is lost.
Does image to video keep the exact product in my photo?
Mostly, for a few seconds, on a rigid object. Fine details such as printed labels, seams and small hardware drift. For anything where the product must be exact, keep the move short and check the last frame before you use it.