A build, not a list of tools
How to make AI videos, shown once from start to finish
Making an AI video is eight steps: idea, script, scene split, prompts, generation, voiceover, edit, export. The generation step is the one everyone writes about and the one that takes the least time. Below is each step in the order it happens, with what it decides and what skipping it costs.
Search how to make AI videos and most of what comes back is a product page answering a tutorial question. This is the tutorial: the eight steps in order, what each one decides, and where the time actually goes. The worked example throughout is a ninety second explainer. The build of that example for this page has not been published yet, so nothing below is dressed up as a measurement. When it publishes, each step gets its real timing log and its real take count, and the slots on this page fill with the takes it produced.
What gets made
The eight steps
Decide what the video is for
One idea, one audience, one length. Written in a sentence before anything else happens. The example carried through this page is a ninety second explainer for a general audience, made to sit on a page rather than in a feed.
Five minutes of work. Skipping it costs an hour later.
Write the script
Written as speech, out loud, in one pass. Ninety seconds is about two hundred and twenty words at a normal reading pace. Write long and cut, because cutting a script is easy and padding one is not.
Split it into scenes
Fourteen scenes for two hundred and twenty words is roughly sixteen words each. One idea per scene. Where a sentence carries two ideas it becomes two scenes, which is how fourteen scenes come out of eleven sentences.
This is the step that decides the video.
Write one prompt per scene
Subject, action, camera move, lens and light, duration. The same lens and light words go in all fourteen prompts, which is the whole trick to making separate generations look like one shoot. The grammar is set out on the prompt page, and the camera move for each scene can be lifted straight from the camera movement library.
Generate takes
Two or three per scene at the lower resolution while choosing, then the winners re run higher. Plan on roughly one usable clip in three on subjects without faces in them, so fourteen scenes is closer to forty takes than to twenty. The two slots below are for the kept take and the take before it, from the build of the example above.
The take that gets usedThe first usable take from the walkthrough build: a train pulling into an empty platform, camera static, evening light. The take before itThe take before it. Same prompt, the platform sign turns into nonsense letters and the train doubles. Kept in on purpose. Record the voiceover
Whole script in one pass. Recording line by line produces a voiceover whose pacing wanders, and pacing is most of what makes narration sound finished. Record two passes and use the second one, because the first is a warm up.
Cut the picture to the voice
Voice track down first, clips placed under the lines they belong to. Straight cuts, no transitions. Expect to drop a scene or two here because the video is better without them, which is normal and worth planning for.
Add captions, check, export
Captions burned in, one pass watching it end to end with sound on, then H.264 MP4 at 1080p. The end to end watch is not optional. Problems that are invisible clip by clip are obvious in sequence.
The accounting
The thing worth noticing about a build like this is that generation, the step the whole category sells, is not where the session goes. Writing, splitting and choosing take most of it, and none of those are things a model does for you.
A per step timing table belongs in this section and it is deliberately not here yet. Publishing invented minutes would make the rest of the page worth less. When the build of the example above runs, this section gets the real minutes for each of the eight steps, measured once on one video, so you can hold it against your own session rather than read it as a promise.
What goes wrong
- Signs and screens rewrite themselves. A scene with readable text in frame burns takes and never settles. Reframe the prompt so the text is out of shot rather than re rolling it.
- The look drifts. Colour wanders as you work down the scene list. Copy the lighting words out of scene one into every later prompt instead of writing fresh ones.
- The first cut is too fast. Fourteen scenes across ninety seconds is a cut every six seconds, which reads as frantic. Drop two scenes and hold two longer.
- Frame rates do not match. Clips from different tools come back at different rates and stutter at the cut. Normalise at import, not after the edit.
Do the second one faster
Almost all of that first session is first time cost. Keep the prompts, keep the lens and light words, and keep the scene lengths that worked. A second video on the same look is a fraction of the first, and most of what is left is writing.
Four things carry over and they are worth saving deliberately rather than rediscovering.
- The look sentence. The lens, light and colour words that appear in every prompt. Paste it into every new prompt before you write anything else.
- The scene lengths. Once you know your delivery runs at about fifteen words per four seconds, a script converts into a scene count immediately.
- The model choices. Which model handled which kind of shot. That mapping is the most expensive thing you learned and it takes one line to write down.
- The rejected takes. Keep them in a folder. Plenty of failed hero shots work as a two second cutaway later.
Doing it on a free tier
The steps above assume you can generate freely. On three generations a day it is still possible and it takes longer, so the order changes. Do not try to make the video in order.
Spend the first week on prompts, not clips
Nine seconds a day is enough to learn whether a prompt structure works. Test structures, not scenes, and write the results down.
Write the script to the shots you can get
Once you know which subjects come back reliably, write a script built out of those. That is backwards from how you would normally work and it is the right order when takes are scarce.
Generate the hardest scenes first
If a scene is going to fail, find out on day one rather than after you have built the rest of the video around it.
Reuse shots
The same clip can appear twice in ninety seconds if it is short and the second use is reversed or cropped differently. Nobody notices, and it halves your generation count.
Adapting the same build for a short
If the output is vertical rather than a page video, three things change and the rest is identical. Generate at 9:16 rather than cropping, cut the script to about seventy words, and move the strongest shot to the front instead of using it as a reveal. The full format rules are on the vertical page.
Where to go next
For the input mode that starts from a photo instead of a sentence, see image to video. For filling the gaps in a cut, see generated B-roll. For the script side of the same workflow, see script to video. For what the models can and cannot do at all, the overview is the video AI generator page.
What you need at each step
Deliberately generic. Every step here can be done with several tools and none of them is the point.
| Step | What you need | What you do not need |
|---|---|---|
| Script | A text file and somewhere quiet | A script generator. Writing two hundred words is faster than fixing generated ones. |
| Scene split | A table with four columns | Any software at all |
| Generation | A tool with access to more than one model | The most expensive model, at least while you are choosing |
| Voice | A microphone, or a voice generator | A treated room. A quiet one is enough. |
| Edit | Any timeline editor | Transitions, effects, colour grading suites |
| Captions | Automatic captions plus a read through to fix names | Hand typing them |
The checklist
Print this once and you will not need it again after the second video.
- The script is written and read out loud before any prompt is written.
- Every scene is under six seconds.
- Every prompt names a camera move.
- The lens and lighting words are identical across every prompt.
- No prompt asks for readable text in frame.
- Every clip is normalised to one frame rate at import.
- The voice track is one continuous recording.
- The whole video has been watched end to end with sound before export.
- The AI content label is set wherever you are uploading it.
What this costs
In money it depends entirely on which model you pick and how many takes you need, so any figure here would be wrong within a month. In takes it is steadier, and takes are the number to plan with, because a take you throw away costs exactly what a take you keep costs.
Keep rate is what decides that number and it moves with the subject rather than with the tool. Landscape, texture and abstract shots come back usable far more often than products and interiors, those come back more often than faceless human activity, and faces, crowds and fast motion are worse than all of them. A measured keep rate table, generated on this site rather than estimated, goes here once the build has run. Until then, plan on several takes for every clip you keep and generate the hardest subjects first so you find out early.
The subject that should change your plan is faces. If your video needs them, either film them or rewrite the shots so the faces are small, brief, or turned away. It is cheaper to change the script than to fight the model.
Questions about the process
How do you make an AI video?
Write a short script, split it into three to six second scenes, write one prompt per scene, generate two or three takes of each, record a voiceover, cut the clips against the voice, and export. The generation is the fast part. The scene splitting and the picking are where the time goes.
How long does it take to make a one minute AI video?
Plan on a few hours of working time for ninety seconds the first time round, most of it spent generating takes and choosing between them rather than on the generation itself. That drops sharply on the second video because your prompts and your look are already settled.
Do I need editing software?
You need something that can lay clips against an audio track. Any timeline editor does it, including the free ones. What you do not need is anything advanced, because an AI video edit is straight cuts and a voice track.
How many takes does each shot need?
Two or three on easy subjects like landscapes, texture and products. Five or more on anything with a face or fast motion. Budget in takes rather than in clips, because that is what you pay for.
Can one prompt make a whole video?
No. Current models generate four to ten seconds at a time, so every longer video is several generations cut together. Any tool that appears to make a whole video from one prompt is splitting it into scenes for you behind the scenes.
What is the most common mistake first time?
Writing prompts before splitting the script. If you prompt from the script directly you get scenes that are too long and shots that do not cut together. Split first, prompt second.
Do I have to disclose that a video is AI generated?
Most platforms now expect it and provide a toggle or a label field when you upload. Use it. The cost of labelling is nothing and the cost of being caught not labelling is your channel.