Vertical format
AI video generator for TikTok: the actual spec
Every page about using an AI video generator for TikTok says viral and stops there. The spec is simple and nobody prints it: 1080 by 1920, motion already happening in the first frame, sound on, the bottom quarter and right edge kept clear, and four to eight generated clips cut into twenty to thirty five seconds.
Vertical is not landscape turned sideways. The composition rules are different, the attention curve is different, and a chunk of the frame belongs to the app rather than to you. Generated footage makes all of that harder, because the model composes for the frame you asked for and knows nothing about the buttons that will sit on top of it. This page is the spec sheet, then the prompt changes that follow from it.
The spec
- Frame1080 by 1920, 9:16. Generate at this ratio, do not crop into it.
- Frame rate30fps is the safe default. Match every clip before you cut them together.
- First frameMotion already in progress. Never a static establishing shot.
- SoundAlways. Native audio, a voice track, or at minimum an ambience bed.
- LengthTwenty to thirty five seconds for one idea. Four to eight generated clips.
- Safe areasKeep the bottom quarter and the right edge free of anything that matters.
- Subject placementUpper middle third. That is the part of the frame nothing covers.
- ExportH.264 MP4, no watermark from whatever tool made it.
Safe areas, honestly
TikTok does not publish exact safe area figures and the layout shifts between app versions and device sizes. What follows is the working approximation, and the right approach is to leave more room than the numbers suggest.
| Edge | Roughly how much | What sits there |
|---|---|---|
| Bottom | The lower quarter of the frame | Caption, account name, audio strip |
| Right | A strip down the right edge, about a tenth of the width | Like, comment, share, profile |
| Top | A shallow band at the top | Search and the feed tabs |
| What is left | The middle and upper middle of the frame | Your subject goes here |
The practical effect on prompting is that you should ask for the subject centred and slightly high, and ask for empty space at the bottom. Models compose to the centre by default, which puts your subject exactly where the caption goes.
Vertical examples
Three clips generated at 9:16 rather than cropped into it. The prompt under each one includes the framing instruction, which is the part most people leave out.
Generating vertical against cropping wide
| Generated at 9:16 | Generated at 16:9 and cropped | |
|---|---|---|
| Pixels kept | All of them | About a third of the frame |
| Composition | The model composes for the tall frame | The model composed for a wide frame you then destroyed |
| Subject framing | Fits | Often cut in half, or off centre after the crop |
| Effective resolution | Full | Lower, because you are enlarging a crop |
| When it is fine | Always | Wide landscapes and abstract texture, where the crop takes nothing important |
The first frame
Vertical feeds are scrolled fast, so your first frame is competing before your first second has played. Two prompt habits fix most of it.
- Describe the action as already happening. Write the hand slams the phone down rather than a hand reaches for a phone.
- Ask for the subject in frame at the start. A subject that walks in at second two means two dead seconds.
Sound
Silent vertical video reads as broken to a viewer whose phone is not muted. Either generate with a model that emits native audio, which is covered on the generated sound page, or lay your own audio over a silent generation. Both work. Neither is optional.
Mistakes that are specific to this format
- Text burned into the middle of the frame where the caption will land on top of it.
- One long generated clip instead of several cuts. Vertical rewards cutting.
- A watermark from whichever generator made the clip, left in the corner.
- Frame rates that do not match across clips, which shows up as a stutter at every cut.
- Landscape B-roll pillarboxed into a vertical timeline, with black bars top and bottom.
Related work
Animating a still photo into vertical is a slightly different job and it is covered on the image to video page. The general capabilities and limits of these models are on the video AI generator overview. This page covers organic posting only.
A thirty second structure that holds
Generated clips are short, so a vertical post is a sequence rather than a shot. This is a structure that works with four to eight generations and does not depend on a face.
| Time | What is on screen | Job |
|---|---|---|
| 0 to 2s | The most striking shot you generated, already in motion | Stop the scroll. This is the only job of the first two seconds. |
| 2 to 5s | A shot that states the subject plainly | Tell them what this is before they decide to leave. |
| 5 to 20s | Three or four clips carrying the middle | Deliver the one idea. Cut every three to five seconds. |
| 20 to 27s | The best remaining shot | Pay off the opening. Ideally it rhymes with the first clip. |
| 27 to 30s | A held frame with the closing line | Give them somewhere to land. No hard cut to black. |
Note where the good shots go. First and near last. The middle carries information and does not need to be beautiful, which is useful because most of your takes will not be.
Captions
Burn them in rather than relying on the platform, because auto captions sit where the platform wants them and that is often on top of your subject. Burned in captions belong in the upper middle of the frame, not the lower third, which is a habit carried over from landscape video and puts your words exactly where the platform caption goes.
- Large type. It is being read on a phone at arm's length while moving.
- Two or three words a line, four lines maximum on screen at once.
- High contrast with a subtle shadow or box. Generated backgrounds change brightness mid clip.
- Never over the last quarter of the frame height.
Before you post
Watch it on a phone
Every framing mistake is invisible on a monitor. Send it to yourself and watch it at the size people will see it.
Watch it with the sound on and then off
It needs to work with sound, and it needs to not be confusing without it. Those are different checks.
Check for another tool's watermark
Corner marks and intro cards from whichever generator made the clip. They are easy to miss on a three second insert and they undo the whole post.
Confirm the frame rates match
Clips from different models come back at different rates. A single mismatched clip reads as a glitch at its cut.
Use the AI content label
The platform provides one. Using it costs nothing and not using it is the kind of problem that affects the whole account rather than the post.
Openings that work with generated footage
Hook advice is normally written for people holding a camera. These are the ones that survive when your first shot is a four second generation.
- Motion into stillness. Something moving fast that settles by second two. The eye follows movement and then has somewhere to rest.
- An unexpected scale. Macro on something normally seen at arm's length, or a wide on something normally seen close. Generated footage does both cheaply.
- A camera move that reveals. A tilt up or a pan that finishes on the subject. It buys you two seconds of attention for free.
- A held frame with one thing wrong. Static shot, one element behaving oddly. Cheap to generate and hard to scroll past.
- Direct statement over a simple shot. If your first line is genuinely interesting, the picture only has to not get in the way.
What does not work is the establishing shot. A slow wide of a landscape is a landscape video habit and in a vertical feed it is two seconds of nothing at the exact moment you cannot afford it.
The same video on other vertical feeds
The spec on this page is close enough to work everywhere vertical, with two adjustments worth making rather than exporting once and posting three times.
| Platform | What to change |
|---|---|
| Reels | Safe areas differ slightly. Keep captions higher than you would for TikTok. |
| Shorts | Longer attention window in practice. A slightly slower cut rhythm survives here where it would not elsewhere. |
| All of them | Remove any platform watermark before reposting. Reposted files carrying another app's mark get treated as reposts. |
What does not work in this format
- One long generated shot. Even a good ten second clip feels static in a feed built on cuts.
- Generated faces talking. Lip sync problems are magnified on a phone held close. Use a voiceover over footage instead.
- Text generated inside the clip. It will not be readable. Add every word in the edit.
- Slow openings. An establishing shot is a landscape habit. Vertical starts on the subject.
- Landscape footage in a vertical timeline. Black bars top and bottom read as a reposted video and get treated like one.
Vertical video questions
What size should an AI generated TikTok video be?
1080 by 1920 pixels, which is 9:16. Generate at that ratio rather than generating wide and cropping, because cropping a 16:9 generation throws away most of the frame and usually cuts the subject in half.
Can AI video generators output vertical directly?
Most current models do, and you get a better result asking for 9:16 in the generation than reframing afterwards. Say vertical 9:16 in the prompt as well as setting the ratio, because the composition changes when the model knows the shape.
How long should a generated TikTok video be?
Long enough to say one thing, which for most posts is twenty to thirty five seconds. Generated clips run four to ten seconds each, so that is four to eight clips cut together rather than one long generation.
Where are the TikTok safe areas?
Roughly the bottom quarter of the frame and a strip down the right edge, where the caption and the buttons sit. TikTok does not publish exact figures and they move between app versions, so leave more room than you think you need.
Does the first frame matter that much?
Yes. The first frame is what people see while deciding whether to keep watching, and a static opening frame reads as a still image. Prompt for motion that has already started rather than motion that begins.
Should generated TikTok videos have sound?
Yes. The platform plays with sound on and silent posts feel broken. Either use a model that generates native audio or add a voice, a music bed, or ambience yourself before posting.
Will TikTok penalise AI generated video?
The platform expects AI generated content to be labelled and provides a toggle for it. Labelling is the requirement. Watermarks from other apps are a separate issue and are worth removing before posting.