videoaigenerator.ai

Vertical format

AI video generator for TikTok: the actual spec

Short answer

Every page about using an AI video generator for TikTok says viral and stops there. The spec is simple and nobody prints it: 1080 by 1920, motion already happening in the first frame, sound on, the bottom quarter and right edge kept clear, and four to eight generated clips cut into twenty to thirty five seconds.

Vertical is not landscape turned sideways. The composition rules are different, the attention curve is different, and a chunk of the frame belongs to the app rather than to you. Generated footage makes all of that harder, because the model composes for the frame you asked for and knows nothing about the buttons that will sit on top of it. This page is the spec sheet, then the prompt changes that follow from it.

The spec

  • Frame1080 by 1920, 9:16. Generate at this ratio, do not crop into it.
  • Frame rate30fps is the safe default. Match every clip before you cut them together.
  • First frameMotion already in progress. Never a static establishing shot.
  • SoundAlways. Native audio, a voice track, or at minimum an ambience bed.
  • LengthTwenty to thirty five seconds for one idea. Four to eight generated clips.
  • Safe areasKeep the bottom quarter and the right edge free of anything that matters.
  • Subject placementUpper middle third. That is the part of the frame nothing covers.
  • ExportH.264 MP4, no watermark from whatever tool made it.

Safe areas, honestly

TikTok does not publish exact safe area figures and the layout shifts between app versions and device sizes. What follows is the working approximation, and the right approach is to leave more room than the numbers suggest.

Approximate areas of a 1080 by 1920 frame that the interface covers. Treat as guidance, not as published values.
EdgeRoughly how muchWhat sits there
BottomThe lower quarter of the frameCaption, account name, audio strip
RightA strip down the right edge, about a tenth of the widthLike, comment, share, profile
TopA shallow band at the topSearch and the feed tabs
What is leftThe middle and upper middle of the frameYour subject goes here

The practical effect on prompting is that you should ask for the subject centred and slightly high, and ask for empty space at the bottom. Models compose to the centre by default, which puts your subject exactly where the caption goes.

Vertical examples

Three clips generated at 9:16 rather than cropped into it. The prompt under each one includes the framing instruction, which is the part most people leave out.

Placeholder frame. This clip has not been generated and uploaded yet. Prompt: Vertical 9:16. A hand slams a phone face down on a desk, camera snaps in fast, harsh top light, first frame already mid action.
First frame already in motionVertical 9:16. A hand slams a phone face down on a desk, camera snaps in fast, harsh top light, first frame already mid action.Seedance 2.09:16clip not uploaded yet
Placeholder frame. This clip has not been generated and uploaded yet. Prompt: Vertical 9:16. Point of view walking into a bright studio, handheld, subject enters frame at second one, safe margins kept clear.
Subject enters earlyVertical 9:16. Point of view walking into a bright studio, handheld, subject enters frame at second one, safe margins kept clear.Veo 39:16clip not uploaded yet
Placeholder frame. This clip has not been generated and uploaded yet. Prompt: Vertical 9:16. A tin of matcha rotating on a turntable, centre framed, top and bottom thirds left empty for captions.
Centred with margins kept clearVertical 9:16. A tin of matcha rotating on a turntable, centre framed, top and bottom thirds left empty for captions.Kling 2.59:16clip not uploaded yet

Generating vertical against cropping wide

Why the ratio belongs in the generation rather than the edit.
Generated at 9:16Generated at 16:9 and cropped
Pixels keptAll of themAbout a third of the frame
CompositionThe model composes for the tall frameThe model composed for a wide frame you then destroyed
Subject framingFitsOften cut in half, or off centre after the crop
Effective resolutionFullLower, because you are enlarging a crop
When it is fineAlwaysWide landscapes and abstract texture, where the crop takes nothing important

The first frame

Vertical feeds are scrolled fast, so your first frame is competing before your first second has played. Two prompt habits fix most of it.

  • Describe the action as already happening. Write the hand slams the phone down rather than a hand reaches for a phone.
  • Ask for the subject in frame at the start. A subject that walks in at second two means two dead seconds.

Sound

Silent vertical video reads as broken to a viewer whose phone is not muted. Either generate with a model that emits native audio, which is covered on the generated sound page, or lay your own audio over a silent generation. Both work. Neither is optional.

Mistakes that are specific to this format

  • Text burned into the middle of the frame where the caption will land on top of it.
  • One long generated clip instead of several cuts. Vertical rewards cutting.
  • A watermark from whichever generator made the clip, left in the corner.
  • Frame rates that do not match across clips, which shows up as a stutter at every cut.
  • Landscape B-roll pillarboxed into a vertical timeline, with black bars top and bottom.

Related work

Animating a still photo into vertical is a slightly different job and it is covered on the image to video page. The general capabilities and limits of these models are on the video AI generator overview. This page covers organic posting only.

A thirty second structure that holds

Generated clips are short, so a vertical post is a sequence rather than a shot. This is a structure that works with four to eight generations and does not depend on a face.

One idea, thirty seconds, six generated clips.
TimeWhat is on screenJob
0 to 2sThe most striking shot you generated, already in motionStop the scroll. This is the only job of the first two seconds.
2 to 5sA shot that states the subject plainlyTell them what this is before they decide to leave.
5 to 20sThree or four clips carrying the middleDeliver the one idea. Cut every three to five seconds.
20 to 27sThe best remaining shotPay off the opening. Ideally it rhymes with the first clip.
27 to 30sA held frame with the closing lineGive them somewhere to land. No hard cut to black.

Note where the good shots go. First and near last. The middle carries information and does not need to be beautiful, which is useful because most of your takes will not be.

Captions

Burn them in rather than relying on the platform, because auto captions sit where the platform wants them and that is often on top of your subject. Burned in captions belong in the upper middle of the frame, not the lower third, which is a habit carried over from landscape video and puts your words exactly where the platform caption goes.

  • Large type. It is being read on a phone at arm's length while moving.
  • Two or three words a line, four lines maximum on screen at once.
  • High contrast with a subtle shadow or box. Generated backgrounds change brightness mid clip.
  • Never over the last quarter of the frame height.

Before you post

  1. Watch it on a phone

    Every framing mistake is invisible on a monitor. Send it to yourself and watch it at the size people will see it.

  2. Watch it with the sound on and then off

    It needs to work with sound, and it needs to not be confusing without it. Those are different checks.

  3. Check for another tool's watermark

    Corner marks and intro cards from whichever generator made the clip. They are easy to miss on a three second insert and they undo the whole post.

  4. Confirm the frame rates match

    Clips from different models come back at different rates. A single mismatched clip reads as a glitch at its cut.

  5. Use the AI content label

    The platform provides one. Using it costs nothing and not using it is the kind of problem that affects the whole account rather than the post.

Openings that work with generated footage

Hook advice is normally written for people holding a camera. These are the ones that survive when your first shot is a four second generation.

  • Motion into stillness. Something moving fast that settles by second two. The eye follows movement and then has somewhere to rest.
  • An unexpected scale. Macro on something normally seen at arm's length, or a wide on something normally seen close. Generated footage does both cheaply.
  • A camera move that reveals. A tilt up or a pan that finishes on the subject. It buys you two seconds of attention for free.
  • A held frame with one thing wrong. Static shot, one element behaving oddly. Cheap to generate and hard to scroll past.
  • Direct statement over a simple shot. If your first line is genuinely interesting, the picture only has to not get in the way.

What does not work is the establishing shot. A slow wide of a landscape is a landscape video habit and in a vertical feed it is two seconds of nothing at the exact moment you cannot afford it.

The same video on other vertical feeds

The spec on this page is close enough to work everywhere vertical, with two adjustments worth making rather than exporting once and posting three times.

What changes between vertical platforms.
PlatformWhat to change
ReelsSafe areas differ slightly. Keep captions higher than you would for TikTok.
ShortsLonger attention window in practice. A slightly slower cut rhythm survives here where it would not elsewhere.
All of themRemove any platform watermark before reposting. Reposted files carrying another app's mark get treated as reposts.

What does not work in this format

  • One long generated shot. Even a good ten second clip feels static in a feed built on cuts.
  • Generated faces talking. Lip sync problems are magnified on a phone held close. Use a voiceover over footage instead.
  • Text generated inside the clip. It will not be readable. Add every word in the edit.
  • Slow openings. An establishing shot is a landscape habit. Vertical starts on the subject.
  • Landscape footage in a vertical timeline. Black bars top and bottom read as a reposted video and get treated like one.

Vertical video questions

What size should an AI generated TikTok video be?

1080 by 1920 pixels, which is 9:16. Generate at that ratio rather than generating wide and cropping, because cropping a 16:9 generation throws away most of the frame and usually cuts the subject in half.

Can AI video generators output vertical directly?

Most current models do, and you get a better result asking for 9:16 in the generation than reframing afterwards. Say vertical 9:16 in the prompt as well as setting the ratio, because the composition changes when the model knows the shape.

How long should a generated TikTok video be?

Long enough to say one thing, which for most posts is twenty to thirty five seconds. Generated clips run four to ten seconds each, so that is four to eight clips cut together rather than one long generation.

Where are the TikTok safe areas?

Roughly the bottom quarter of the frame and a strip down the right edge, where the caption and the buttons sit. TikTok does not publish exact figures and they move between app versions, so leave more room than you think you need.

Does the first frame matter that much?

Yes. The first frame is what people see while deciding whether to keep watching, and a static opening frame reads as a still image. Prompt for motion that has already started rather than motion that begins.

Should generated TikTok videos have sound?

Yes. The platform plays with sound on and silent posts feel broken. Either use a model that generates native audio or add a voice, a music bed, or ambience yourself before posting.

Will TikTok penalise AI generated video?

The platform expects AI generated content to be labelled and provides a toggle for it. Labelling is the requirement. Watermarks from other apps are a separate issue and are worth removing before posting.