Image to video

First frame, last frame: how an AI commercial starts as stills

For the brand manager or agency producer approving an AI commercial, who is shown stills before any motion and wants to know what to look for in them.

The short answer

In image-to-video, the model starts a shot from a still you give it, the first frame, and if you give it a second still, the last frame, it generates the motion between the two. Google's Veo 3.1 and Gemini Omni Flash, Kling 3.0, Luma and MiniMax all offer it. It lets a brand approve the cast, the product and the composition on a still that costs cents, before paying for motion at Google's published $0.05 to $0.60 a second of video, and it is the most direct control a studio has over where a shot begins and ends.

The modes

Four ways to start an AI video shot

ModeWhat you give the modelWhat you controlGood for
Text to videoA promptThe idea; the model invents the frameExploring ideas, backgrounds, crowds
Reference imagesA prompt plus pictures of the person or productWho and what appears; the model frames itKeeping a face or a product across new scenes
First frameA prompt plus the opening stillThe composition, cast, product and light of the first frameMost commercial shots
First and last frameA prompt plus the opening and closing stillsWhere the shot starts and where it landsReveals, transitions, landing on a pack shot
How much of the frame you decide in each mode. The more you decide before generation, the fewer takes you throw away.

A commercial usually mixes all four. A studio might generate a mood shot of weather from text alone, then build every shot with a face or the product from an approved first frame, so the parts of the ad that carry the brand start from a picture someone signed off.

The models

What the main video models support

ModelFirst and last frameClip lengthPublished price
Veo 3.1 (Google)A starting image, plus a lastFrame image for interpolation4, 6 or 8 seconds at 24 fps; 8 with reference images, 1080p or 4K$0.40 a second with audio at 720p or 1080p; Fast $0.10 to $0.30; Lite $0.05 to $0.08
Gemini Omni Flash (Google)Images tagged FIRST_FRAME and LAST_FRAMEExtensions of 3 to 10 seconds, up to 40 seconds in totalAbout $0.10 a second at 720p
Kling 3.0Start & End Frames-to-Video3 to 15 seconds, 720p or 1080pApp credits: 6 to 12 a second; no dollar price on the pages checked
Runway Gen-4.5A first frame only, in the API2 to 10 seconds$0.12 a second in the API
Luma Ray3.2Start and end frames, and several keyframes5 or 10 seconds; 10-second clips take no start or end frame$1.20 for a 5-second clip at 1080p
MiniMax H3First-frame and last-frame roles4 to 15 seconds, 768P or 2K$0.08 a second at 768P; $0.13 at 2K
From each vendor's documentation and pricing page, checked September 2026: Veo 3.1, Gemini Omni Flash, Gemini API pricing, Kling 3.0, Runway API pricing, Luma API pricing and MiniMax video generation. Prices in USD.

Google now names Gemini Omni Flash its default video model and keeps Veo 3.1 for scene extension and last-frame control. Both put an invisible SynthID watermark in every video they make. Veo 3.1 allows only adults in image-to-video, interpolation and reference-image shots, which matters for any campaign with children in it.

Kling's guide to start and end frames has the most useful rule for commercials: pick two similar images, because large differences between the first and last frame can make the model cut to a new shot instead of moving smoothly. It also notes that many users make similar images first, with an image model, and then use the feature.

Read the terms as closely as the features. OpenAI shut down its Sora 2 models and Videos API on September 24, 2026. Kling's user terms, effective April 21, 2026, say outputs may not be used for any commercial purpose without Kling's written permission, so get that permission, or a plan whose terms grant it, before Kling footage goes into an ad.

The reason

Why AI commercials start from approved stills

Money. At Google's published prices, a Nano Banana 2 still costs $0.067 at 1K and $0.101 at 2K. An 8-second 1080p shot from Veo 3.1 with audio costs $3.20. Most takes get thrown away: the director of the Kalshi ad that aired during the 2025 NBA Finals reported 300 to 400 generations to get 15 usable clips (Mashable, June 2025). A rejected still costs cents, and a rejected shot costs dollars.

Approval. A frame is quick to approve: the product, the label, the cast, the wardrobe and the light are all visible and still. A legal or MLR reviewer can check a claim on a frame without scrubbing through a clip.

Control. Google's own Veo documentation shows this workflow: make an image with Nano Banana 2, then use it as the starting frame for Veo 3.1. Anything the first frame leaves open, from the background to the wardrobe, is left to the model.

The checklist

What makes a clean first frame

  1. Made at the delivery ratio. Veo 3.1 generates 16:9 or 9:16. Make a separate first frame for each ratio you deliver; a 16:9 frame cropped to 9:16 usually cuts the product or a face.
  2. At or above the output resolution. Google's Omni guide asks for high-resolution images. A 1080p shot needs a first frame of at least 1920 by 1080 pixels.
  3. Room to move. Leave space on the side the subject moves toward. A push-in needs detail at the center of the frame; a pull-out needs a world around the edges.
  4. The product right in frame one. The model carries what it sees, so check the label, the logo and the shape against the real product before the still is approved. See product and label accuracy in AI commercials.
  5. No type in the frame. Supers, prices, claims and logos go on in the edit, from the brand's real files. A headline painted into a first frame is redrawn by the model in every frame after it, where it can warp. Overs' guide to getting text right in AI images makes the same case for stills: render clean, set the words in a second pass.
  6. Light that matches the next shot. If the sun is at frame left in shot 4, it is at frame left in the first frame of shot 5.
  7. The moment before the action. One action per shot. The first frame shows the runner about to tie the lace, not halfway through it.
  8. A last frame close to the first. For interpolation, the end frame should share the place, the light and the lens with the start frame, so the model moves between them instead of cutting.

The prompt

What to write with the frame

The still fixes what the shot looks like. The motion prompt says what happens: the camera move, the action, the timing, what stays still and what the shot sounds like. Google's Omni guide puts it plainly: vague prompts like “make it move” produce less compelling results than detailed descriptions of the camera movement, the subject's motion and the surroundings.

An example motion prompt, written for this page for a made-up shoe ad:

Slow push-in from a medium shot to a close-up over six seconds. The runner ties the lace on the left shoe, then looks up toward frame right. Leaves move in a light wind. The shoe and its logo stay sharp and do not change. Forest ambience, footsteps on gravel, no music.

Every line of it can be checked against the result, and the line about the shoe names what must not change.

The stills

Where Overs comes in

The first frames of an AI commercial are stills, and often the campaign stills themselves, which is where Overs comes in. It is the sister product from the same founder: an AI tool that makes product and campaign photos that match a brand, without a photo shoot, built on the process 100creatives used to make hundreds of photos a week for hundreds of brands.

Two of its features go straight at the parts of a first frame that break most often, the person and the product. For a product that is worn or held, Overs makes a character sheet first, one model from several angles wearing the exact product, and every later photo with a person uses that same model. Each photo gets only the reference pictures it needs, in order, and if a needed reference does not exist, Overs skips it instead of guessing with the wrong one. You see the shot plan and an estimated cost before anything renders, and you approve, change or reject every photo. For each photo, it can write a motion prompt for Seedance or Veo, so each still comes with a written starting point for the motion. Overs makes no video itself; the motion happens in the video model.

Stills are the cheap end of the job. In 4 campaigns 100creatives made in 2026, a usable photo took about 8 renders, which comes to $0.27 to $1.69 in AI fees per usable photo at published model prices. The free plan makes 40 photos a month, with the AI billed to your own OpenRouter key at a few cents a photo and no markup from Overs. Try Overs free at www.overs.studio and approve your first frames before a single second of video is paid for.

Should Ruminate X make your commercial?

Hire Ruminate X if

  • A brand that needs a full commercial where every shot with the product or a face is planned, boarded and checked.
  • A marketing team with approved campaign stills that wants a film made from the same world.
  • A product that has to stay exact across many shots, where each take has to be checked against the real thing.

Hire someone else if

  • You need one five-second loop from a product photo: make the still yourself and animate it in Veo or Gemini Omni Flash.
  • The ad needs real footage of your store, team or customers: hire a crew.
  • You need dozens of quick variations a week for testing: a self-serve tool costs less than a studio.

Questions buyers ask

What is first frame last frame in AI video?

It is an image-to-video mode where you give the model two stills: the frame the shot starts on and the frame it ends on. The model generates the motion between them. With only a first frame, the model starts from your still and decides where the shot goes. Google calls the two-frame version interpolation; Kling calls it Start and End Frames.

Which AI video models support first and last frames?

As of September 2026: Google's Veo 3.1 (a starting image plus a lastFrame image), Google's Gemini Omni Flash (images tagged FIRST_FRAME and LAST_FRAME), Kling 3.0 (Start & End Frames-to-Video), Luma Ray3.2 (start and end frames) and MiniMax H3 (first-frame and last-frame roles). Runway's Gen-4.5 API takes a first frame only. Check each vendor's documentation and terms before a project, because limits change with every model version.

Why do AI commercials start from stills?

Because a still is cheap to remake and quick to approve, and a second of video is neither. At Google's published September 2026 prices, a 2K still from Nano Banana 2 costs about $0.10, and an 8-second 1080p shot from Veo 3.1 with audio costs $3.20. Approving the cast, the product and the composition on a still before paying for motion keeps rejected takes cheap.

How long can a first and last frame clip be?

Short. Veo 3.1 makes clips of 4, 6 or 8 seconds; Kling 3.0 and MiniMax H3 go up to 15 seconds, though Kling's guide says start and end frames work best for transitions within 5 seconds. Longer shots are built by extending a clip or by cutting several shots together in the edit, each with its own first frame.

Can I use my campaign photos as first frames?

Often, if the photo is at the aspect ratio and resolution of the video, has no text or logo set into it, and leaves room in the frame for the planned camera move. A 4:5 feed photo cropped to 16:9 usually loses the top of the product or the head of the person, so make a separate first frame for each video ratio. You also need the rights to the photo, and people in it must be cleared for the new use.

How does Ruminate X plan the shots of an AI commercial?

Every shot is boarded before generation, with its framing, movement and duration, and the client sees the beat sheet and shot list before a single frame is generated. Each shot is then generated, checked against its board and generated again until it holds; the pack shot, faces and hands get the most passes. Every frame Ruminate X delivers is made with generative AI, with no crew, set or location shoot.

Send the frames you have

Send your campaign stills, the product and where the ad will run. Every frame Ruminate X delivers is made with generative AI. There are no film crews, sets or location shoots.