How to make a video between two images

Give the model the first frame and the last frame and it renders the journey between them. It is the most controllable video lane there is, because you have fixed both ends of the shot and left the model only the middle to invent. On OpenModels it runs on our own GPUs, which means the free green wallet pays for it.

Why two frames beats one

With a single start frame you are asking the model to invent everything after the first twenty-fifth of a second, and it will — including where the shot ends up. With two frames you have specified the destination, so the model is solving a much smaller problem: how to get from A to B in the time given. The failure modes shrink accordingly.

This is what makes it the right lane for a transition, a product turn, a before-and-after, or any shot where the last frame has to match something you already have — the first frame of the next clip, for instance.

Which plan: Free plan and up. LTX 2.3 First / Last Frame is a comfy lane on our own GPUs, so it is payable with the green wallet a free account refills to ⚡20 a day. Free renders are clamped to a 480 px short side; any paid plan removes the clamp.

The envelope

LTX 2.3 First / Last Frame
PropertyValue
Images requiredexactly two — a start frame and an end frame
Clip lengths4, 8 or 16 seconds
Resolutions480p, 720p, 1080p — clamped to 480 px short side on the free plan
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Wallet⚡ green, from ⚡6 per started five-second block at 720p
Planevery plan, including Free

Both images are required, not optional. Attaching one and pressing Generate is refused in the composer rather than quietly rendering a one-frame job, because a lane that silently drops an input is a lane that bills you for the wrong render.

What it costs: Billing is per started five-second block, quoted at 720p, with a 0.75 multiplier at 480p. So at the free plan's clamped size: about ⚡4.5 for four seconds, ⚡9 for eight, ⚡18 for sixteen. Own-GPU prices are set at runtime and can carry a demand multiplier, so treat those as floors and read the quote on the button.

Making the clip

  1. Prepare two frames that belong to the same shot. Same subject, same lens, same light, same aspect ratio. The model interpolates a camera move and a subject action — it does not interpolate a change of location or a different person.
  2. Open the composer and pick the lane. Go to /video and choose LTX 2.3 First / Last Frame in the model picker. The media row will then ask for two images rather than one.
  3. Attach them in order. First frame, then last frame. The order is the shot's direction; swapping them renders the reverse move, which is occasionally what you want.
  4. Write what happens in between. The two frames already say where it starts and ends. Spend the prompt on the path: "the camera arcs left as she turns her head to follow", not another description of either image.
  5. Choose a length that suits the distance. A small move over sixteen seconds looks like a freeze; a large move over four looks like a whip. Match the length to how far apart the two frames actually are.
  6. Read the quote and generate. The button carries the price and the wallet before you press it. Four seconds is the cheap way to check whether the pair works at all.

Open the composer — Opens the video composer. Pick the First / Last Frame lane and attach both images.

Choosing a pair that works

Almost every bad result from this lane comes from a pair of images that could not belong to the same continuous shot. The model will still try, and what it produces is a morph rather than a move.

Watch out: The lane renders a single continuous shot. It cannot cut, and asking it to "then cut to" produces a smear across the cut rather than an edit. Render two clips and join them.

Where the second image should come from

The hardest part of this lane is not the render, it is producing an end frame that could genuinely be the same shot four seconds later. Prompting a second image from scratch will not do it — two independent generations of the same description are two different people in two different rooms.

Edit the start frame instead. Qwen Image Edit runs on our own GPUs and takes a written instruction against an attached picture, on the same green wallet, from ⚡1: attach the start frame, ask for the one change you want at the other end — a turned head, a raised cup, a wider framing — and use the result as frame two. Everything you did not ask to change stays, which is exactly the property this lane needs. Krea 2 Identity is the other own-GPU option when the change is a restage of the same person or product rather than a small edit.

For a change the edit models cannot make, the fallback is to shoot both frames from one source: generate a batch from the same prompt and seed, pick two that share a subject, and accept that you will spend a few green credits finding the pair. It is still cheaper than a bad sixteen-second render.

Chaining clips into a longer sequence

Because you control both ends, this lane is the one that lets separate renders join up. The method is mechanical:

  1. Render clip one from frame A to frame B.
  2. Take frame B — the actual last frame of that render, not the image you fed in — as the start frame of clip two.
  3. Make frame C with an edit against frame B, and render clip two from B to C.
  4. Repeat. Each render is a self-contained job on the green wallet, and the joins are continuous because each new clip literally begins on the previous one's last frame.

Four chained 16-second LTX renders is just over a minute of continuous video, built four blocks at a time. On a free account that is four days of the daily wallet, which is a real constraint rather than a rhetorical one — but it is a minute of video for no money, and the clamped 480 px size is fine for deciding whether the sequence works before a paid plan renders it properly.

Is first-and-last-frame video free here?

It runs on our own GPUs and is paid with the green wallet, which a free account refills to ⚡20 a day. At the free plan's clamped 480 px size a four-second render starts at about ⚡4.5, so roughly four a day.

How long can the clip be?

4, 8 or 16 seconds. Billing is per started five-second block, so those three options are one, two and four blocks.

Do the two images have to be the same size?

They should share an aspect ratio. The lane offers 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, and a mismatched pair forces a reframe that shows.

Can I do this with a hosted model instead?

Yes, in gold, from Starter up. Kling 3.0 Pro and Seedance 2.0 Mini both take an optional end frame alongside a start frame. They cost per second rather than per block and render at higher resolutions.

Why does the middle of my clip look like a morph?

Because the two frames could not belong to one continuous shot — a different subject, a different light, or a different framing. Generate the end frame by editing the start frame rather than prompting it separately.