How to make faceless videos, end to end

A faceless video is four separate jobs pretending to be one: pictures, motion, a voice, and captions. Each of those runs in a different part of this platform, out of a different wallet, on a different plan gate — which is the part every other guide leaves out. This page is the chain, in order, with the gate named at every link so you know before you start where your account stops.

The chain, and what each link costs

There is no single button for this and pretending otherwise would waste your time. Four steps, four different surfaces. The useful thing to know up front is that the first two are cheap and reachable on a free account, and the second two are not.

The four steps, their surface and their gate
StepWherePlanWallet
Stills for each beat/imageFree plan and up⚡ green on our own GPUs
Motion on each still/video, GenerateFree plan and up⚡ green on our own GPUs
Voice-over/audio, SpeechStarter and up✦ gold
Burnt-in captions/audio, CaptionsStarter and upno credits, but the tier gate applies

Which plan: The Audio Studio is Starter and up and the gate is enforced on the server, not just in the interface — every generating audio route refuses a free account outright. Every audio engine is provider-billed, so green credits cannot reach the studio either.

Step one: the stills

Write the script first, in whatever you like, and break it into beats — one picture per beat, roughly one every five to eight seconds of finished video. A three-minute video is about twenty-five stills.

Generate them on the own-GPU image lane, where images start at ⚡1 each and a free account's wallet refills to ⚡20 a day. Keep one visual rule across the whole set — the same lens, the same palette, the same light — because consistency is what makes a faceless video read as one piece rather than a slideshow. The free image generator guide covers the models and the limits.

Step two: motion on every still

Take each still into the video composer and animate it. LTX 2.3 Image + Text → Video is the lane: one start frame, one sentence of motion, 4, 8 or 16 seconds, paid with the same green wallet. On a free account the render is clamped to a 480 px short side; any paid plan removes the clamp.

Prompt the camera, not the picture. The still already carries the subject. "Slow push in" and "static camera, only the smoke moves" are the two moves that carry most faceless work, and the second one is cheaper to get right.

What it costs: At the free plan's clamped size an LTX clip starts at about ⚡4.5 for four seconds and ⚡18 for sixteen, against a ⚡20 day. Own-GPU prices are set at runtime and can carry a demand multiplier when the queue is deep, so treat those as floors. Twenty-five clips is a paid plan's job, not a free day's.

Step three: the voice-over

This is where the chain leaves the free plan. Open the Audio Studio and use the Speech tab: paste the script, choose a preset voice or one you have cloned from your own recording, and render. Speech is billed per thousand characters in gold — from ✦3.4 on the fast engine, ✦6.7 on the high-definition and expressive ones, and ✦1.7 on a voice you cloned yourself. One request takes up to 30,000 characters, which is a long script by any measure, and billing is in whole thousand-character blocks.

Cloning your own voice first is worth it if you are making more than a handful of these: the cloned rate is half the preset rate, and the sample is a single upload under 20 MB. The voice cloning guide covers it.

Step four: captions

Faceless video is watched muted, so the captions are not a nicety. The Captions tab burns subtitle lines into a clip, up to 500 lines, and costs no credits — but it lives inside the Audio Studio, so the Starter gate applies to it like everything else there.

Assembling the thing

  1. Lay the script against the stills. Count the seconds of voice-over each beat needs, and generate the clip for that beat at the nearest length the model offers. Matching lengths at generation time saves the trimming later.
  2. Render the voice-over in one pass. One request up to 30,000 characters keeps the pacing and the tone consistent. Splitting the script into twenty requests does not.
  3. Burn the captions per clip. Captions go onto the clip they belong to, before assembly, so a re-order later does not desynchronise anything.
  4. Join the clips. The clips are ordinary assets and can be assembled in whatever editor you already use. The in-app Timeline editor does this too, but it is an Ultra-only section — do not plan the workflow around it unless you are on Ultra.
  5. Re-cut for shorts if you want them. Once the long video exists, Analyze can index it and pull vertical shorts out of it, in gold, from Starter up.

Start with the video composer — Opens the composer for step two. Nothing runs until you press Generate.

Watch out: Output on the Free and Starter plans carries an OpenModels watermark on every asset. The switch that turns it off arrives at Creator, and it is re-read every time an asset is saved rather than at generation time.

What a three-minute faceless video actually costs

Worth doing before you start, because the two wallets run out at different rates and only one of them refills. Take a three-minute video at roughly one beat every seven seconds: about twenty-five stills, twenty-five short clips, and around 3,000 characters of script.

A three-minute faceless video, by step
StepQuantityWalletRough cost
Stills25 images⚡ greenfrom ⚡25
Clips25 renders of 4 s⚡ greenfrom about ⚡113 at 480p
Voice-over3,000 characters✦ goldabout ✦5 on a cloned voice, ✦10 on the fast preset engine
Captionsone pass per clipnone0 credits, Starter gate applies

So the green side of the bill is around ⚡140 — a Creator day, seven free days, or a couple of Starter days — and the gold side is around ✦5 on a cloned voice, or ✦10 to ✦20 on a preset one. Own-GPU prices are set at runtime and can carry a demand multiplier when the render queue is deep, so read those green numbers as floors and the button as the truth.

The two useful levers are both on the clip step, because it dominates. Shorter clips cost less per beat and hold together better. And a still that needs no motion at all — a title card, a chart — can simply be a still.

What not to bother with

What does the whole chain need, at minimum?

Starter. The picture and motion steps run on a free account, but the voice-over and captions are Audio Studio features and the studio refuses non-paid tiers at the server, not just in the interface.

Can I do it entirely on the free plan?

You can make the visuals — stills and animated clips, at a 480 px short side — but not the voice-over or the burnt-in captions. Every audio engine is provider-billed and the green wallet cannot pay for any of them.

How many credits is a three-minute faceless video?

Roughly twenty-five stills at ⚡1 and twenty-five short clips at about ⚡4.5 each on the own-GPU lane, so on the order of ⚡140 of green — a Creator day. The voice-over adds gold: about ✦5 for 3,000 characters on a cloned voice, or ✦10 on the fast preset engine.

Do I need the Timeline editor?

No, and you should not plan on it: Timeline is an Ultra-only section and is not visible below that plan. Clips come out as ordinary video assets that any editor can join.

Which model should the clips use?

LTX 2.3 Image + Text → Video for anything on the green wallet — it is the only own-GPU lane that takes a literal start frame and reaches sixteen seconds. Kling 3.0 Pro or Veo 3.1 if you are paying gold for 1080p with sound.