How to make a face match an audio track

Two different jobs share one tab. Hand the studio a still photo and an audio track and you get a talking portrait that was never filmed. Hand it a video that already has a face in it and the face is re-synced to new audio while the rest of the frame is left alone. The models differ for each, the prices differ, and so does the minimum billable length — Sync Lipsync-2 Pro starts at one second where everything else starts at five.

Two modes: a photo, or a clip you already have

Photo mode takes a single still and animates the mouth, the jaw and a little of the head to match speech. It invents a performance that was never filmed, which is the harder problem and the reason it costs more per second.

Video mode takes footage that already contains a talking face and re-times the mouth to different audio. Everything else in the frame is left as shot. This is what you want for re-voicing an existing clip, fixing a flubbed line, or putting a swapped voice track back over a performance.

Which plan: The Audio Studio opens on the Starter plan. That gate lives on the server: every generating audio endpoint answers 403 to a free account, so it is not a greyed-out button you can work around. Everything behind these tabs runs on a paid provider, which means it is billed in ✦ gold credits — ⚡ green credits, the ones that refill daily, cannot pay for any of it.

The five models and what they cost

Prices are per second of the speech audio, floored and capped where the provider floors and caps.
ModelInputCredits a secondBillable floorCeiling
InfiniteTalkA photo✦1.15 seconds600 seconds
InfiniteTalk FastA photo✦0.555 seconds600 seconds
LatentSyncA video✦0.355 seconds600 seconds
InfiniteTalk V2VA video✦1.15 seconds600 seconds
Sync Lipsync-2 ProA video✦1.71 second120 seconds

What it costs: Sync Lipsync-2 Pro's minimum is one second, not five. On short work that matters more than the per-second rate: a two-second line is ✦3.4 on Lipsync-2 Pro against ✦1.8 on LatentSync's five-second floor, and by about five seconds the comparison has flipped. Its ceiling is also lower — 120 seconds, where the others run to 600.

The three video lanes do genuinely different things to a clip. LatentSync re-syncs the mouth and leaves everything else untouched, which is the least invasive and by a distance the cheapest. InfiniteTalk V2V regenerates the performance, so the head and the expression move too. Lipsync-2 Pro is the studio-grade option and prices like one.

How to run it

  1. Open the Lipsync tab. Go to /audio and pick Lipsync in the rail. The first choice on the screen is photo or video, because it changes which models are offered.
  2. Upload the face. A portrait for photo mode, a clip for video mode. Front-on, well lit, mouth visible and unobstructed. A face at an angle or half in shadow is the most common reason a result looks wrong.
  3. Give it the audio. Upload a track, record one on the spot, or type a line and pick a voice. A typed line is capped at 2,000 characters and is spoken first — on ElevenLabs Flash, or on your cloned voice — then handed to the lipsync model.
  4. Pick the model. For a photo, InfiniteTalk Fast is half the price of InfiniteTalk and worth trying first. For a clip, start with LatentSync: it is the cheapest by a wide margin and often all that is needed.
  5. Check the quote, then generate. The cost is computed from the length of the speech audio and shown before you commit. If you typed the line, the speech generation is added into the same quote.
  6. Collect it from your assets. The job runs asynchronously and the result lands in your assets library. Long audio takes proportionally longer — these models work through a clip rather than in one pass.

Where the audio comes from

Three doors, and the third is the one people miss. You can upload an audio file. You can record straight from the browser microphone, which is transcoded to mp3 server-side so the format never becomes your problem. Or you can skip having audio at all: type the line, pick a preset or cloned voice, and the speech is generated first and fed into the lipsync model in the same request.

That third route bills twice, and the quote says so — once for the speech at ✦3.4 per started 1,000 characters, or ✦1.7 with a cloned voice, and once for the lipsync at its per-second rate. Typed lines are capped at 2,000 characters, which is roughly two minutes of speech and comfortably inside every model's ceiling except Lipsync-2 Pro's.

The other talking-head lane, on the video page

There is a second way to make a portrait talk, and it is not in the Audio Studio. Wan InfiniteTalk on /video takes an image and an audio track and runs on our own GPUs. It opens on the Creator plan rather than Starter, and below the Ultra plan it is billed in ✦ gold anyway — the Wan family is one of the few own-GPU lanes green credits cannot pay for. On Starter, the Audio Studio's Lipsync tab is your lane.

What makes a source that works

More lipsync results are spoiled by the input than by the choice of model. What every one of these engines needs is an unambiguous mouth and a clean voice:

Where lipsync stops

Open the Lipsync tab — Opens the studio on Lipsync. Uploads and the model choice happen there; nothing is charged until you generate.

Can I lipsync a still photo, or does it need video?

Either. Photo mode animates a still with InfiniteTalk or InfiniteTalk Fast. Video mode re-syncs a face already in footage, using LatentSync, InfiniteTalk V2V or Sync Lipsync-2 Pro.

What does lipsync cost?

Per second of speech audio: ✦0.35 on LatentSync, ✦0.55 on InfiniteTalk Fast, ✦1.1 on InfiniteTalk and InfiniteTalk V2V, ✦1.7 on Sync Lipsync-2 Pro. All in ✦ gold credits.

Is there really a five-second minimum?

On four of the five models, yes — a shorter clip is still billed as five seconds. Sync Lipsync-2 Pro is the exception: its floor is one second, though its ceiling is 120 seconds rather than 600.

Do I need to record the audio first?

No. You can type up to 2,000 characters and pick a preset or cloned voice; the speech is generated and fed straight into the lipsync model. That adds the speech cost to the same quote.

What plan do I need?

Starter or above. The Audio Studio is gated at the API, so a free account is refused by the server rather than by the interface, and every lipsync model bills ✦ gold credits.