Two different jobs share one tab. Hand the studio a still photo and an audio track and you get a talking portrait that was never filmed. Hand it a video that already has a face in it and the face is re-synced to new audio while the rest of the frame is left alone. The models differ for each, the prices differ, and so does the minimum billable length — Sync Lipsync-2 Pro starts at one second where everything else starts at five.
Photo mode takes a single still and animates the mouth, the jaw and a little of the head to match speech. It invents a performance that was never filmed, which is the harder problem and the reason it costs more per second.
Video mode takes footage that already contains a talking face and re-times the mouth to different audio. Everything else in the frame is left as shot. This is what you want for re-voicing an existing clip, fixing a flubbed line, or putting a swapped voice track back over a performance.
Which plan: The Audio Studio opens on the Starter plan. That gate lives on the server: every generating audio endpoint answers 403 to a free account, so it is not a greyed-out button you can work around. Everything behind these tabs runs on a paid provider, which means it is billed in ✦ gold credits — ⚡ green credits, the ones that refill daily, cannot pay for any of it.
| Model | Input | Credits a second | Billable floor | Ceiling |
|---|---|---|---|---|
| InfiniteTalk | A photo | ✦1.1 | 5 seconds | 600 seconds |
| InfiniteTalk Fast | A photo | ✦0.55 | 5 seconds | 600 seconds |
| LatentSync | A video | ✦0.35 | 5 seconds | 600 seconds |
| InfiniteTalk V2V | A video | ✦1.1 | 5 seconds | 600 seconds |
| Sync Lipsync-2 Pro | A video | ✦1.7 | 1 second | 120 seconds |
What it costs: Sync Lipsync-2 Pro's minimum is one second, not five. On short work that matters more than the per-second rate: a two-second line is ✦3.4 on Lipsync-2 Pro against ✦1.8 on LatentSync's five-second floor, and by about five seconds the comparison has flipped. Its ceiling is also lower — 120 seconds, where the others run to 600.
The three video lanes do genuinely different things to a clip. LatentSync re-syncs the mouth and leaves everything else untouched, which is the least invasive and by a distance the cheapest. InfiniteTalk V2V regenerates the performance, so the head and the expression move too. Lipsync-2 Pro is the studio-grade option and prices like one.
Three doors, and the third is the one people miss. You can upload an audio file. You can record straight from the browser microphone, which is transcoded to mp3 server-side so the format never becomes your problem. Or you can skip having audio at all: type the line, pick a preset or cloned voice, and the speech is generated first and fed into the lipsync model in the same request.
That third route bills twice, and the quote says so — once for the speech at ✦3.4 per started 1,000 characters, or ✦1.7 with a cloned voice, and once for the lipsync at its per-second rate. Typed lines are capped at 2,000 characters, which is roughly two minutes of speech and comfortably inside every model's ceiling except Lipsync-2 Pro's.
There is a second way to make a portrait talk, and it is not in the Audio Studio. Wan InfiniteTalk on /video takes an image and an audio track and runs on our own GPUs. It opens on the Creator plan rather than Starter, and below the Ultra plan it is billed in ✦ gold anyway — the Wan family is one of the few own-GPU lanes green credits cannot pay for. On Starter, the Audio Studio's Lipsync tab is your lane.
More lipsync results are spoiled by the input than by the choice of model. What every one of these engines needs is an unambiguous mouth and a clean voice:
Open the Lipsync tab — Opens the studio on Lipsync. Uploads and the model choice happen there; nothing is charged until you generate.
Either. Photo mode animates a still with InfiniteTalk or InfiniteTalk Fast. Video mode re-syncs a face already in footage, using LatentSync, InfiniteTalk V2V or Sync Lipsync-2 Pro.
Per second of speech audio: ✦0.35 on LatentSync, ✦0.55 on InfiniteTalk Fast, ✦1.1 on InfiniteTalk and InfiniteTalk V2V, ✦1.7 on Sync Lipsync-2 Pro. All in ✦ gold credits.
On four of the five models, yes — a shorter clip is still billed as five seconds. Sync Lipsync-2 Pro is the exception: its floor is one second, though its ceiling is 120 seconds rather than 600.
No. You can type up to 2,000 characters and pick a preset or cloned voice; the speech is generated and fed straight into the lipsync model. That adds the speech cost to the same quote.
Starter or above. The Audio Studio is gated at the API, so a free account is refused by the server rather than by the interface, and every lipsync model bills ✦ gold credits.