Voice Swap takes speech that already exists and replaces the voice performing it. The words stay the same, the timing stays the same, and on a video the picture is never touched — ffmpeg pulls the audio out, the provider re-voices it, and ffmpeg puts it back over the original video stream. That precision is also the limit, and it is the first thing this page says out loud: this is a voice change, not a translation.
Give it a clip of somebody talking and pick a different voice. What comes back is the same performance — same words, same pace, same pauses — spoken by somebody else. On a video only the audio track is regenerated: the picture that goes back out is the one you uploaded, because the video stream is copied rather than re-encoded around the new audio.
There is a second door that is easy to miss. You do not need a video at all. An audio file, or a recording made on the spot from the browser microphone, goes straight through and comes back as an mp3 in the new voice. That is the shortest path from "record yourself reading it" to "a narrator reads it", and it skips speech synthesis entirely.
Which plan: The Audio Studio opens on the Starter plan. That gate lives on the server: every generating audio endpoint answers 403 to a free account, so it is not a greyed-out button you can work around. Everything behind these tabs runs on a paid provider, which means it is billed in ✦ gold credits — ⚡ green credits, the ones that refill daily, cannot pay for any of it.
Watch out: Voice Swap does not translate. It re-performs the words that are already there, in a different voice. If you want a clip in another language the words have to change, and that is a different job — the three steps are below.
| Input | Cap | Price | What comes back |
|---|---|---|---|
| Video file | 300 seconds | ✦0.17 a second, 5-second minimum | The same video, new voice on the audio track |
| Audio file | 300 seconds | ✦0.17 a second, 5-second minimum | An mp3 in the new voice |
| Mic recording | 300 seconds | ✦0.17 a second, 5-second minimum | An mp3 in the new voice |
At ✦0.17 a second it is roughly half the price of the cheapest lipsync model, because it only has to rebuild the audio. If the face on screen does not have to match the new voice — a voiceover, a screen recording, a podcast clip — this is the tab, not Lipsync.
This is the request behind most searches for dubbing, so here is the honest answer: there is no one-button translation lane on this platform. What exists is the three steps that make one, and each is a separate charge.
It is more work than one button, and it is what actually works today. Voice Swap stays the right tool when the words do not change — a different narrator, one brand voice across clips recorded by different people, or anonymising a recording.
The two tabs overlap enough to be confusing, and picking the wrong one costs money. The question to ask is whether a mouth has to match.
| The situation | The tab | Roughly |
|---|---|---|
| A voiceover, a screen recording, a podcast cut — nobody is on camera | Voice Swap | ✦0.17 a second |
| A face on camera, and the words are staying the same | Voice Swap, then Lipsync in video mode | ✦0.17 plus ✦0.35 a second |
| A face on camera, and the words are changing | Speech, then Lipsync in video mode | Per 1,000 characters, plus ✦0.35 a second |
| A still photo that has to speak | Lipsync in photo mode | From ✦0.55 a second |
Voice Swap on its own leaves the original mouth movement in place. That reads fine when the speaker is off camera or in a wide shot, and reads wrong in a close-up. That is the whole decision.
The minimum charge is worth planning around too: anything up to about six seconds costs ✦1 whatever its actual length, so a handful of short lines is cheaper re-voiced as one continuous take and cut afterwards than as five separate jobs.
Open the Voice Swap tab — Opens the studio on Voice Swap. Upload or record there; the cost is quoted before anything runs.
No. It replaces the voice and keeps the words exactly as they are. To change the language you need a translated script, speech generated from it, and usually a lipsync pass to match the mouth.
✦0.17 per second of media, with a five-second minimum. A thirty-second clip is ✦5.1 and the five-minute maximum is ✦51, all in ✦ gold credits.
Up to 300 seconds — five minutes — measured on the file you upload. The same cap applies whether you send a video or an audio file.
Yes. Cloned voices appear in the same picker as the ten presets, and the per-second price is identical either way.
No. The audio is extracted, re-voiced and muxed back over the original video stream, so the picture you get back is the picture you uploaded.
No. Voice Swap sits behind the Audio Studio's Starter gate, which is enforced by the server, and it bills the ✦ gold wallet.