How to re-voice a video in another voice

Voice Swap takes speech that already exists and replaces the voice performing it. The words stay the same, the timing stays the same, and on a video the picture is never touched — ffmpeg pulls the audio out, the provider re-voices it, and ffmpeg puts it back over the original video stream. That precision is also the limit, and it is the first thing this page says out loud: this is a voice change, not a translation.

What Voice Swap actually does

Give it a clip of somebody talking and pick a different voice. What comes back is the same performance — same words, same pace, same pauses — spoken by somebody else. On a video only the audio track is regenerated: the picture that goes back out is the one you uploaded, because the video stream is copied rather than re-encoded around the new audio.

There is a second door that is easy to miss. You do not need a video at all. An audio file, or a recording made on the spot from the browser microphone, goes straight through and comes back as an mp3 in the new voice. That is the shortest path from "record yourself reading it" to "a narrator reads it", and it skips speech synthesis entirely.

Which plan: The Audio Studio opens on the Starter plan. That gate lives on the server: every generating audio endpoint answers 403 to a free account, so it is not a greyed-out button you can work around. Everything behind these tabs runs on a paid provider, which means it is billed in ✦ gold credits — ⚡ green credits, the ones that refill daily, cannot pay for any of it.

Watch out: Voice Swap does not translate. It re-performs the words that are already there, in a different voice. If you want a clip in another language the words have to change, and that is a different job — the three steps are below.

How to re-voice a clip

  1. Open the Voice Swap tab. Go to /audio and pick Voice Swap in the rail. It is labelled Swap on a narrow screen.
  2. Upload the source. A video, an audio file, or press record. A browser recording arrives as webm or ogg and is transcoded to mp3 server-side before anything else happens to it.
  3. Stay under five minutes. 300 seconds is the cap, measured on the media you upload, and it applies to both doors. Longer material is refused with a message rather than silently truncated.
  4. Pick the new voice. Any of the ten preset voices, or any voice you have cloned. A cloned voice costs the same here as a preset one — the per-second rate does not move with the voice.
  5. Check the quote. The cost is computed from the length of the media, at ✦0.17 a second with a five-second floor. A thirty-second clip is ✦5.1; the full five minutes is ✦51.
  6. Generate and collect. A video comes back as a video, with the new audio muxed over the untouched picture. An audio source comes back as an mp3. Both land in your assets library.

The numbers

Voice Swap is the cheapest per-second lane in the studio. Everything is billed in ✦ gold.
InputCapPriceWhat comes back
Video file300 seconds✦0.17 a second, 5-second minimumThe same video, new voice on the audio track
Audio file300 seconds✦0.17 a second, 5-second minimumAn mp3 in the new voice
Mic recording300 seconds✦0.17 a second, 5-second minimumAn mp3 in the new voice

At ✦0.17 a second it is roughly half the price of the cheapest lipsync model, because it only has to rebuild the audio. If the face on screen does not have to match the new voice — a voiceover, a screen recording, a podcast clip — this is the tab, not Lipsync.

Dubbing into another language

This is the request behind most searches for dubbing, so here is the honest answer: there is no one-button translation lane on this platform. What exists is the three steps that make one, and each is a separate charge.

  1. Get the words. Caption the clip in the Captions tab to pull a timed transcript out of it, then translate the text — in the chat section on /llms, or anywhere else you like.
  2. Speak the translation. In the Speech tab, ElevenLabs Multilingual HD is the engine built for non-English scripts, at ✦6.7 per started 1,000 characters. A cloned voice will read any language you give it too, at ✦1.7.
  3. Put it back on the face. If the speaker is on camera, the new track needs the mouth to match: take it to the Lipsync tab in video mode, from ✦0.35 a second on LatentSync.

It is more work than one button, and it is what actually works today. Voice Swap stays the right tool when the words do not change — a different narrator, one brand voice across clips recorded by different people, or anonymising a recording.

Voice Swap or Lipsync?

The two tabs overlap enough to be confusing, and picking the wrong one costs money. The question to ask is whether a mouth has to match.

The situationThe tabRoughly
A voiceover, a screen recording, a podcast cut — nobody is on cameraVoice Swap✦0.17 a second
A face on camera, and the words are staying the sameVoice Swap, then Lipsync in video mode✦0.17 plus ✦0.35 a second
A face on camera, and the words are changingSpeech, then Lipsync in video modePer 1,000 characters, plus ✦0.35 a second
A still photo that has to speakLipsync in photo modeFrom ✦0.55 a second

Voice Swap on its own leaves the original mouth movement in place. That reads fine when the speaker is off camera or in a wide shot, and reads wrong in a close-up. That is the whole decision.

The minimum charge is worth planning around too: anything up to about six seconds costs ✦1 whatever its actual length, so a handful of short lines is cheaper re-voiced as one continuous take and cut afterwards than as five separate jobs.

Where Voice Swap stops

Open the Voice Swap tab — Opens the studio on Voice Swap. Upload or record there; the cost is quoted before anything runs.

Does Voice Swap translate my video?

No. It replaces the voice and keeps the words exactly as they are. To change the language you need a translated script, speech generated from it, and usually a lipsync pass to match the mouth.

What does it cost?

✦0.17 per second of media, with a five-second minimum. A thirty-second clip is ✦5.1 and the five-minute maximum is ✦51, all in ✦ gold credits.

How long a video can I re-voice?

Up to 300 seconds — five minutes — measured on the file you upload. The same cap applies whether you send a video or an audio file.

Can I use my own cloned voice?

Yes. Cloned voices appear in the same picker as the ten presets, and the per-second price is identical either way.

Does the video get re-encoded?

No. The audio is extracted, re-voiced and muxed back over the original video stream, so the picture you get back is the picture you uploaded.

Can I do this on a free plan?

No. Voice Swap sits behind the Audio Studio's Starter gate, which is enforced by the server, and it bills the ✦ gold wallet.