A cloned voice here is a reference recording, not a training run. You hand the studio one sample of somebody speaking, and from then on any text you write can be spoken in that voice — in the Speech tab, across an audiobook, over a lipsynced portrait, or onto a video that already exists. There is no training queue to sit through. This page covers what a usable sample looks like, what speaking with the clone costs, and the five places the voice can be used.
Most tools that say "voice cloning" mean a training job: upload thirty minutes of audio, wait, get a checkpoint. This is the other kind. Your sample is stored as a file, and every generation sends that file to a zero-shot cloning engine alongside the text. Nothing is trained, nothing is queued, and the voice is usable the second the upload finishes.
That has one consequence worth knowing before you record anything. A clone is only ever as good as its sample, because the sample is what gets copied — room echo, a laptop fan, a phone speaker in the background and the fact that you read the line flatly all come along with the voice. The engine cannot add a warmth that is not in the recording.
Which plan: The Audio Studio opens on the Starter plan. That gate lives on the server: every generating audio endpoint answers 403 to a free account, so it is not a greyed-out button you can work around. Everything behind these tabs runs on a paid provider, which means it is billed in ✦ gold credits — ⚡ green credits, the ones that refill daily, cannot pay for any of it.
Note: Listening is not gated. The voice list and each voice's preview sit outside the paid gate, which is POST-only, so any signed-in account can play the ten preset voices before paying for anything. Generating is the part that needs the plan.
Watch out: A sample over 20 MB is refused outright, with a message saying so. Trim it rather than compressing it into mush: the engine copies artefacts as faithfully as it copies timbre.
Cloning itself carries no per-clone charge — there is no credit line on the upload, because nothing is generated. You pay when the voice speaks. And a clone is the cheapest voice in the studio, at half the rate of the cheapest preset engine, because it runs on a different provider that has not raised its price.
| Voice | Credits per 1,000 characters | Engine behind it |
|---|---|---|
| Cloned voice | ✦1.7 | Qwen3 TTS, zero-shot from your sample |
| ElevenLabs Flash | ✦3.4 | Preset voices — the default, fast and natural |
| ElevenLabs Multilingual HD | ✦6.7 | Preset voices — highest fidelity |
| ElevenLabs v3 Expressive | ✦6.7 | Preset voices — most emotional range |
What it costs: Speech bills per STARTED 1,000 characters, because that is how the provider bills us. A 1,001-character script costs two blocks — the same as a 2,000-character one. It is worth padding a script up to a boundary rather than spilling one character past it.
One thing the interface does not spell out: the engine dropdown has no effect on a cloned voice. A clone always runs on the zero-shot lane at ✦1.7 whatever the picker says, because the sample is the model. Selecting v3 Expressive next to a cloned voice buys you nothing, and costs you nothing either.
Once the voice exists it appears in every picker in the studio that takes a voice. Each surface has its own cap on how much text it will accept in one go, and the caps differ for good reasons — a mix is a scene, an audiobook is a book.
| Tab | What it does with the voice | Text cap in one request |
|---|---|---|
| Speech | Reads a script as one audio file | 30,000 characters |
| Audiobook | Splits a manuscript on blank lines, narrates chapter by chapter | 30,000 characters a chapter |
| Mix | Layers a spoken track under or over sound effects | 5,000 characters a speech track |
| Lipsync | Speaks the line, then drives a face with it | 2,000 characters |
| Voice Swap | Replaces the voice on a video or a recording | Not text — up to 300 seconds of audio |
The sample is the only variable you control, so it is worth thirty seconds of care. Roughly in order of how much difference each one makes:
Being straight about this saves a wasted upload. Four places it ends:
Open the Voices tab — Opens the studio on the voice library. Nothing is generated or charged until you press a generate button.
As long as the upload. There is no training step — the sample is stored and used as a reference on every generation, so the voice is selectable the moment the file finishes uploading.
No. There is no charge on the clone itself. You are charged when the voice speaks, at ✦1.7 per started 1,000 characters, which is half the rate of the cheapest preset engine.
No. The whole Audio Studio opens on the Starter plan and the check is server-side, so a free account is refused by the API rather than by the interface. Signed in, you can still play the preset voice previews on any plan.
Thirty to sixty seconds of clean speech is enough, and the upload is capped at 20 MB. Sample quality matters far more than sample length — the engine copies whatever is in the recording, including the room.
Yes. Cloned voices appear in the Speech, Audiobook, Mix, Lipsync and Voice Swap pickers. Lipsync caps the typed line at 2,000 characters and a mix caps each speech track at 5,000.
Yes, from the Voices tab. Deleting removes it from every picker; anything you already generated with it stays in your assets library.