Three ElevenLabs engines and any voice you have cloned, from one text box. The differences between the engines are real and priced accordingly: one is fast and half the price, one is the fidelity option, and one performs bracketed stage directions inline instead of reading them out loud. This page covers which to pick, what each costs per 1,000 characters, how long text is handled, and where the ceiling sits.
There is no quality slider here — the engine is the choice, and the price gap between them is exactly two to one. Flash is the default because it is the one most scripts should use. The other two are worth double when you need what they do.
| Engine | Credits per 1,000 characters | Pick it when |
|---|---|---|
| ElevenLabs Flash | ✦3.4 | Default. Fast and natural — most narration, every draft |
| ElevenLabs Multilingual HD | ✦6.7 | Highest fidelity, and the engine for a script that is not in English |
| ElevenLabs v3 Expressive | ✦6.7 | The widest emotional range, and the only one that performs bracketed tags |
| Any cloned voice | ✦1.7 | Cheapest of all — a clone always runs on its own zero-shot lane |
Which plan: The Audio Studio opens on the Starter plan. That gate lives on the server: every generating audio endpoint answers 403 to a free account, so it is not a greyed-out button you can work around. Everything behind these tabs runs on a paid provider, which means it is billed in ✦ gold credits — ⚡ green credits, the ones that refill daily, cannot pay for any of it.
What it costs: Billing is per STARTED 1,000 characters, matching the provider. A 1,001-character script costs two blocks — the same as a 2,000-character one. Trimming a script back under a boundary is the cheapest optimisation available here.
ElevenLabs v3 performs bracketed directions inline rather than reading them out. Write [whispers] before a line and the line is whispered. This is the studio's most expressive feature and it is easy to miss, so the tag palette appears in the composer whenever v3 is selected — and is hidden otherwise, because Flash and Multilingual will happily say the word "whispers" out loud.
The Sound group is the voice acting out an effect, not an effect layered underneath it. For real ambience under a line, generate it in the Sounds tab and combine the two in Mix.
Thirty thousand characters is roughly five thousand words, which covers most scripts and no books. Text past a few thousand characters is split into chunks, generated in order and rejoined into one continuous file, so what comes back is a single mp3 rather than a folder of fragments.
For anything longer, the Audiobook tab is the right door. Paste a manuscript and separate chapters with a blank line; each block becomes its own chapter with its own audio file. Chapters are narrated one at a time in sequence, with a progress bar showing which one is in flight — keep the tab open while it runs. The per-request character cap applies per chapter, and pricing is identical to the Speech tab.
Every preset voice has a short spoken preview that is generated once, globally, and cached — the same bytes for everyone, which is exactly why it can be free. Cloned voices need none of that machinery: their preview is the reference sample you uploaded, played straight back.
Previews sit outside the paid gate deliberately. If you are deciding whether the studio is worth a subscription, listening to the voices is the part you should not have to pay for. You do need to be signed in.
Text that looks fine on a page often reads badly out loud, and the fix is always in the script rather than in the engine. Four things carry most of the difference:
| Limit | The number | Why it is there |
|---|---|---|
| Characters a request | 30,000 | The provider's own request ceiling |
| Characters on a speech track in a mix | 5,000 | A mix is a scene, not an audiobook |
| Characters on a lipsync line | 2,000 | A face has to speak it |
| Speed | 0.5x to 2x | Applied to the rendered audio, pitch-corrected |
| Plan | Starter and up | Enforced at the route, not in React |
Watch out: There is no music generation on this platform — no tab, no model row, nothing behind a plan. Speech and sound effects are the two things the studio makes from a prompt. If you need a bed under narration, generate ambience in the Sounds tab and layer it in Mix.
Open the Speech tab — Opens the studio on Speech with nothing typed. The cost is quoted on the button before anything runs.
✦3.4 on ElevenLabs Flash, ✦6.7 on Multilingual HD or v3 Expressive, and ✦1.7 with a voice you have cloned. Billing is per started block, so 1,001 characters costs two blocks.
30,000 characters per request. Longer text is chunked and rejoined into one file automatically. For book-length material use the Audiobook tab, which splits on blank lines and narrates chapter by chapter.
Flash unless you have a reason not to — it is half the price and natural enough for most narration. Multilingual HD for non-English scripts and final renders, v3 Expressive when you need it to act.
Select ElevenLabs v3 and write the direction in brackets inline, such as [whispers] or [laughs]. Only v3 performs them — the other engines read the bracketed word aloud.
No. The Audio Studio opens on the Starter plan and the gate is enforced by the server. Signed in, you can play the preset voice previews on any plan, but generating needs a paid plan and pays in ✦ gold credits.