How to turn text into speech

Three ElevenLabs engines and any voice you have cloned, from one text box. The differences between the engines are real and priced accordingly: one is fast and half the price, one is the fidelity option, and one performs bracketed stage directions inline instead of reading them out loud. This page covers which to pick, what each costs per 1,000 characters, how long text is handled, and where the ceiling sits.

The three engines, and which to pick

There is no quality slider here — the engine is the choice, and the price gap between them is exactly two to one. Flash is the default because it is the one most scripts should use. The other two are worth double when you need what they do.

Every price is per STARTED 1,000-character block.
EngineCredits per 1,000 charactersPick it when
ElevenLabs Flash✦3.4Default. Fast and natural — most narration, every draft
ElevenLabs Multilingual HD✦6.7Highest fidelity, and the engine for a script that is not in English
ElevenLabs v3 Expressive✦6.7The widest emotional range, and the only one that performs bracketed tags
Any cloned voice✦1.7Cheapest of all — a clone always runs on its own zero-shot lane

Which plan: The Audio Studio opens on the Starter plan. That gate lives on the server: every generating audio endpoint answers 403 to a free account, so it is not a greyed-out button you can work around. Everything behind these tabs runs on a paid provider, which means it is billed in ✦ gold credits — ⚡ green credits, the ones that refill daily, cannot pay for any of it.

What it costs: Billing is per STARTED 1,000 characters, matching the provider. A 1,001-character script costs two blocks — the same as a 2,000-character one. Trimming a script back under a boundary is the cheapest optimisation available here.

How to generate speech

  1. Open the Speech tab. Go to /audio. Speech is the tab the page opens on; the rail switches between it and Audiobook, Voices, Sounds, Mix, Lipsync, Voice Swap and Captions.
  2. Pick a voice. Ten preset voices ship with the studio — Aria, Roger, Sarah, George, Charlie, Liam, Charlotte, Alice, Brian and Lily — alongside anything you have cloned. Every preset has a free preview you can play first.
  3. Pick the engine. Flash unless you have a reason. Multilingual HD for a non-English script or a render that is going into something finished, and v3 when the delivery matters more than the words.
  4. Paste the text. Up to 30,000 characters in one request. The cost is computed from the character count and shown before you press anything, so there is no guessing at it.
  5. Set the speed if you need to. Between 0.5x and 2x. It is applied to the rendered audio rather than asked of the engine, so it changes pace without dragging the pitch with it.
  6. Generate, then find it in your assets. The result is an mp3 stored against your account, listed in the Audio tab of /assets and reusable as the audio track for Lipsync or Voice Swap.

Directing v3 with bracketed tags

ElevenLabs v3 performs bracketed directions inline rather than reading them out. Write [whispers] before a line and the line is whispered. This is the studio's most expressive feature and it is easy to miss, so the tag palette appears in the composer whenever v3 is selected — and is hidden otherwise, because Flash and Multilingual will happily say the word "whispers" out loud.

The Sound group is the voice acting out an effect, not an effect layered underneath it. For real ambience under a line, generate it in the Sounds tab and combine the two in Mix.

Long text, chapters and audiobooks

Thirty thousand characters is roughly five thousand words, which covers most scripts and no books. Text past a few thousand characters is split into chunks, generated in order and rejoined into one continuous file, so what comes back is a single mp3 rather than a folder of fragments.

For anything longer, the Audiobook tab is the right door. Paste a manuscript and separate chapters with a blank line; each block becomes its own chapter with its own audio file. Chapters are narrated one at a time in sequence, with a progress bar showing which one is in flight — keep the tab open while it runs. The per-request character cap applies per chapter, and pricing is identical to the Speech tab.

Voices, and hearing them before you pay

Every preset voice has a short spoken preview that is generated once, globally, and cached — the same bytes for everyone, which is exactly why it can be free. Cloned voices need none of that machinery: their preview is the reference sample you uploaded, played straight back.

Previews sit outside the paid gate deliberately. If you are deciding whether the studio is worth a subscription, listening to the voices is the part you should not have to pay for. You do need to be signed in.

Writing text that reads well

Text that looks fine on a page often reads badly out loud, and the fix is always in the script rather than in the engine. Four things carry most of the difference:

Where text to speech stops

LimitThe numberWhy it is there
Characters a request30,000The provider's own request ceiling
Characters on a speech track in a mix5,000A mix is a scene, not an audiobook
Characters on a lipsync line2,000A face has to speak it
Speed0.5x to 2xApplied to the rendered audio, pitch-corrected
PlanStarter and upEnforced at the route, not in React

Watch out: There is no music generation on this platform — no tab, no model row, nothing behind a plan. Speech and sound effects are the two things the studio makes from a prompt. If you need a bed under narration, generate ambience in the Sounds tab and layer it in Mix.

Open the Speech tab — Opens the studio on Speech with nothing typed. The cost is quoted on the button before anything runs.

How much does 1,000 characters of speech cost?

✦3.4 on ElevenLabs Flash, ✦6.7 on Multilingual HD or v3 Expressive, and ✦1.7 with a voice you have cloned. Billing is per started block, so 1,001 characters costs two blocks.

What is the longest text I can send at once?

30,000 characters per request. Longer text is chunked and rejoined into one file automatically. For book-length material use the Audiobook tab, which splits on blank lines and narrates chapter by chapter.

Which engine should I use?

Flash unless you have a reason not to — it is half the price and natural enough for most narration. Multilingual HD for non-English scripts and final renders, v3 Expressive when you need it to act.

How do I make the voice whisper or laugh?

Select ElevenLabs v3 and write the direction in brackets inline, such as [whispers] or [laughs]. Only v3 performs them — the other engines read the bracketed word aloud.

Can I generate speech on a free account?

No. The Audio Studio opens on the Starter plan and the gate is enforced by the server. Signed in, you can play the preset voice previews on any plan, but generating needs a paid plan and pays in ✦ gold credits.