Model chooser  /  Generate voiceover and speech  /  One-off

For a one-off voiceover, use ElevenLabs v3 in ElevenLabs Studio.

For a single voiceover, the tool matters as much as the model. ElevenLabs has the most complete app for non-developers: a voice library, voice cloning, dubbing and a studio editor, on a model that ranks in the top tier of the speech arena.

Why ElevenLabs v3

The reasons it wins for this job.

  • ElevenLabs v3 scores Elo 1197 on Artificial Analysis's text-to-speech arena, in the leading group.
  • The Studio app lets you pick a voice, adjust pronunciation, and regenerate one sentence without re-recording the whole script.
  • Voice cloning and dubbing into other languages are in the same product, which covers most marketing voiceover needs.
  • This is a judgment on the tooling. On voice quality alone, Cartesia Sonic 3.6 ranks higher; see the alternatives.

How to use it

4 steps to a first result.

  1. 1Write for the earShort sentences, numbers written as spoken ("twenty percent"), no parentheses.
  2. 2Audition three voicesGenerate the first paragraph in three voices and pick one before doing the rest.
  3. 3Fix pronunciation onceAdd product and company names to the pronunciation settings.
  4. 4Export per paragraphSeparate files make editing against video much easier.

Script prompt for a text model

Start from this.

Edit the parts in capitals, then run it.

Rewrite the text below as a voiceover script for a LENGTH-second VIDEO TYPE.

Rules for the ear:
- Sentences under 15 words
- Write numbers and symbols as they are spoken
- No parentheses, no bullet points
- Mark natural pauses with a line break
- About 150 words per minute of video

Text:
PASTE HERE
Nothing is sent anywhere. It copies to your clipboard.

Alternatives that also work

If ElevenLabs v3 is not an option.

Cartesia Sonic 3.6Official page →

Cartesia · closed · cost: medium

Pick it when you want the best-sounding result and do not need a studio app. It ranks first on the speech arena.

Qwen3-TTS 1.7BHugging Face →

Alibaba Qwen · open weights · cost: free weights; you pay for the hardware · 2.6M HF downloads/mo

Pick it when you want to clone a voice locally from a few seconds of audio. Apache 2.0.

hexgrad · open weights · cost: free weights; you pay for the hardware · 12M HF downloads/mo

Pick it when you need decent speech for free on a laptop, such as internal training clips.

Watch out for

Doing this again and again? For speech at volume or in a voice agent, use Cartesia Sonic 3.6. Sonic 3.6 ranks first on Artificial Analysis's text-to-speech arena, is built for low latency, and is priced per character, which covers both high-volume narration and live voice agents.

Sources

Checked . Models change monthly; we re-check this page when they do.

Picking the model is the easy part.

Wiring it into a workflow that runs every week, with evals, fallbacks and a cost you can predict, is the work. Fifteen minutes, no deck.

Book fifteen minutes →