Model chooser  /  Transcribe audio, calls and meetings  /  One-off

To transcribe a recording once, use ElevenLabs Scribe v2.

Scribe v2 has one of the lowest word error rates on Artificial Analysis's speech-to-text benchmark from a major vendor, labels who is speaking, and runs in a web app where you upload a file and download the transcript.

Why ElevenLabs Scribe v2

The reasons it wins for this job.

  • It scores 2.2% word error rate on Artificial Analysis's speech-to-text benchmark, among the lowest listed and the lowest from a vendor with a consumer app.
  • It labels speakers, which is what makes a sales call or interview transcript usable.
  • The ElevenLabs web app takes a file upload and returns a transcript with no code.
  • The same model has an API, so a one-off that becomes a weekly job does not need a new vendor.

How to use it

4 steps to a first result.

  1. 1Use the original fileUpload the recording itself, not a re-recording from speakers. Audio quality is the biggest factor in accuracy.
  2. 2Turn on speaker labelsName the speakers after it separates them.
  3. 3Skim the proper nounsProduct names, people and companies are where transcripts go wrong. Fix them once, search-and-replace the rest.
  4. 4Summarise with a text modelPaste the transcript into Claude or ChatGPT for notes, action items or quotes.

Prompt for summarising the transcript

Start from this.

Edit the parts in capitals, then run it.

Below is a transcript of a CALL TYPE between SPEAKER A (us) and SPEAKER B (them).

Give me:
1. Their situation and the problem they described, in their words where possible
2. Objections or concerns they raised, each with a direct quote
3. What was agreed and who owns each next step
4. Anything they said that contradicts what is in our CRM: PASTE CRM NOTES

Quote exactly. If something is ambiguous in the transcript, say so.

TRANSCRIPT:
PASTE HERE
Nothing is sent anywhere. It copies to your clipboard.

Alternatives that also work

If ElevenLabs Scribe v2 is not an option.

Qwen3-ASR 1.7BHugging Face →

Alibaba Qwen · open weights · cost: free weights; you pay for the hardware · 2.5M HF downloads/mo

Pick it when the recording is confidential and must stay on your machine. It is the best open model on the Hugging Face Open ASR leaderboard (4.31 average WER).

Gemini 3.8 FlashOfficial page →

Google · closed · cost: low

Pick it when you want the transcript and the summary in one step. It takes audio directly.

Deepgram Nova-3Official page →

Deepgram · closed · cost: low

Pick it when you need the transcript in seconds rather than minutes. It is the fastest model on Artificial Analysis's benchmark.

Watch out for

Doing this again and again? To transcribe calls every week, use the ElevenLabs Scribe v2 API. A call-notes pipeline needs accurate transcripts with speaker labels at a price per hour you can scale. Scribe v2 combines one of the lowest error rates on Artificial Analysis's benchmark with a low per-hour price that ElevenLabs cut further in August 2026.

Sources

Checked . Models change monthly; we re-check this page when they do.

Picking the model is the easy part.

Wiring it into a workflow that runs every week, with evals, fallbacks and a cost you can predict, is the work. Fifteen minutes, no deck.

Book fifteen minutes →