All evals
ElevenLabs

Eval directory · Voice AI

Evals for ElevenLabs

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for ElevenLabs AI products.

About ElevenLabs

ElevenLabs is a voice AI platform that generates natural, expressive speech for any use case — from conversational agents and audiobooks to real-time dubbing and voice cloning. Its API powers production voice experiences across more than 30 languages.

Employees

~200

Industry

Voice AI

Headquarters

New York, NY

Use the eval library for ElevenLabs

All 70 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Related in Voice AI

All evals →

More Voice AI eval libraries

Coverage map

What would you measure for ElevenLabs?

7 areas · 70 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Conversational Ai Agents Tool Use

Evaluates ElevenLabs' Conversational AI Agents & Tool Use across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI platform eval coverage.

Mapped capabilities

10 scenarios

  • turn-taking/barge-in
  • tool-use audit
  • KB grounding refusal

Public sample case

Input
Conversational AI WebSocket session; user interrupts mid-sentence; half-duplex bridge.
Expected behavior
Cancel in-flight TTS on barge-in event; capture new STT utterance; resume dialog state without repeating cancelled phrase.
Check
Pass / fail check

02

Dubbing Multilingual Localization

Evaluates ElevenLabs' Dubbing & Multilingual Localization across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI platform eval coverage.

Mapped capabilities

9 scenarios

  • RTL Arabic
  • CJK timing agent behavior
  • lip-sync retry NOT MOS scoring

Public sample case

Input
Source English video 12 min; target ar; subtitles must render RTL in player bundle.
Expected behavior
Create dubbing project with target ar; verify RTL metadata in subtitle export; adjust TTS pace for MSA clarity.
Check
Pass / fail check

03

Speech To Text Scribe

Evaluates ElevenLabs' Speech-to-Text (Scribe) across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI platform eval coverage.

Mapped capabilities

10 scenarios

  • diarization
  • code-switch
  • medical/legal vocabulary

Public sample case

Input
45-minute WAV; Scribe STT with diarization enabled; output for CMS chapter markers.
Expected behavior
POST /v1/speech-to-text with diarization flags; return segments with speaker_id labels; preserve timestamps.
Check
Pass / fail check

04

Trust Safety Moderation Provenance

Evaluates ElevenLabs' Trust, Safety, Moderation & Provenance across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI platform eval coverage.

Mapped capabilities

9 scenarios

  • profanity input moderation
  • hate speech refusal
  • C2PA/watermark agent check

05

Tts Generation Models Streaming

Evaluates ElevenLabs' TTS Generation, Models & Streaming across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI platform eval coverage.

Mapped capabilities

10 scenarios

  • model selection
  • SSML/break tags
  • streaming latency

07

Voice Design Voice Library

Evaluates ElevenLabs' Voice Design & Voice Library across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI platform eval coverage.

Mapped capabilities

10 scenarios

  • brief fidelity
  • library sharing
  • voice_settings stability

Frequently asked questions

What do the Corsac evals for ElevenLabs test?+

Each eval pack tests ElevenLabs's public product surface — including Conversational Ai Agents Tool Use, Dubbing Multilingual Localization, and Speech To Text Scribe — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the ElevenLabs evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 70 ElevenLabs cases — from Voice Cloning Consent Abuse Resistance (12 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the ElevenLabs library.

How many test cases does the ElevenLabs library include?+

The ElevenLabs eval library includes 70 graded test cases across 7 eval packs, the largest being Voice Cloning Consent Abuse Resistance with 12 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against ElevenLabs or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 7 ElevenLabs packs — Conversational Ai Agents Tool Use and Dubbing Multilingual Localization and the rest — against ElevenLabs or your own agent with your own data.