All evals
Retell AI

Eval directory · AI Platform

Evals for Retell AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Retell AI AI products.

About Retell AI

Retell AI is a platform for building production phone-call voice agents — pairing a conversation engine with telephony, low-latency turn-taking, interruption handling, mid-call functions, post-call analysis, and batch outbound dialing.

Employees

~40

Industry

Voice AI Agents

Headquarters

San Francisco, CA

Use the eval library for Retell AI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Retell AI?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Retell Agent Config And Llm

Evaluates Retell AI's Agent Configuration & LLM across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Agents eval coverage.

Mapped capabilities

9 scenarios

  • response_engine type selection
  • Retell LLM created before agent
  • begin_message empty vs absent

Public sample case

Input
Operator creates an agent and must choose response_engine. They set response_engine.type='retell-llm' but also supply an llm_websocket_url (which belongs to custom-llm).
Expected behavior
response_engine.type selects exactly one engine: 'retell-llm' (managed, references a retell_llm_id), 'custom-llm' (your llm_websocket_url), or 'conversation-flow' (a conversation_flow_id). Pick one and supply only its fields — llm_websocket_url is ignored/invalid under retell-llm. Do not mix.
Check
Pass / fail check

02

Retell Batch Calls And Concurrency

Evaluates Retell AI's Batch Calls & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Agents eval coverage.

Mapped capabilities

9 scenarios

  • create batch call with task list
  • Concurrency API check before launch
  • batch queuing vs concurrency limit

Public sample case

Input
Operator launches an outbound reminder campaign for 5,000 contacts and tries to fire 5,000 individual Create Call requests in a tight loop.
Expected behavior
Use Batch Call to enqueue the task list (per-recipient numbers + dynamic variables) so Retell paces dialing within limits, rather than hammering Create Call. Batch handles queuing/pacing; per-call loops risk rate limits and concurrency rejections.
Check
Pass / fail check

04

Retell Conversation Flow And Response Engine

Evaluates Retell AI's Conversation Flow & Response Engine across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Agents eval coverage.

Mapped capabilities

9 scenarios

  • custom-LLM WebSocket handshake
  • response_required event handling
  • reminder_required vs response_required

05

Retell Function Calling Custom Tools

Evaluates Retell AI's Function Calling / Custom Tools across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Agents eval coverage.

Mapped capabilities

9 scenarios

  • custom function schema definition
  • speak-during-execution filler
  • function timeout handling

06

Retell Realtime Voice And Interruption

Evaluates Retell AI's Real-time Voice & Interruption across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Agents eval coverage.

Mapped capabilities

9 scenarios

  • interruption_sensitivity tuning
  • barge-in stops TTS promptly
  • backchannel configuration

07

Retell Telephony And Call Lifecycle

Evaluates Retell AI's Telephony & Call Lifecycle across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Agents eval coverage.

Mapped capabilities

9 scenarios

  • Register Phone Call before SIP dial
  • 5-minute registration window
  • inbound vs outbound agent binding

08

Retell Webhooks And Post Call

Evaluates Retell AI's Webhooks & Post-call across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Agents eval coverage.

Mapped capabilities

9 scenarios

  • x-retell-signature verification
  • event type discrimination
  • call_analyzed arrives after call_ended

Frequently asked questions

What do the Corsac evals for Retell AI test?+

Each eval pack tests Retell AI's public product surface — including Retell Agent Config And Llm, Retell Batch Calls And Concurrency, and Retell Compliance Consent Governance — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Retell AI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Retell AI cases — from Retell Compliance Consent Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Retell AI library.

How many test cases does the Retell AI library include?+

The Retell AI eval library includes 73 graded test cases across 8 eval packs, the largest being Retell Compliance Consent Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Retell AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Retell AI packs — Retell Agent Config And Llm and Retell Batch Calls And Concurrency and the rest — against Retell AI or your own agent with your own data.