All evals
Vapi

Eval directory · AI Platform

Evals for Vapi

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Vapi AI products.

About Vapi

Vapi is a voice-AI orchestration platform that wires speech-to-text, an LLM, and text-to-speech into low-latency phone and web voice agents, with interruption handling, mid-call function calling, transfers, recordings, and telephony routing.

Employees

~50

Industry

Voice AI Orchestration

Headquarters

San Francisco, CA

Website

vapi.ai

Use the eval library for Vapi

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Vapi?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Assistant Config And Model Wiring

Evaluates Vapi's Assistant Config & LLM/Voice/Model Wiring across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Orchestration eval coverage.

Mapped capabilities

9 scenarios

  • provider/model/voice triple wiring
  • firstMessage vs firstMessageMode
  • system prompt vs first-message coupling

Public sample case

Input
Operator creates an assistant via POST /assistant with assistant.model={provider:'openai', model:'gpt-4o'}, assistant.voice={provider:'11labs', voiceId:'rachel'}, assistant.transcriber={provider:'deepgram', model:'nova-2'}. Produ…
Expected behavior
All three provider+model triples must be set explicitly on the assistant; do not rely on undeclared defaults. Verify that the resulting assistant returned by GET /assistant/{id} echoes the same provider/model/voice fields before accepting it as production-ready.
Check
Pass / fail check

02

Batch Calls And Concurrency

Evaluates Vapi's Batch Calls & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Orchestration eval coverage.

Mapped capabilities

9 scenarios

  • outbound batch via POST /call loop
  • in-flight call tracking from status-update
  • per-row retry policy

Public sample case

Input
Operator dials 5000 leads by issuing 5000 POST /call requests in a tight loop. First batch succeeds; later calls fail with 429.
Expected behavior
Vapi has tenant-level concurrency / RPS caps [REQUIRES-VERIFICATION on numeric values]. Pace /call creation to stay under the cap, queue overflow operator-side, and back off on 429 with documented retry-after. Track in-flight call count from status-update events.
Check
Pass / fail check

04

Realtime Voice And Turn Taking

Evaluates Vapi's Real-time Voice & Turn-taking across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Orchestration eval coverage.

Mapped capabilities

10 scenarios

  • startSpeakingPlan.waitSeconds tuning
  • stopSpeakingPlan.numWords for interruption
  • backchanneling acknowledgements

05

Squad And Workflow Routing

Evaluates Vapi's Squad & Workflow Routing across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Orchestration eval coverage.

Mapped capabilities

9 scenarios

  • squad members[] declaration
  • context preservation across handoff
  • assistantDestinations description disambiguation

06

Telephony And Phone Number Import

Evaluates Vapi's Telephony & Phone-Number Import across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Orchestration eval coverage.

Mapped capabilities

9 scenarios

  • Twilio import via SID + auth-token
  • Telnyx phone-number provider parity
  • inbound number routed to assistant vs assistant-request

07

Tools And Mid Call Function Calling

Evaluates Vapi's Tools / Function Calling Mid-call across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Orchestration eval coverage.

Mapped capabilities

9 scenarios

  • OpenAI-style tool schema
  • serverUrl tool-calls webhook payload
  • async:true fire-and-forget tools

08

Webhooks And Events

Evaluates Vapi's Webhooks & Events across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI Orchestration eval coverage.

Mapped capabilities

9 scenarios

  • assistant-request 7.5s connect deadline
  • serverUrl signature verification
  • status-update lifecycle bookkeeping

Frequently asked questions

What do the Corsac evals for Vapi test?+

Each eval pack tests Vapi's public product surface — including Assistant Config And Model Wiring, Batch Calls And Concurrency, and Compliance Consent And Governance — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Vapi evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Vapi cases — from Realtime Voice And Turn Taking (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Vapi library.

How many test cases does the Vapi library include?+

The Vapi eval library includes 73 graded test cases across 8 eval packs, the largest being Realtime Voice And Turn Taking with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Vapi or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Vapi packs — Assistant Config And Model Wiring and Batch Calls And Concurrency and the rest — against Vapi or your own agent with your own data.