All evals
LiveKit

Eval directory · AI Platform

Evals for LiveKit

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for LiveKit AI products.

About LiveKit

LiveKit is open-source real-time voice/video infrastructure used to build voice agents and live experiences — a WebRTC SFU, telephony (SIP), recording/egress, and the LiveKit Agents framework for STT→LLM→TTS pipelines, available as LiveKit Cloud and self-hosted.

Employees

~50

Industry

Voice AI Infrastructure

Headquarters

New York, NY

Website

livekit.io

Use the eval library for LiveKit

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for LiveKit?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Auth And Tokens

Evaluates LiveKit's Auth & Tokens across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

9 scenarios

  • API key + secret signing
  • short TTL on access tokens
  • video grants — can_publish/can_subscribe

Public sample case

Input
Frontend code includes the LiveKit API secret to mint JWTs client-side for rapid prototyping.
Expected behavior
API secret MUST stay server-side. Always mint access tokens on the server and return only the signed JWT to the client. Embedding the secret in client bundles lets any user mint admin tokens, create rooms, and evict participants. Rotate the secret immediately on suspected leak.
Check
Pass / fail check

02

Cloud Vs Self Host And Scaling

Evaluates LiveKit's Cloud vs Self-host & Scaling across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

9 scenarios

  • Cloud vs self-host decision
  • multi-node self-host requires Redis
  • egress worker as separate process

Public sample case

Input
Operator wants the cheapest 30-participant beta with minimal ops; later expects to scale to 5000-room production.
Expected behavior
LiveKit Cloud provides the SFU + edge mesh + SIP + egress as a managed service with regions; self-host gives full control on infrastructure cost but requires k8s, Redis, egress workers, TURN/TLS, observability, and capacity planning. Recommend Cloud for the beta and revisit only if cost or data-res…
Check
Pass / fail check

03

Livekit Agents Framework

Evaluates LiveKit's LiveKit Agents Framework across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

9 scenarios

  • agent worker registration
  • VoicePipelineAgent STT→LLM→TTS chain
  • VAD-based turn detection

Public sample case

Input
Operator runs `python agent.py start` to register a worker against LiveKit Cloud. Worker prints 'registered' but no agent ever joins a room.
Expected behavior
Worker registration only advertises availability — agents join rooms via (a) automatic dispatch matching the worker's room-name pattern, or (b) explicit AgentDispatchService.CreateDispatch from server code. Verify the dispatch path is wired; do not assume 'registered' implies 'joined'.
Check
Pass / fail check

04

Recording And Egress

Evaluates LiveKit's Recording & Egress across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

9 scenarios

  • room composite vs track egress
  • S3 sink credentials lifecycle
  • segmented HLS playlist

05

Rooms Participants And Tracks

Evaluates LiveKit's Rooms, Participants & Tracks across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

9 scenarios

  • participant identity collision
  • room metadata is opaque string
  • empty_timeout and departure_timeout

06

Safety Compliance And Governance

Evaluates LiveKit's Safety, Compliance & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

10 scenarios

  • AI identity disclosure
  • PII / PHI in transcripts
  • GDPR right to erasure

07

Sfu And Media Transport

Evaluates LiveKit's SFU & Media Transport across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

9 scenarios

  • ICE/TURN fallback
  • codec preference negotiation
  • region routing on Cloud

08

Telephony Sip

Evaluates LiveKit's Telephony (SIP) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Real-time Voice & Video Infra eval coverage.

Mapped capabilities

9 scenarios

  • inbound SIP trunk dispatch rule
  • outbound SIP CreateSIPParticipant
  • DTMF over RFC 2833

Frequently asked questions

What do the Corsac evals for LiveKit test?+

Each eval pack tests LiveKit's public product surface — including Auth And Tokens, Cloud Vs Self Host And Scaling, and Livekit Agents Framework — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the LiveKit evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 LiveKit cases — from Safety Compliance And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the LiveKit library.

How many test cases does the LiveKit library include?+

The LiveKit eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Compliance And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against LiveKit or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 LiveKit packs — Auth And Tokens and Cloud Vs Self Host And Scaling and the rest — against LiveKit or your own agent with your own data.