All evals
Cartesia

Eval directory · AI Platform

Evals for Cartesia

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Cartesia AI products.

About Cartesia

Cartesia builds real-time generative voice — its Sonic model delivers ultra-low-latency, high-fidelity text-to-speech with streaming, voice cloning, and prosody control for production voice agents and interactive audio experiences.

Employees

~40

Industry

Voice AI

Headquarters

San Francisco, CA

Use the eval library for Cartesia

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Cartesia?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Audio Formats And Encoding

Evaluates Cartesia's Audio Formats & Encoding across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • output_format container selection
  • sample_rate match on decode
  • raw PCM encoding (bit depth / signedness)

Public sample case

Input
Agent requests output_format=raw PCM but writes the bytes to a '.wav' file and serves it as audio/wav.
Expected behavior
Match the container to how the bytes will be consumed: request a wav container (with header) when serving a .wav file, or raw PCM only when the consumer knows the encoding/sample_rate out of band. Do not label raw PCM as wav — players will misread the missing header.
Check
Pass / fail check

02

Auth Keys Rate Limits Concurrency

Evaluates Cartesia's Auth, Keys, Rate Limits & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • API key auth placement
  • Cartesia-Version header
  • 429 backoff handling

Public sample case

Input
Agent embeds the Cartesia API key in client-side browser JS to call /tts directly from the user's browser.
Expected behavior
Keep the API key server-side; authenticate /tts requests with the documented key header from a backend (or use a documented short-lived/scoped token mechanism for client-side streaming if available [REQUIRES-VERIFICATION]). Never ship the long-lived key to the browser.
Check
Pass / fail check

03

Realtime Agents Integration

Evaluates Cartesia's Realtime / Agents Integration across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • barge-in / interruption handling
  • duplex audio pipeline
  • latency budget allocation

Public sample case

Input
In a voice agent, the caller starts speaking while Sonic is mid-utterance; the stack keeps playing the synthesized audio over the caller.
Expected behavior
On detected user speech (barge-in), immediately stop local playback AND cancel the in-flight TTS context so the agent yields the floor. Resume the dialog from the interrupted point. Treat interruption as a first-class control path, not an afterthought.
Check
Pass / fail check

05

Sonic Tts Synthesis

Evaluates Cartesia's Sonic TTS Synthesis across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • POST /tts/bytes required fields
  • model_id pinning vs latest
  • /tts/bytes vs /tts/sse selection

06

Streaming Tts And Websocket

Evaluates Cartesia's Streaming TTS & WebSocket across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • WebSocket connect and auth
  • context_id for continuations
  • chunked transcript streaming

07

Voice Cloning And Voice Library

Evaluates Cartesia's Voice Cloning & Voice Library across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • create voice from clip
  • voice id reuse across synth calls
  • voice embedding vs voice id

08

Voice Control And Prosody

Evaluates Cartesia's Voice Control & Prosody across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.

Mapped capabilities

9 scenarios

  • speed control range
  • emotion / expressivity controls
  • pronunciation / custom dictionary

Frequently asked questions

What do the Corsac evals for Cartesia test?+

Each eval pack tests Cartesia's public product surface — including Audio Formats And Encoding, Auth Keys Rate Limits Concurrency, and Realtime Agents Integration — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Cartesia evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Cartesia cases — from Safety Consent And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Cartesia library.

How many test cases does the Cartesia library include?+

The Cartesia eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Consent And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Cartesia or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Cartesia packs — Audio Formats And Encoding and Auth Keys Rate Limits Concurrency and the rest — against Cartesia or your own agent with your own data.