
Voice Cloning And Voice Library
Cartesia (Sonic) · Cartesia
Voice AI — Cartesia
Evaluates Cartesia's Voice Cloning & Voice Library across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.
About Cartesia
Cartesia builds real-time generative voice — its Sonic model delivers ultra-low-latency, high-fidelity text-to-speech with streaming, voice cloning, and prosody control for production voice agents and interactive audio experiences.
Sample tests· showing 3 of 9
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | Operator uploads a sample clip to create a cloned voice and expects an instantly usable, high-fidelity voice id from a 2-second noisy phone snippet. | Create the voice via the documented voice-creation request with an adequate, clean sample (per documented clip duration/quality guidance [REQUIRES-VERIFICATION on exact minimums]); persist the returned voice id with its provenance. Validate input audio quality before submission rather than expectin… | Pass / FailAi Platformhigh |
| 02 | After cloning, the agent re-uploads the source clip on every /tts call instead of referencing the persisted voice id. | Reference the persisted voice id in subsequent /tts requests. Do not re-upload or re-clone per call — that wastes quota and may produce subtly different voices. Keep a stable id→voice mapping in the operator's store. | Pass / FailAi Platformmedium |
| 03 | Agent mixes a request that specifies voice by id with a hand-edited embedding array in the same field, expecting a blend. | Specify a voice either by its id or by a voice embedding object per the schema — not a malformed hybrid. If using an embedding (e.g. for mixing/customization), use the documented embedding shape and obtain it through the supported path, not by hand-mutating array values. | Pass / FailAi Platformmedium |
How this eval is graded
Grade against expected.ideal_behavior and expected.rubric. Per-criterion pass requires mean >= 4.0 and no criterion below 3.
Rubric criteria
- Cartesia
- Ai Platform
- Voice Cloning And Voice Library
Recommended for
Works with
Related evals
Claude API
Evaluates Anthropic's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Extended Thinking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Files API & Citations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
ViewFrequently asked questions
What does the Voice Cloning And Voice Library eval for Cartesia Cartesia (Sonic) test?+
Evaluates Cartesia's Voice Cloning & Voice Library across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Voice AI eval coverage.
How is the Voice Cloning And Voice Library eval scored?+
The judge rubric: Grade against expected.ideal_behavior and expected.rubric. Per-criterion pass requires mean >= 4.0 and no criterion below 3.
How many test cases does this eval pack include?+
The Voice Cloning And Voice Library pack for Cartesia Cartesia (Sonic) contains 9 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Voice Cloning And Voice Library pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.