
Eval directory
Evals for Deepgram
12 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Deepgram AI products.
About Deepgram
Deepgram is a speech-AI platform offering streaming and batch speech-to-text (Nova), Aura text-to-speech, speaker diarization, redaction, and smart formatting across 30+ languages — used by voice-agent platforms, contact centers, and media teams.
How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
12/12 packs
Scoring metrics
0/12 packs
Judge rubrics
12/12 packs
Use-case maps
0/12 packs
Pack context
12/12 packs
Available eval packs for Deepgram
12 packs ready to run.
Batch Stt Async Callbacks
Transcription AccuracyEvaluates Deepgram's Batch STT & Async Callbacks across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Model Selection Language Detection
LanguageEvaluates Deepgram's Model Selection & Language Detection across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Pii Phi Redaction
PII LeakageEvaluates Deepgram's PII/PHI Redaction across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Smart Formatting Profanity
Evaluates Deepgram's Smart Formatting & Profanity Filter across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Aura Tts
Evaluates Deepgram's Aura TTS across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Auth Rate Limits And Concurrency
Evaluates Deepgram's Auth, Rate Limits & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Batch Stt Rest
Transcription AccuracyEvaluates Deepgram's Batch STT (REST) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Safety Pii Redaction And Governance
PII LeakageEvaluates Deepgram's Safety, PII Redaction & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Speaker Diarization And Channels
Evaluates Deepgram's Speaker Diarization & Channels across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Streaming Stt Websocket
Transcription AccuracyEvaluates Deepgram's Streaming STT (WebSocket) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Transcript Features
Evaluates Deepgram's Transcript Features across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Voice Agent Api
Evaluates Deepgram's Voice Agent API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
Why eval Deepgram AI
Deepgram's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Deepgram measures four dimensions teams care about most when deploying ai platform agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Deepgram's public product surface and runnable in Corsac with your own data.