All evals
Deepgram

Eval directory · AI Platform

Evals for Deepgram

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Deepgram AI products.

About Deepgram

Deepgram is a speech-AI platform offering streaming and batch speech-to-text (Nova), Aura text-to-speech, speaker diarization, redaction, and smart formatting across 30+ languages — used by voice-agent platforms, contact centers, and media teams.

Employees

~150

Industry

Speech AI

Headquarters

San Francisco, CA

Use the eval library for Deepgram

All 106 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Deepgram?

12 areas · 106 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Batch Stt Async Callbacks

Evaluates Deepgram's Batch STT & Async Callbacks across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

8 scenarios

  • Callback webhook delivery

Public sample case

Input
Async transcription for 45-minute podcast; customer server flaky during deploy; Deepgram retries callback delivery.
Expected behavior
Accept initial 200 with request_id; implement idempotent webhook handler keyed by request_id; expect up to 10 retries per docs; persist transcript once successfully.
Check
Pass / fail check

02

Model Selection Language Detection

Evaluates Deepgram's Model Selection & Language Detection across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

8 scenarios

  • Nova-3 vs Nova-2 vs Whisper routing

Public sample case

Input
New wss://api.deepgram.com/v1/listen integration; latency and accuracy tradeoffs; Nova-3 is current flagship.
Expected behavior
Default to model=nova-3 for English streaming agent; document fallback path to nova-2 if SKU constraints; measure WER/latency empirically per deployment.
Check
Pass / fail check

03

Pii Phi Redaction

Evaluates Deepgram's PII/PHI Redaction across 8 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

8 scenarios

  • Multi-redact query params

Public sample case

Input
Agent discusses patient name and diagnosis; policy requires redact=pii&redact=phi in query string.
Expected behavior
Pass redact=pii and redact=phi query params per docs; verify entity tags in transcript; never store cleartext in analytics warehouse.
Check
Pass / fail check

04

Smart Formatting Profanity

Evaluates Deepgram's Smart Formatting & Profanity Filter across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • Numerals dates currency

05

Aura Tts

Evaluates Deepgram's Aura TTS across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • REST /v1/speak with text body
  • voice model selection
  • encoding/container/sample_rate

06

Auth Rate Limits And Concurrency

Evaluates Deepgram's Auth, Rate Limits & Concurrency across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • Token auth header scheme
  • project keys with scoped roles
  • temporary scoped tokens for browser clients

07

Batch Stt Rest

Evaluates Deepgram's Batch STT (REST) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • POST /v1/listen with audio URL
  • POST /v1/listen with raw audio body
  • Nova model variant selection

08

Safety Pii Redaction And Governance

Evaluates Deepgram's Safety, PII Redaction & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

10 scenarios

  • redact=pii on transcripts
  • redact=pci for card numbers
  • redact=ssn / numbers

09

Speaker Diarization And Channels

Evaluates Deepgram's Speaker Diarization & Channels across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • diarize=true single-channel
  • diarize_model selection
  • streaming diarization caveat

10

Streaming Stt Websocket

Evaluates Deepgram's Streaming STT (WebSocket) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • WebSocket connect with Authorization Token
  • interim_results vs final transcripts
  • endpointing and utterance_end_ms

11

Transcript Features

Evaluates Deepgram's Transcript Features across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • punctuate=true
  • smart_format=true
  • numerals=true

12

Voice Agent Api

Evaluates Deepgram's Voice Agent API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • agent WebSocket configuration
  • turn-taking and barge-in
  • function call dispatch

Frequently asked questions

What do the Corsac evals for Deepgram test?+

Each eval pack tests Deepgram's public product surface — including Batch Stt Async Callbacks, Model Selection Language Detection, and Pii Phi Redaction — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Deepgram evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 106 Deepgram cases — from Safety Pii Redaction And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Deepgram library.

How many test cases does the Deepgram library include?+

The Deepgram eval library includes 106 graded test cases across 12 eval packs, the largest being Safety Pii Redaction And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Deepgram or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 12 Deepgram packs — Batch Stt Async Callbacks and Model Selection Language Detection and the rest — against Deepgram or your own agent with your own data.