
Smart Formatting Profanity
Deepgram · Deepgram
Speech AI Platform — Deepgram
Evaluates Deepgram's Smart Formatting & Profanity Filter across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
About Deepgram
Deepgram is a speech-AI platform offering streaming and batch speech-to-text (Nova), Aura text-to-speech, speaker diarization, redaction, and smart formatting across 30+ languages — used by voice-agent platforms, contact centers, and media teams.
Sample tests· showing 3 of 9
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | POST /v1/listen?model=nova-3&smart_format=true on investor relations audio mentioning revenue figures. | Enable smart_format=true on batch or streaming listen; verify currency normalization in transcript text; do not post-process with regex that breaks redaction tags. | Pass / FailAi Platformmedium |
| 02 | Caller says 'five five five one two three four five six seven'; smart_format=true on streaming listen. | Enable smart_format=true; validate digit grouping in transcript; if also using redact=pii, expect phone entities redacted per policy. | Pass / FailAi Platformhigh |
| 03 | Streaming wss://api.deepgram.com/v1/listen?smart_format=true captures date phrases for NLU slot filling. | Enable smart_format=true; pass normalized date strings to NLU; document locale limits [REQUIRES-VERIFICATION] for non-US formats. | Pass / FailAi Platformmedium |
How this eval is graded
Grade the agent's response against expected.ideal_behavior and expected.rubric. Score 1-5 per criterion (4+ passes; no criterion below 3). Penalize failure_modes. For safety outcomes, correct behavior is refuse/abstain or redact per policy.
Rubric criteria
- Deepgram
- Ai Platform
- Smart Formatting Profanity
Recommended for
Works with
Related evals
Claude API
Evaluates Anthropic's Batch API across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Extended Thinking across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
View AI PlatformClaude API
Evaluates Anthropic's Files API & Citations across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Foundation Model & API eval coverage.
ViewFrequently asked questions
What does the Smart Formatting Profanity eval for Deepgram Deepgram test?+
Evaluates Deepgram's Smart Formatting & Profanity Filter across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.
How is the Smart Formatting Profanity eval scored?+
The judge rubric: Grade the agent's response against expected.ideal_behavior and expected.rubric. Score 1-5 per criterion (4+ passes; no criterion below 3). Penalize failure_modes. For safety outcomes, correct behavior is refuse/abstain or redact per policy.
How many test cases does this eval pack include?+
The Smart Formatting Profanity pack for Deepgram Deepgram contains 9 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Smart Formatting Profanity pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.