All evals
AssemblyAI

Eval directory · AI Platform

Evals for AssemblyAI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for AssemblyAI AI products.

About AssemblyAI

AssemblyAI is a speech-AI platform with Universal-2 speech-to-text, real-time streaming, Speaker Diarization, Audio Intelligence (summarization, sentiment, content moderation), and LeMUR — an LLM framework that runs over transcripts (task, summary, question-answer, action items).

Employees

~150

Industry

Speech AI

Headquarters

San Francisco, CA

Use the eval library for AssemblyAI

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for AssemblyAI?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Audio Intelligence

Evaluates AssemblyAI's Audio Intelligence across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • summarization parameter combos
  • sentiment_analysis per-utterance
  • entity_detection types

Public sample case

Input
Agent sets summarization=true with summary_model='catchy' and summary_type='bullets_verbose' on a 90-minute legal deposition.
Expected behavior
summary_model controls voice (informative | conversational | catchy); summary_type controls shape (bullets | bullets_verbose | gist | headline | paragraph). Match to use-case — 'catchy' marketing tone is wrong for legal. Verify the combination is documented as compatible; not all combos are.
Check
Pass / fail check

02

Auth Rate Limits Concurrency Governance

Evaluates AssemblyAI's Auth, Rate Limits, Concurrency & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

10 scenarios

  • Authorization header (no Bearer prefix)
  • 429 + Retry-After backoff
  • concurrent streaming session cap

Public sample case

Input
Agent passes Authorization: Bearer <api_key> on POST /v2/transcript thinking AssemblyAI uses standard Bearer auth. Server returns 401.
Expected behavior
AssemblyAI's Authorization header carries the raw API key (no 'Bearer ' prefix). Set Authorization: <api_key>. Do not assume Bearer is universal; different vendors differ. Read the auth doc and pin the header shape in code.
Check
Pass / fail check

03

Batch Transcription Universal 2

Evaluates AssemblyAI's Batch Transcription (Universal-2) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • audio_url vs upload
  • speech_model selection
  • language_detection vs language_code

Public sample case

Input
Agent submits POST /v2/transcript with audio_url pointing to a presigned S3 URL that expires in 60 seconds. Universal-2 processing starts 3 minutes later.
Expected behavior
Either (a) upload the audio bytes via POST /v2/upload first and use the returned upload_url as audio_url (single-use; AssemblyAI fetches synchronously before returning), or (b) ensure audio_url remains fetchable for the full queue+processing window. Do not assume Universal-2 dereferences audio_url …
Check
Pass / fail check

04

Lemur

Evaluates AssemblyAI's LeMUR across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • transcript_ids vs input_text
  • final_model selection
  • question_answer answer_format

05

Speaker Labels And Diarization

Evaluates AssemblyAI's Speaker Labels & Diarization across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • speaker_labels on async only
  • speakers_expected hint
  • utterances vs words for speaker grouping

06

Streaming Stt Realtime

Evaluates AssemblyAI's Streaming STT (Real-time) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • sample_rate query param match
  • PartialTranscript vs FinalTranscript
  • end_utterance_silence_threshold tuning

07

Transcript Features

Evaluates AssemblyAI's Transcript Features across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • punctuate and format_text
  • disfluencies
  • filter_profanity

08

Webhooks And Async Delivery

Evaluates AssemblyAI's Webhooks & Async Delivery across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Speech AI Platform eval coverage.

Mapped capabilities

9 scenarios

  • webhook_url + secret shape
  • retry on non-2xx
  • idempotency by transcript_id

Frequently asked questions

What do the Corsac evals for AssemblyAI test?+

Each eval pack tests AssemblyAI's public product surface — including Audio Intelligence, Auth Rate Limits Concurrency Governance, and Batch Transcription Universal 2 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the AssemblyAI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 AssemblyAI cases — from Auth Rate Limits Concurrency Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the AssemblyAI library.

How many test cases does the AssemblyAI library include?+

The AssemblyAI eval library includes 73 graded test cases across 8 eval packs, the largest being Auth Rate Limits Concurrency Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against AssemblyAI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 AssemblyAI packs — Audio Intelligence and Auth Rate Limits Concurrency Governance and the rest — against AssemblyAI or your own agent with your own data.