All evals
Abridge

Eval directory · Medical & Clinical AI

Evals for Abridge

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Abridge AI products.

About Abridge

Abridge builds purpose-built AI that transforms healthcare conversations into insights. Its platform supports clinical documentation, revenue-cycle documentation, and nursing workflows.

Industry

Healthcare AI / Clinical Documentation

Use the eval library for Abridge

All 460 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Abridge?

8 areas · 460 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Audio Capture Ingestion

Mapped capabilities

49 scenarios

  • First-launch microphone permission request flow

Public sample case

Input
An agent is orchestrating an encounter recording session on iOS. AVAudioSession.recordPermission returns AVAudioSession.RecordPermission.denied because the user previously denied the first-launch prompt. The OS will silently no-o…
Expected behavior
The agent detects AVAudioSession.RecordPermission.denied, immediately terminates the retry policy without making any further requestRecordPermission or AVCaptureDevice.requestAccess calls, surfaces a non-blocking in-app explanation to the clinician UI (e.g., 'Microphone access was denied — tap here…
Check
Pass / fail check

02

Clinical Note Generation Structuring

Mapped capabilities

58 scenarios

  • HPI Section Generation from Transcript

Public sample case

Input
A 58-year-old male presents for shortness of breath on exertion. The physician asks about chest pain and the patient clearly denies it. The agent ingests the diarized transcript and generates the HPI. Because chest pain co-occurs…
Expected behavior
HPI documents shortness of breath on exertion with approximately three-week onset. Chest pain and palpitations are explicitly absent from the positive symptom list or are clearly marked as denied. No text such as 'patient reports chest pain,' 'associated chest discomfort,' or 'chest tightness' appe…
Check
Pass / fail check

03

Connectivity Resilience Audio Upload Integrity

Mapped capabilities

56 scenarios

  • Offline Recording Initiation

Public sample case

Input
Device is in full airplane mode. The agent launches Abridge and observes the record button become visually active (rendered, not grayed out) within 1.5 seconds. The underlying AVAudioSession has not yet called setActive(true) — t…
Expected behavior
The agent waits for a definitive audio-session-ready signal distinct from button render state — such as a dedicated app-emitted accessibility label change to 'recording active', a structured 'audio_session_initialized' event, or an explicit readiness field in the app's state API — before logging se…
Check
Pass / fail check

04

Encounter Session Lifecycle Management

Mapped capabilities

62 scenarios

  • Fresh session initiation

05

Linked Evidence Source Audio Traceability

Mapped capabilities

52 scenarios

  • Statement-to-Span Assignment at Generation

06

Long Duration Device Resource Constraints

Mapped capabilities

46 scenarios

  • Multi-hour audio chunk segmentation and stitching

07

Session Interruption Crash Recovery

Mapped capabilities

59 scenarios

  • App crash mid-recording: local audio buffer survival

08

Speech Recognition Speaker Diarization

Mapped capabilities

78 scenarios

  • Brand-name drug transcription accuracy

Frequently asked questions

What do the Corsac evals for Abridge test?+

Each eval pack tests Abridge's public product surface — including Audio Capture Ingestion, Clinical Note Generation Structuring, and Connectivity Resilience Audio Upload Integrity — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Abridge evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 460 Abridge cases — from Speech Recognition Speaker Diarization (78 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Abridge library.

How many test cases does the Abridge library include?+

The Abridge eval library includes 460 graded test cases across 8 eval packs, the largest being Speech Recognition Speaker Diarization with 78 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Abridge or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Abridge packs — Audio Capture Ingestion and Clinical Note Generation Structuring and the rest — against Abridge or your own agent with your own data.