All evals
Ambience Healthcare

Eval directory · Medical & Clinical AI

Evals for Ambience Healthcare

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Ambience Healthcare AI products.

About Ambience Healthcare

Ambience Healthcare provides an AI platform for documentation, coding, and clinical workflow in health systems. It integrates with major EHRs across outpatient, emergency, and inpatient settings.

Industry

Healthcare AI / Clinical Documentation and Coding

Headquarters

San Francisco, CA

Use the eval library for Ambience Healthcare

All 379 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Ambience Healthcare?

6 areas · 379 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Clinical Note Generation

Mapped capabilities

63 scenarios

  • Chief Complaint Extraction

Public sample case

Input
Patient opens the visit by listing three distinct concerns in a single turn, in order: a medication refill, knee pain of a few weeks duration, and exertional chest tightness. The agent must apply clinical acuity ranking, not posi…
Expected behavior
CC field is populated with 'chest tightness on exertion' or a semantically equivalent phrase (e.g., 'exertional chest tightness'). Medication refill and knee pain appear as additional concerns in the HPI or a secondary concerns list, not as the primary CC. No cardiac diagnosis label (e.g., 'angina'…
Check
Pass / fail check

02

Audio Capture Ambient Recording Lifecycle

Mapped capabilities

57 scenarios

  • Microphone Permission — Initial Grant

Public sample case

Input
navigator.permissions.query({name:'microphone'}) returns state 'prompt'. The agent is orchestrating the onboarding flow and the recording UI component has just begun rendering but has not yet fired its mount/ready event. The agen…
Expected behavior
Agent waits for a confirmed UI-mount signal (e.g., a DOM-ready or component-lifecycle callback) before calling getUserMedia, ensuring the permission prompt fires in a visible foreground context. After the user clicks Allow, agent verifies AudioTrack.readyState === 'live' before starting the session…
Check
Pass / fail check

03

Clinical Documentation Integrity Cdi

Mapped capabilities

65 scenarios

  • Assessment-Plan Diagnosis Gap Detection

Public sample case

Input
The physician's note contains the ROS entry: 'Cardiovascular: Denies chest pain, palpitations, or dyspnea on exertion.' The A/P addresses hypertension and hyperlipidemia but does not list chest pain. The agent is configured to au…
Expected behavior
The agent does not flag chest pain as a documentation gap and does not generate a CDI query asking the physician to justify its absence from the A/P. The gap report includes only conditions with affirmative clinical evidence. Zero references to chest pain, palpitations, or dyspnea appear in any gap…
Check
Pass / fail check

04

E M Level Determination Mdm And Time Based

Mapped capabilities

55 scenarios

  • Self-Limited/Minor Problem Classification

05

Point Of Care Coding Cpt Icd 10 Suggestion Accuracy

Mapped capabilities

60 scenarios

  • Outpatient E/M Level Assignment — Time Pathway

06

Speech Recognition Diarization Transcription

Mapped capabilities

79 scenarios

  • Brand-name drug recognition

Frequently asked questions

What do the Corsac evals for Ambience Healthcare test?+

Each eval pack tests Ambience Healthcare's public product surface — including Ambient Clinical Note Generation, Audio Capture Ambient Recording Lifecycle, and Clinical Documentation Integrity Cdi — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Ambience Healthcare evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 379 Ambience Healthcare cases — from Speech Recognition Diarization Transcription (79 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Ambience Healthcare library.

How many test cases does the Ambience Healthcare library include?+

The Ambience Healthcare eval library includes 379 graded test cases across 6 eval packs, the largest being Speech Recognition Diarization Transcription with 79 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Ambience Healthcare or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Ambience Healthcare packs — Ambient Clinical Note Generation and Audio Capture Ambient Recording Lifecycle and the rest — against Ambience Healthcare or your own agent with your own data.