All evals
Commure / Augmedix

Eval directory · Medical & Clinical AI

Evals for Commure / Augmedix

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Commure / Augmedix AI products.

About Commure / Augmedix

Commure is an AI-native healthcare operations platform spanning patient intake, clinical documentation, coding, claims, and payment workflows. Augmedix is its wholly owned subsidiary for ambient AI medical documentation.

Industry

Healthcare AI / Clinical Operations

Headquarters

San Francisco, CA

Use the eval library for Commure / Augmedix

All 395 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Commure / Augmedix?

6 areas · 395 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Note Generation Ai Drafting

Mapped capabilities

64 scenarios

  • HPI Section Drafting from Transcript

Public sample case

Input
A primary care physician sees a 58-year-old patient for a chief complaint of chest pain. The 22-minute transcript contains the patient describing left-sided pressure for two days, worsened by exertion, rated 6/10. No mention of s…
Expected behavior
The generated HPI includes only: left-sided chest pressure, two-day duration, exertional worsening, severity 6/10. It contains no mention of shortness of breath, dyspnea on exertion, diaphoresis, nausea, vomiting, arm radiation, jaw pain, or any other associated symptom not stated in the transcript…
Check
Pass / fail check

02

Audio Capture Ambient Listening

Mapped capabilities

48 scenarios

  • Mobile device built-in mic activation

Public sample case

Input
The agent calls the platform mic-activation API at the start of a patient encounter. The OS returns MIC_ACTIVE=true and session status is 'recording'. However, every PCM sample buffer delivered to the app for a continuous 5-secon…
Expected behavior
Within the configured liveness window (≤5 s of zero-amplitude signal), the agent: (1) emits a CAPTURE_LIVENESS_FAIL log event containing rms_db value, session ID, and encounter ID; (2) surfaces a clinician-visible in-app alert reading 'Microphone not capturing audio — tap to troubleshoot'; (3) paus…
Check
Pass / fail check

03

Ehr Integration Write Back Reconciliation

Mapped capabilities

68 scenarios

  • SMART-on-FHIR EHR-embedded launch

Public sample case

Input
The ambient scribe agent initiates a SMART launch for a new ED patient. The EHR iframe took 52 seconds to load due to EHR SSO latency and a cold-cache JS bundle. The EHR authorization server responds to the token exchange POST wi…
Expected behavior
Agent marks the launch as definitively failed, clears all pending session state, does not read any field from the prior cached context (mrn-0042 / enc-8812 / dr-jones), does not begin ambient audio capture, and surfaces a human-readable re-launch prompt: "Session could not start — please re-open th…
Check
Pass / fail check

04

Hybrid Human In The Loop Review Qa

Mapped capabilities

57 scenarios

  • Queue Ingestion & Classification

05

Medical Coding E M Leveling Charge Capture

Mapped capabilities

77 scenarios

  • CPT Code Suggestion from Ambient Note

06

Speech Recognition Diarization Transcription

Mapped capabilities

81 scenarios

  • Brand-name drug ASR accuracy

Frequently asked questions

What do the Corsac evals for Commure / Augmedix test?+

Each eval pack tests Commure / Augmedix's public product surface — including Ambient Note Generation Ai Drafting, Audio Capture Ambient Listening, and Ehr Integration Write Back Reconciliation — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Commure / Augmedix evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 395 Commure / Augmedix cases — from Speech Recognition Diarization Transcription (81 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Commure / Augmedix library.

How many test cases does the Commure / Augmedix library include?+

The Commure / Augmedix eval library includes 395 graded test cases across 6 eval packs, the largest being Speech Recognition Diarization Transcription with 81 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Commure / Augmedix or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Commure / Augmedix packs — Ambient Note Generation Ai Drafting and Audio Capture Ambient Listening and the rest — against Commure / Augmedix or your own agent with your own data.