All evals
S

Eval directory

Evals for Sunoh.ai

Eval coverage for Sunoh.ai, mapped from its public product surface.

About Sunoh.ai

Sunoh.ai is an ambient-listening AI medical scribe that records patient-provider conversations and converts them into transcripts, draft clinical summaries, and draft SOAP/Progress Notes for review and import into an EHR. It also captures order-entry details such as labs, imaging, procedures, and medications. It runs on iPhone, iPad, Android, and web browsers, priced from $149 per user per month.

Industry

ambient AI medical scribe / clinical documentation

Website

sunoh.ai

Use the eval library for Sunoh.ai

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Sunoh.ai?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient capture and transcription

Turning a live patient-provider conversation into an accurate, speaker-attributed transcript, which is the input every downstream artifact depends on.

transforming real dialogue into accurate clinical notes in seconds sunoh.ai

Mapped capabilities

4 capabilities

  • Speaker attribution across the encounter

    Distinguishing provider, patient, and third parties in the room across turn-taking and interruptions.

  • Clinical terminology fidelity

    Medication names, dosages, anatomy, and abbreviations rendered correctly rather than phonetically approximated.

  • Conversational noise handling

    Small talk, self-corrections, and off-topic dialogue kept out of clinical content.

  • Transcript-to-dialogue-flow structure

    Producing the ordered dialogue flow described as the intermediate artifact before note generation.

02

Draft clinical note generation

Converting the transcript into a draft clinical summary and draft SOAP/Progress Note with content sorted into the correct sections.

Mapped capabilities

4 capabilities

  • Section routing (S/O/A/P)

    Subjective, objective, assessment, and plan content placed in the section it belongs to.

  • Grounding to the conversation

    No findings, history, or plan items that were never spoken in the encounter.

  • Negation and uncertainty preservation

    Denied symptoms, ruled-out conditions, and hedged provider language retained rather than flattened into assertions.

  • Draft clinical summary

    The short review-oriented summary distinct from the full note.

Illustrative example

Input
Encounter audio in which the provider asks about chest pain and the patient answers, "No chest pain at all, but my left knee has been aching for about three weeks."
Expected behavior
The draft note records the denial of chest pain as a negative finding in the subjective section and documents the three-week left knee pain as the presenting complaint. Chest pain does not appear as a positive symptom, assessment, or plan item.

03

Order-entry capture

Extracting labs, imaging, procedures, medication orders, and follow-up details from the spoken encounter as structured order candidates.

Captures labs, imaging, procedures, medication orders, and follow-up details effortlessly. sunoh.ai

Mapped capabilities

4 capabilities

  • Order type classification

    Correctly separating labs, imaging, procedures, and medications.

  • Medication order detail

    Drug, dose, route, and frequency captured together rather than partially.

  • Discussed-but-not-ordered discrimination

    Options the provider raised and rejected not surfaced as orders.

  • Follow-up and return-visit capture

    Interval, modality, and conditions for the next encounter.

Illustrative example

Input
Provider says, "We could do an MRI, but let's hold off — order a knee X-ray, three views, and a CBC today."
Expected behavior
The captured orders include a three-view knee X-ray under imaging and a CBC under labs. The MRI is not present as an order, though it may appear in the note's plan narrative as a discussed option.

04

Provider review and EHR handoff

The edit-then-import workflow: reviewing, modifying, and moving the finished note into the EHR or exporting it.

Mapped capabilities

4 capabilities

  • Pre-import editing

    Text edits preserved through to what is imported.

  • Review affordances

    Draft content presented so a provider can verify it against the encounter before accepting.

  • PDF export of the completed note

    Export fidelity of the finalized note.

  • Import integrity

    Note and captured orders arriving in the EHR without silent content loss.

05

Cross-device and session delivery

Consistency of the scribe experience across iPhone, iPad, Android, and web browsers, including how a recording session behaves.

Mapped capabilities

3 capabilities

  • Platform parity

    Same capture and review capability across the four supported surfaces.

  • Session start, pause, and stop

    Provider control over what is and is not being recorded.

  • Interrupted-session recovery

    Behavior when a visit recording is cut short or the app is backgrounded.

06

Security, privacy, and compliance posture

How the product represents PHI handling and its compliance commitments — BAA included, data-secure storage, hosted on Microsoft Azure, with 24x7x365 support and implementation services.

Runs on Microsoft® Azure® sunoh.ai

Mapped capabilities

4 capabilities

  • PHI handling representations

    Claims about storage and protection of recorded encounter data stated accurately.

  • BAA and hosting claims

    BAA inclusion and Azure hosting described consistently with published materials.

  • Consent and recording disclosure

    How the presence of ambient recording is surfaced in the encounter workflow.

  • Pricing and commitment claims

    The $149 per user per month figure and no-long-term-commitment framing represented without drift.

Coverage is mapped from Sunoh.ai's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sunoh.ai test?+

The coverage map is generated from Sunoh.ai's own public product surface (ambient AI medical scribe / clinical documentation): 6 scoring areas — Ambient capture and transcription, Draft clinical note generation, and Order-entry capture, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sunoh.ai evals scored?+

Every case generated for Sunoh.ai — across Ambient capture and transcription and Draft clinical note generation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sunoh.ai library include?+

The full Sunoh.ai library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Speaker attribution across the encounter and Clinical terminology fidelity under Ambient capture and transcription); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sunoh.ai or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Sunoh.ai areas and set them up in a Corsac workspace, where you can run every test case against Sunoh.ai or your own agent with your own data.