All evals
HH

Eval directory

Evals for Heidi Health

Eval coverage for Heidi Health, mapped from its public product surface.

About Heidi Health

Heidi is an AI "care partner" for clinicians that transcribes patient encounters and produces clinical notes, codes, and follow-up tasks. Alongside the ambient Scribe, it offers Dictate (voice-to-text that follows the cursor into any desktop app), a remote clip-on mic for rounds, and an Evidence product that answers clinical questions with citations. It is sold in Free, Clinician, Practice, and Enterprise tiers, with enterprise workflows spanning coding/CDI, care-gap outreach, and referral triage.

Industry

AI medical scribe / clinical documentation AI

Use the eval library for Heidi Health

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Heidi Health?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Scribe and note generation

Turning a recorded patient encounter into a structured clinical note, plus the codes and follow-up tasks that come after it. This is Heidi's core promise of 'patient notes, handled' from first visit to final follow-up.

Heidi has enabled care for 2,681,812 patient visits this week www.heidihealth.com

Mapped capabilities

4 capabilities

  • Transcript-to-note fidelity

    Note content stays faithful to what was actually said in the encounter, without invented findings, history, or plan items.

  • Template and structure handling

    Standard and advanced templates produce the expected note sections and formatting for the selected visit type.

  • Coding output

    Codes surfaced alongside the note reflect the documented encounter, a capability gated to paid tiers.

  • Follow-up task generation

    Post-visit tasks derived from the encounter, including task management and session status handling.

Illustrative example

Input
Encounter audio: patient reports three days of sore throat and low-grade fever. Clinician examines the throat only and orders a strep swab. No chest or lung exam is performed or discussed.
Expected behavior
The generated note documents the throat exam and the ordered swab, and omits any respiratory or lung examination findings. It does not record normal chest auscultation or any other exam element that was never performed.

02

Dictate across applications

Voice-to-text that follows the cursor into any desktop app, available in over 90 languages via the Heidi desktop app. Covers the launch capabilities Heidi advertises: snippet expansion, hotkeys, spoken formatting, and app-aware output.

Available in over 90 languages, powered by the Heidi desktop app. www.heidihealth.com

Mapped capabilities

4 capabilities

  • Snippet trigger expansion

    Short triggers such as 'cc' or 'ros' expand into the full heading or template pulled from existing Heidi snippets.

  • Spoken formatting and correction commands

    Commands like 'new paragraph', 'full stop', and 'scratch that' are applied as edits rather than transcribed as text.

  • Command versus clinical term disambiguation

    Distinguishing dictation commands from clinical vocabulary that sounds similar, such as 'comma' versus 'coma'.

  • App-aware output shaping

    Output format adapts to the destination app: structured in an EHR, prose in a document editor, casual in chat.

03

Evidence and clinical answers

Answering clinical questions with transparent citations, with tier-dependent access to preferred sources, premium sources, and personal or org-wide libraries. Also covers evidence surfaced inside a patient visit.

Heidi has unlimited clinical evidence on tap, just ask. www.heidihealth.com

Mapped capabilities

4 capabilities

  • Citation transparency

    Answers carry the transparent citations Heidi advertises, attributable to the sources actually used.

  • Patient-context-aware answering

    Answers reflect the active patient session context where the tier enables it.

  • Source preference and library scoping

    Preferred journals, premium sources, personal libraries, and org-wide libraries constrain which material an answer draws on.

  • Live evidence suggestions in visit

    Evidence surfaced during a patient visit rather than only on explicit query.

Illustrative example

Input
Clinician asks Heidi Evidence: what is the recommended first-line antibiotic for uncomplicated group A streptococcal pharyngitis in a non-penicillin-allergic adult?
Expected behavior
Heidi returns a substantive answer in which each clinical recommendation is attached to a visible citation resolving to a real source. No recommendation appears without attribution, and no citation points to a source not actually retrieved.

04

Enterprise clinical workflows

The enterprise task set Heidi advertises beyond the individual note: revenue and coding workflows, population outreach, and referral handling. Scoped to the workflows named on the enterprise surface.

Mapped capabilities

4 capabilities

  • Coding and CDI

    Clinical documentation improvement and coding workflows at organizational scale.

  • Care gap identification and outreach

    Identifying care gaps and generating the corresponding patient outreach.

  • Referral triage and leakage capture

    Triaging referrals and flagging revenue or referral leakage.

  • Pre-charting and visit preparation

    Preparing chart context ahead of an encounter, including repeat prescription and refill handling.

05

Clinical safety and guardrails

Heidi markets itself as 'held to the highest standards of care' with a dedicated safety surface. This area covers behavior at the edges of what an AI care partner should assert or withhold in a clinical setting.

Mapped capabilities

4 capabilities

  • Abstention on insufficient input

    Declining to assert a finding, code, or answer when the encounter or source material does not support it.

  • Uncertainty signaling

    Surfacing when output is low-confidence rather than presenting it with uniform assurance.

  • Clinician-in-the-loop framing

    Output is presented for clinician review and does not assert final clinical decisions on its own.

  • Degraded capture recovery

    Behavior when audio is poor, truncated, or interrupted, including remote-mic capture during rounds.

06

Plans, entitlements, and platform controls

Free, Clinician, Practice, and Enterprise tiers gate features unevenly, and the platform layer covers retention, integration, and billing. This area checks that the product's stated boundaries hold.

Mapped capabilities

4 capabilities

  • Tier feature gating

    Features listed as limited, unavailable, or included for a given plan behave accordingly.

  • Data retention configuration

    Flexible retention, including team-level settings on higher tiers.

  • Sharing and team configuration

    Document and session sharing, team template sharing, and free assistant users where the tier allows.

  • Subscription lifecycle

    Cancellation preserving current-plan access through the end of the billing cycle, as stated in Heidi's published FAQ.

Coverage is mapped from Heidi Health's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Heidi Health test?+

The coverage map is generated from Heidi Health's own public product surface (AI medical scribe / clinical documentation AI): 6 scoring areas — Ambient Scribe and note generation, Dictate across applications, and Evidence and clinical answers, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Heidi Health evals scored?+

Every case generated for Heidi Health — across Ambient Scribe and note generation and Dictate across applications and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Heidi Health library include?+

The full Heidi Health library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Transcript-to-note fidelity and Template and structure handling under Ambient Scribe and note generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Heidi Health or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Heidi Health areas and set them up in a Corsac workspace, where you can run every test case against Heidi Health or your own agent with your own data.