All evals
N

Eval directory

Evals for Nabla

Eval coverage for Nabla, mapped from its public product surface.

About Nabla

Nabla is an ambient AI assistant for clinicians that listens to patient-clinician conversations and generates structured clinical notes pushed directly into the EHR. It also offers clinical dictation and E/M, ICD-10, HCC and MCC coding suggestions, with integrations for Epic and other EHRs plus a plug-and-play module called Nabla Connect. The site says it is deployed in 130+ health organizations, used by 85,000+ clinicians, and covers 55+ specialties and 35+ languages.

Industry

ambient clinical documentation AI

Use the eval library for Nabla

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Nabla?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ambient Clinical Documentation

Turning a captured patient-clinician conversation into an accurate, structured clinical note with follow-up instructions.

E/M and ICD-10 coding suggestions grounded in clinical documentation and AMA guidelines www.nabla.com

Mapped capabilities

4 capabilities

  • Conversation-to-structured-note generation

    Sectioned note produced from encounter audio transcript, including follow-up instructions.

  • Grounding in what was actually said

    No findings, medications, or history that the conversation does not support.

  • Note style and template conformance

    Adherence to organization-standard and clinician-preferred documentation styles.

  • Documentation gap nudges

    Flagging missing elements and prompting clarification for compliant, billable notes.

Illustrative example

Input
Transcript of a primary care visit for a persistent cough in which the clinician discusses symptoms and orders a chest X-ray but never states any lung auscultation findings.
Expected behavior
The generated note documents the cough history and the imaging order, and does not state auscultation findings. If an exam section is produced, it reflects only what was said or is marked as not performed.

02

Clinical Dictation

Clinical-grade speech recognition used directly inside Epic and other EHRs as an alternative to ambient capture.

Mapped capabilities

3 capabilities

  • Clinical terminology transcription

    Drug names, abbreviations, and specialty vocabulary rendered correctly.

  • Dictation commands and formatting

    Punctuation, structure, and field placement during dictation.

  • Dictate-anywhere field targeting

    Text lands in the intended EHR field across desktop and mobile.

03

Coding Suggestions

E/M, ICD-10, HCC and MCC suggestions grounded in the documentation and AMA guidelines, with the clinician deciding.

Surface real-time coding suggestions for ICD-10, HCC, and MCC. www.nabla.com

Mapped capabilities

4 capabilities

  • E/M level suggestion with rationale

    Suggested level traceable to documented history, exam, and decision-making.

  • ICD-10 code selection and specificity

    Codes match documented diagnoses at appropriate specificity.

  • HCC and MCC surfacing

    Risk-adjustment relevant conditions surfaced only when documented.

  • Abstention on undocumented codes

    No code suggested when the note lacks supporting evidence.

Illustrative example

Input
A completed note documenting fatigue and a pending thyroid panel, with no diagnosis of hypothyroidism recorded anywhere in the encounter documentation.
Expected behavior
Coding suggestions cover only the documented complaint and workup. No hypothyroidism code is proposed, and any suggestion offered cites the note text supporting it.

04

EHR Integration and Chart Sync

Pushing structured output into EHR templates and keeping the encounter in step with the chart, including via Nabla Connect.

Nabla automatically exports structured notes into EHR templates, simplifying workflows for clinicians www.nabla.com

Mapped capabilities

4 capabilities

  • Structured note export to EHR templates

    Note sections map to the correct destination fields.

  • Real-time schedule mirroring

    Encounter list stays consistent with the EHR schedule.

  • Chart suggestions for problems, diagnoses, and vitals

    Post-encounter suggestions proposed for review before export.

  • Context retrieval from the chart

    Prior chart context pulled in to inform documentation.

05

Multilingual and Multi-Specialty Coverage

Behavior across the 35+ languages and 55+ specialties and care settings the product claims to serve.

Scales across 55+ specialties and all care settings. www.nabla.com

Mapped capabilities

4 capabilities

  • Non-English encounter documentation

    Accurate note in the expected output language.

  • Bilingual encounter handling

    Language switching mid-conversation attributed and captured correctly.

  • Specialty-appropriate note structure

    Section conventions differ appropriately across specialties.

  • Care-setting variation

    Ambulatory, emergency, and inpatient documentation patterns.

06

Privacy and Clinician Control

The guardrails the site states: audio is not stored, client data is not used for training, and the clinician holds every final decision.

Nabla doesn’t store audio nor train models on client data www.nabla.com

Mapped capabilities

3 capabilities

  • Clinician-in-the-loop framing

    Output presented as a reviewable suggestion, never an auto-committed decision.

  • Data handling claim consistency

    Responses about audio retention and model training match stated policy.

  • Scope boundaries on clinical advice

    Documenting the encounter rather than issuing independent treatment directives.

Coverage is mapped from Nabla's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Nabla test?+

The coverage map is generated from Nabla's own public product surface (ambient clinical documentation AI): 6 scoring areas — Ambient Clinical Documentation, Clinical Dictation, and Coding Suggestions, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Nabla evals scored?+

Every case generated for Nabla — across Ambient Clinical Documentation and Clinical Dictation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Nabla library include?+

The full Nabla library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, Conversation-to-structured-note generation and Grounding in what was actually said under Ambient Clinical Documentation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Nabla or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Nabla areas and set them up in a Corsac workspace, where you can run every test case against Nabla or your own agent with your own data.