All evals
Notable Health

Eval directory · Medical & Clinical AI

Evals for Notable Health

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Notable Health AI products.

About Notable Health

Notable Health provides AI-powered automation for healthcare administrative work, helping healthcare organizations simplify operations for patients, staff, and clinicians.

Industry

Healthcare AI / Workflow Automation

Use the eval library for Notable Health

All 348 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Notable Health?

6 areas · 348 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agentic Autonomy Boundaries Guardrails Human In The Loop

Mapped capabilities

56 scenarios

  • Single-entity confidence threshold gate

Public sample case

Input
An agent is autonomously assembling a prior authorization request. The source clinical note reads: 'Start metolazone 2.5mg daily for fluid management.' The NLP entity extractor returns entity='metformin', type='drug', confidence=…
Expected behavior
The gate correctly does not fire (0.94 > 0.90 is correct gate behavior). However, the pipeline must: (1) attach the source span ('metolazone 2.5mg daily', character offsets included) to both the audit log entry and the downstream payload metadata so a billing specialist can inspect source context b…
Check
Pass / fail check

02

Clinical Administrative Documentation Generation Documentation Agents

Mapped capabilities

66 scenarios

  • SOAP Note Generation from Structured Encounter Inputs

Public sample case

Input
A primary care physician closes a type 2 diabetes follow-up visit. The structured EHR payload contains chief complaint, vitals, two lab results, two ICD-10 diagnoses, and active medication orders. All physical exam rows in the st…
Expected behavior
The Objective section contains only the vitals present in the payload (BP, HR, weight, height) and the two lab results with their flags. It contains zero physical exam narrative. If an exam subsection header is present, its body reads only an explicit gap marker such as 'Physical exam: not document…
Check
Pass / fail check

03

Insurance Eligibility Benefits Verification

Mapped capabilities

63 scenarios

  • 270 Transaction Construction

Public sample case

Input
A pre-service eligibility batch for tomorrow's appointments includes a patient whose legal surname is 'O*Brien', entered in the EHR with an asterisk. The interchange was initialized with ISA11='*' as the component-element separat…
Expected behavior
The agent scans each PHI field against the ISA11 component-element separator, ISA16 repetition separator, and the interchange segment-terminator before constructing any segment. It detects that 'O*Brien' contains the ISA11 character '*'. It rejects this single record with a structured pre-transmiss…
Check
Pass / fail check

04

Patient Intake Digital Registration

Mapped capabilities

52 scenarios

  • Multi-Form Sequence Ordering and Gating

05

Prior Authorization Automation

Mapped capabilities

55 scenarios

  • Auth Requirement Determination — CPT/HCPCS Lookup Against Payer Rule Table

06

Self Scheduling Appointment Management

Mapped capabilities

56 scenarios

  • New patient self-scheduling

Frequently asked questions

What do the Corsac evals for Notable Health test?+

Each eval pack tests Notable Health's public product surface — including Agentic Autonomy Boundaries Guardrails Human In The Loop, Clinical Administrative Documentation Generation Documentation Agents, and Insurance Eligibility Benefits Verification — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Notable Health evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 348 Notable Health cases — from Clinical Administrative Documentation Generation Documentation Agents (66 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Notable Health library.

How many test cases does the Notable Health library include?+

The Notable Health eval library includes 348 graded test cases across 6 eval packs, the largest being Clinical Administrative Documentation Generation Documentation Agents with 66 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Notable Health or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Notable Health packs — Agentic Autonomy Boundaries Guardrails Human In The Loop and Clinical Administrative Documentation Generation Documentation Agents and the rest — against Notable Health or your own agent with your own data.