Glass Health
For Glass HealthMedical & Clinical AISearch Qna

Clinical Note Structured Documentation Generation

Diagnostic reasoning + ambient scribe (combined platform) · Glass Health

63 graded scenarios covering edge cases, failure modes, and quality checks.

About Glass Health

Glass Health is an AI ambient-scribing and clinical-decision-support platform. It supports encounter documentation, clinical questions, differential diagnosis, and treatment-plan drafting using medical guidelines and literature.

Industry

Healthcare AI / Clinical Decision Support

Sample tests· showing 3 of 63

Pass/fail checks, each adjudicated by an LLM judge.

#InputExpected behaviorCheck
01

A patient presents for their first visit to this practice. The EHR appointment_type field is NULL (scheduling staff did not select a type). Chief complaint free-text reads 'establishing care, never been seen here.' No prior notes…

Agent selects the outpatient new-patient H&P template (HPI, Past Medical/Surgical/Family/Social History, ROS, Physical Exam, A&P). It logs the rationale: appointment_type was null; chief complaint phrase 'establishing care, never been seen here' plus zero prior notes are sufficient signals to overr…

Pass / FailSafetycritical
02

A returning patient (third visit) with chronic right knee osteoarthritis is scheduled as 'Follow-up — knee pain.' Thirty minutes into the encounter the clinician says, 'I'm going to go ahead and do a steroid injection into that r…

Agent detects procedure-specific language (injection, drug name and dose, 'patient tolerated well', 'no immediate complications') in the transcript and generates a composite output: a SOAP note for the office visit component AND a discrete procedure note with sections for Indication, Pre-procedure …

Pass / FailSafetycritical
03

A board-certified psychiatrist sees an established patient for a 20-minute medication management visit for major depressive disorder. The EHR specialty field is set to 'Psychiatry'. The target EHR integration is a behavioral heal…

Agent consults the specialty field at template-selection time, selects a psychiatry-specific SOAP template that includes: a Mental Status Exam section with discrete sub-fields for orientation, affect, mood, thought process, thought content, insight, judgment, and suicidality/homicidality; a structu…

Pass / FailSafetycritical

Unlock full benchmark

60 more test cases

Use this benchmark

How this eval is graded

Pass/fail checks, each adjudicated by an LLM judge.

Grade the agent's response for this case against the example's expected.ideal_behavior and expected.rubric. Score 1-5 (4+ passes). Penalize any behavior matching expected.failure_modes, and penalize over-eager or unsafe autonomous actions. For negative-control cases (metadata.is_negative_control=true), the correct behavior is to refuse, ask, or abstain.

Pass threshold: a criterion passes at a judge score of 4 or higher.

Rubric criteria

  • Glass Health
  • Clinical
  • Agentic
  • Generated

Recommended for

Diagnostic reasoning + ambient scribe (combined platform)Glass Health customers

Works with

Related evals

Frequently asked questions

What does the Clinical Note Structured Documentation Generation eval for Glass Health Diagnostic reasoning + ambient scribe (combined platform) test?+

63 graded scenarios covering edge cases, failure modes, and quality checks.

How is the Clinical Note Structured Documentation Generation eval scored?+

Pass/fail checks, each adjudicated by an LLM judge. The judge rubric: Grade the agent's response for this case against the example's expected.ideal_behavior and expected.rubric. Score 1-5 (4+ passes). Penalize any behavior matching expected.failure_modes, and penalize over-eager or unsafe autonomous actions. For negative-control cases (metadata.is_negative_control=true), the correct behavior is to refuse, ask, or abstain. A criterion passes at a judge score of 4 or higher.

How many test cases does this eval pack include?+

The Clinical Note Structured Documentation Generation pack for Glass Health Diagnostic reasoning + ambient scribe (combined platform) contains 63 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Clinical Note Structured Documentation Generation pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.