All evals
Bayesian Health

Eval directory · Medical & Clinical AI

Evals for Bayesian Health

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Bayesian Health AI products.

About Bayesian Health

Bayesian Health provides a real-time clinical intelligence platform that continuously monitors patient data, surfaces patients needing attention, and gives clinicians guidance for proactive care.

Industry

Clinical AI / Predictive Analytics

Use the eval library for Bayesian Health

All 359 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Bayesian Health?

6 areas · 359 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Alerting Routing Escalation

Mapped capabilities

65 scenarios

  • Threshold Crossing Detection

Public sample case

Input
Patient PT-4421 had a prior upward crossing at T=09:00 (score 0.63→0.75, threshold 0.70). That alert was acknowledged by the floor nurse. At T=09:22 a new score of 0.62 arrives — a large downward crossing back below threshold. Th…
Expected behavior
Agent identifies the transition as a downward crossing (0.75→0.62, crossing 0.70 from above to below), takes no new alert action, logs a recovery/resolution event for audit purposes only, and does NOT dispatch a rapid-response team, page an attending, or create any new alert entry in the queue.
Check
Pass / fail check

02

Clinical Feature Engineering Data Quality

Mapped capabilities

60 scenarios

  • HR window aggregation derivation

Public sample case

Input
A code review surfaces that the 1-hour HR window is computed as [T−60min, T) (exclusive upper bound) in the real-time inference module but as [T−60min, T] (inclusive upper bound) in the legacy batch-scoring path. The training dat…
Expected behavior
The agent (1) identifies that the training dataset was built with the inclusive boundary [T-60min, T], meaning the model was trained on that distribution; (2) explains that flipping the real-time path to inclusive would match the batch convention but model revalidation is still required because agg…
Check
Pass / fail check

03

Ehr Data Ingestion Interoperability

Mapped capabilities

56 scenarios

  • HL7v2 ADT Message Parsing

Public sample case

Input
An A01 admit message arrives from an Epic installation. PID-3 contains three CX repetitions in this order: a Social Security Number (AA='SSN'), an enterprise ID (AA='ENTERPRISE'), and the site MRN (AA='EPIC-MRN'). The facility's …
Expected behavior
Parser extracts MRN '001234567' with assigning authority 'EPIC-MRN' as the patient's primary census key. The SSN ('987654321') and enterprise ID ('88776655') are stored as secondary identifiers only and are never used as the lookup key for census operations, alert routing, or model enrollment. The …
Check
Pass / fail check

04

Model Lifecycle Governance

Mapped capabilities

78 scenarios

  • Version identifier assignment and uniqueness enforcement

05

Model Performance Monitoring Drift Detection

Mapped capabilities

46 scenarios

  • Rolling AUPRC Computation Correctness

06

Real Time Prediction Risk Scoring Engine

Mapped capabilities

54 scenarios

  • Periodic-cadence scoring

Frequently asked questions

What do the Corsac evals for Bayesian Health test?+

Each eval pack tests Bayesian Health's public product surface — including Alerting Routing Escalation, Clinical Feature Engineering Data Quality, and Ehr Data Ingestion Interoperability — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Bayesian Health evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 359 Bayesian Health cases — from Model Lifecycle Governance (78 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Bayesian Health library.

How many test cases does the Bayesian Health library include?+

The Bayesian Health eval library includes 359 graded test cases across 6 eval packs, the largest being Model Lifecycle Governance with 78 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Bayesian Health or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Bayesian Health packs — Alerting Routing Escalation and Clinical Feature Engineering Data Quality and the rest — against Bayesian Health or your own agent with your own data.