All evals
OpenEvidence

Eval directory · Medical & Clinical AI

Evals for OpenEvidence

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for OpenEvidence AI products.

About OpenEvidence

OpenEvidence is a medical AI platform for clinicians that provides point-of-care answers grounded in peer-reviewed medical literature. Its mission is to organize and expand the world's collective medical knowledge.

Industry

Medical AI / Clinical Decision Support

Use the eval library for OpenEvidence

All 573 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for OpenEvidence?

10 areas · 573 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Authentication Session Account Lifecycle

Mapped capabilities

55 scenarios

  • NPI scan verification — happy path

Public sample case

Input
A newly registered attending physician submits a valid 10-digit Type 1 (individual) NPI during account setup. The NPPES registry returns an active record whose provider name fields match the name supplied at account creation with…
Expected behavior
The agent (1) validates the check digit using the NPI Luhn-variant algorithm before making any external call; (2) queries the NPPES NPI Registry API for NPI 1234567893; (3) confirms the returned record Entity Type Code is '1' (individual); (4) confirms the record status field is 'A' (active); (5) p…
Check
Pass / fail check

02

Citation Grounding Faithfulness

Mapped capabilities

55 scenarios

  • Claim-to-Citation Presence

Public sample case

Input
An agent is generating a draft vancomycin order for a 75 kg adult with suspected MRSA bacteremia. It queries OpenEvidence for weight-based dosing and AUC/MIC targets, then extracts numerical values to pre-populate order fields. I…
Expected behavior
Every discrete numerical dosing claim — loading dose in mg/kg, maintenance dose range, dosing interval, and target AUC/MIC range — is immediately followed by an inline citation to a specific authoritative source (e.g., a published vancomycin consensus guideline or FDA prescribing label [REQUIRES-VE…
Check
Pass / fail check

03

Clinical Question Answering Core Synthesis

Mapped capabilities

67 scenarios

  • Single-Fact Drug Dosing Accuracy

Public sample case

Input
An inpatient medicine resident is treating a 140 kg, 170 cm male patient (BMI ~48) with gram-negative bacteremia. The resident queries OpenEvidence to confirm extended-interval gentamicin dosing before calling pharmacy. The agent…
Expected behavior
The model states the standard extended-interval dose (5–7 mg/kg q24h) AND explicitly instructs that in obese patients (BMI ≥30) the dose must be calculated using adjusted body weight (ABW = IBW + 0.4 × [TBW − IBW]), not total actual body weight, providing or describing the ABW formula. It must stat…
Check
Pass / fail check

04

Clinician Identity Verification Access Gate

Mapped capabilities

54 scenarios

  • NPI optical/camera scan parsing

05

Drug Safety Pharmacovigilance

Mapped capabilities

42 scenarios

  • Known DDI pair retrieval

07

Medical Domain Coverage Correctness

Mapped capabilities

59 scenarios

  • Internal Medicine Core Condition Management

08

Natural Language Clinical Query Answering Core Q A

Mapped capabilities

62 scenarios

  • Clinical synonym resolution

09

Renal Hepatic Weight Based Dose Adjustment

Mapped capabilities

58 scenarios

  • eGFR Threshold Band Classification

10

Retrieval Pipeline Corpus Coverage

Mapped capabilities

50 scenarios

  • Rare-condition recall completeness

Frequently asked questions

What do the Corsac evals for OpenEvidence test?+

Each eval pack tests OpenEvidence's public product surface — including Authentication Session Account Lifecycle, Citation Grounding Faithfulness, and Clinical Question Answering Core Synthesis — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the OpenEvidence evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 573 OpenEvidence cases — from Evidence Retrieval Corpus Search (71 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the OpenEvidence library.

How many test cases does the OpenEvidence library include?+

The OpenEvidence eval library includes 573 graded test cases across 10 eval packs, the largest being Evidence Retrieval Corpus Search with 71 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against OpenEvidence or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 10 OpenEvidence packs — Authentication Session Account Lifecycle and Citation Grounding Faithfulness and the rest — against OpenEvidence or your own agent with your own data.