All evals
Hippocratic AI

Eval directory · Medical & Clinical AI

Evals for Hippocratic AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Hippocratic AI AI products.

About Hippocratic AI

Hippocratic AI builds safety-focused AI agents for healthcare, focused on patient communication, navigation, and triage. Its models are trained with clinical oversight to ensure safe, empathetic interactions that complement clinical care rather than replace it.

Employees

~150

Industry

Healthcare AI

Headquarters

Palo Alto, CA

Use the eval library for Hippocratic AI

All 353 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Hippocratic AI?

6 areas · 353 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Clinical Safety Non Diagnostic Guardrails

Mapped capabilities

53 scenarios

  • Implicit Diagnosis Elicitation Refusal

02

Conversational Voice Core Asr Tts Turn Taking

Mapped capabilities

79 scenarios

  • Baseline Continuous-Speech ASR Accuracy

03

Polaris Constellation Orchestration

Mapped capabilities

64 scenarios

  • Single-domain routing — medication

04

Telephony Call Lifecycle

Mapped capabilities

51 scenarios

  • Outbound call initiation — schedule-triggered dialing

05

Tts Output Quality Drug Name Pronunciation

Mapped capabilities

58 scenarios

  • Brand-name drug pronunciation accuracy

06

Turn Taking Conversational Dynamics

Mapped capabilities

48 scenarios

  • Agent Response to Barge-In — Immediate Speech Halt

Frequently asked questions

What do the Corsac evals for Hippocratic AI test?+

Each eval pack tests Hippocratic AI's public product surface — including Clinical Safety Non Diagnostic Guardrails, Conversational Voice Core Asr Tts Turn Taking, Polaris Constellation Orchestration — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Hippocratic AI evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Hippocratic AI library include?+

The Hippocratic AI eval library includes 353 graded test cases across 6 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Hippocratic AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Hippocratic AI or your own agent with your own data.