All evals
K Health

Eval directory

Evals for K Health

6 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for K Health AI products.

Medical & Clinical AI
Use evals for K Health

About K Health

K Health is a clinical AI platform for health systems that combines clinical AI, patient and provider platforms, and a virtual clinical workforce to support integrated primary care.

Industry

Clinical AI / Virtual Primary Care

Headquarters

New York City, NY

80/ 100
BStrong coverage

How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.

Test datasets

6/6 packs

Scoring metrics

6/6 packs

Judge rubrics

6/6 packs

Use-case maps

0/6 packs

Pack context

6/6 packs

Test datasetsStrong100%
Scoring metricsStrong100%
Judge rubricsStrong100%
Use-case mapsLimited0%
Pack contextStrong100%

Available eval packs for K Health

6 packs ready to run.

Why eval K Health AI

K Health's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.

The Corsac eval library for K Health measures four dimensions teams care about most when deploying medical & clinical ai agents:

  • Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
  • Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
  • Safety gates — does it escalate or refuse when it should, and only then?
  • Operator quality — does it preserve analyst trust by surfacing the right context at the right time?

Every eval pack above is hand-authored against K Health's public product surface and runnable in Corsac with your own data.