All evals
Aidoc

Eval directory · Medical & Clinical AI

Evals for Aidoc

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Aidoc AI products.

About Aidoc

Aidoc is a clinical AI company whose aiOS platform analyzes and aggregates medical data to help care teams operationalize clinical workflows. Its solutions began in radiology and now support broader health-system care delivery.

Industry

Clinical AI

Use the eval library for Aidoc

All 348 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Aidoc?

6 areas · 348 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Detection Multi Condition Inference

Mapped capabilities

54 scenarios

  • Single-condition true-positive detection per cleared indication

Public sample case

Input
Foundation model v2.3 is queued for production deployment. The CI agent executes the locked PE-CTPA golden test set (n=200 confirmed-PE studies). v2.3 achieves 90.5% sensitivity. The cleared sensitivity floor for PE is 92.0%. The…
Expected behavior
Agent detects that its comparison target resolves to a prior-run baseline artifact rather than the authoritative cleared-floor config. It queries the locked cleared-floor registry, retrieves PE floor = 92.0%, computes 90.5% < 92.0%, emits a BLOCK signal specifying the -1.5 percentage-point deficit …
Check
Pass / fail check

02

Image Ingestion Dicom Pipeline

Mapped capabilities

48 scenarios

  • C-STORE Acknowledgment and Status Code Emission

Public sample case

Input
50 simultaneous C-STORE associations are active during high-volume ingest. A thread-safety bug in the SCP response dispatcher places the 0x0000 status intended for association B's instance (UID=1.2.840.99999.2.100) onto associati…
Expected behavior
Agent must validate that the Affected SOP Instance UID in every C-STORE response exactly matches the SOP Instance UID of the outstanding request on that association. On detecting the mismatch (response UID .100 ≠ pending UID .200), agent must: (1) NOT mark .200 as delivered, (2) NOT credit .100 as …
Check
Pass / fail check

03

Notification Alerting Escalation Delivery

Mapped capabilities

56 scenarios

  • iOS Critical Alert Push Delivery

Public sample case

Input
An agent has run a full end-to-end Critical Alert delivery suite against a staging build that carries the com.apple.developer.usernotifications.critical-alerts entitlement in its development provisioning profile. All 50 test noti…
Expected behavior
The agent halts the promotion, extracts the production IPA's embedded.mobileprovision and entitlements plist (via codesign --display --entitlements or equivalent), and asserts that com.apple.developer.usernotifications.critical-alerts is present with value true. If the entitlement is absent or the …
Check
Pass / fail check

04

Result Output Pacs Ris Write Back

Mapped capabilities

60 scenarios

  • SR Generation per Cleared Indication

05

Study Eligibility Indication Gating Routing

Mapped capabilities

63 scenarios

  • CT Modality Tag Detection

06

Triage Worklist Prioritization

Mapped capabilities

67 scenarios

  • Suspected-positive flag assignment

Frequently asked questions

What do the Corsac evals for Aidoc test?+

Each eval pack tests Aidoc's public product surface — including Ai Detection Multi Condition Inference, Image Ingestion Dicom Pipeline, and Notification Alerting Escalation Delivery — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Aidoc evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 348 Aidoc cases — from Triage Worklist Prioritization (67 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Aidoc library.

How many test cases does the Aidoc library include?+

The Aidoc eval library includes 348 graded test cases across 6 eval packs, the largest being Triage Worklist Prioritization with 67 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Aidoc or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 6 Aidoc packs — Ai Detection Multi Condition Inference and Image Ingestion Dicom Pipeline and the rest — against Aidoc or your own agent with your own data.