All evals
OS

Eval directory · Medical & Clinical AI

Evals for Oncology Starter

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Oncology Starter AI products.

About Oncology Starter

Oncology Starter is a Corsac starter benchmark for evidence-based clinical-trial recommendations, with tests for trial-status verification, eligibility completeness, and citation fidelity. It is a library template, not a company.

Industry

Clinical Research Benchmark

Use the eval library for Oncology Starter

All 80 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Oncology Starter?

4 areas · 80 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Oncology Curbside Trials V1

Starter eval pack for evidence-based clinical trial recommendations, emphasizing trial status verification, completeness of eligibility variables, and citation fidelity.

Mapped capabilities

30 scenarios

  • Trial Eligibility Accuracy
  • Citation Fidelity
  • Recommendation Completeness

Public sample case

Input
Produce output JSON that strictly matches the output schema in this suite. Return ONLY valid JSON.
Expected behavior
{"extraction":{"citations":[{"note":"Trial registry/program page used as source of truth.","source_id":"CTG_TAPUR_NCT02693535"},{"note":"Trial registry/program page used as source…
Check
Pass / fail check

Example criterion: The trial-matching agent returns accurate, evidence-grounded recommendations with complete eligibility rationale for clinicians.

02

Oncology Intake Triage V1

Starter eval pack for oncology intake triage with emphasis on oncologic emergencies and escalation safety.

Mapped capabilities

40 scenarios

  • Safety Triage Recall
  • Structured Intake Quality
  • Clinical Actionability

Public sample case

Input
Produce output JSON that strictly matches the output schema in this suite. Return ONLY valid JSON.
Expected behavior
{"extraction":{"citations":[{"note":"Supports urgency/escalation guidance for this scenario.","source_id":"NICE_QS56_MSCC_MRI24H"}],"immediate_actions":["Immediate ED evaluation /…
Check
Pass / fail check

Example criterion: The oncology intake agent safely triages high-risk cases, captures complete structured context, and provides clinically actionable next steps.

03

Oncology Longitudinal Episodes V1

Small longitudinal eval pack (multi-turn) that tests stability, safe escalation, and de-escalation with safety-netting.

Mapped capabilities

5 scenarios

  • Cross-Visit State Consistency
  • Update Correctness
  • Care Narrative Integrity

Public sample case

Input
Given this multi-turn clinical conversation, produce the appropriate triage/recommendation outputs for each turn. Return JSON with turn_outputs array matching the output schema. Return ONLY valid JSON.
Expected behavior
{"extraction":{"turn_outputs":[{"citations":[{"note":"Supports urgency/escalation guidance for this scenario.","source_id":"NICE_QS56_MSCC_MRI24H"}],"immediate_actions":["Immediat…
Check
Pass / fail check

Example criterion: The longitudinal agent preserves accurate cross-visit state, applies updates correctly, and maintains a trustworthy care narrative.

04

Oncology Precision Variants V1

Small eval pack for interpreting common actionable variants with citation fidelity to public variant knowledgebases.

Mapped capabilities

5 scenarios

  • Variant Actionability Classification
  • Evidence Attribution
  • Clinical Safety Boundaries

Example criterion: The precision-variants agent classifies actionable findings correctly, cites evidence faithfully, and stays within clinical safety boundaries.

Frequently asked questions

What do the Corsac evals for Oncology Starter test?+

Each eval pack tests Oncology Starter's public product surface — including Oncology Curbside Trials V1, Oncology Intake Triage V1, and Oncology Longitudinal Episodes V1 — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Oncology Starter evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 80 Oncology Starter cases — from Oncology Intake Triage V1 (40 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Oncology Starter library.

How many test cases does the Oncology Starter library include?+

The Oncology Starter eval library includes 80 graded test cases across 4 eval packs, the largest being Oncology Intake Triage V1 with 40 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Oncology Starter or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 4 Oncology Starter packs — Oncology Curbside Trials V1 and Oncology Intake Triage V1 and the rest — against Oncology Starter or your own agent with your own data.