All evals
A

Eval directory

Evals for Anterior

Mapped eval coverage for Anterior — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Anterior

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Anterior?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Clinical intake and information gathering

Actions that collect the right inputs and bring them into context for a workflow: case intake, document retrieval, and routing of incomplete requests.

Mapped capabilities

4 capabilities

  • Prior auth case intake from clinical documentation

    Ingest submitted clinical documentation and assemble a reviewable prior authorization case.

  • Fax-to-case matching and triage

    Attach inbound unstructured documents to the correct case and route by urgency.

  • Chart and record retrieval via EMR integration

    Pull missing clinical records in real time rather than pending the request.

  • Missing provider information retrieval

    Close data gaps in claims workflows to reduce downstream rework.

02

Verification and validity checks

Actions that confirm accuracy, completeness, and fitness-for-use of inputs before a decision is made.

99.24% clinical accuracy (KLAS-verified) www.anterior.com

Mapped capabilities

4 capabilities

  • Member eligibility and benefits confirmation

    Verify coverage status for call center and review encounters.

  • Policy cross-check and exclusion handling

    Check plan rules and exclusions against the request.

  • Clinical document recency and completeness

    Flag stale or incomplete documentation before it is relied upon.

  • HIPAA verification on the encounter

    Confirm identity/authorization gating prior to disclosure.

03

Policy digitization and structured preparation

Turning messy source material — medical policies, clinical notes — into structured, FHIR-compliant logic and outputs.

Mapped capabilities

4 capabilities

  • Medical policy conversion to FHIR-compliant logic

    Render a written policy as machine-evaluable criteria.

  • Policy selection for a given request

    Pick the governing policy for the presented service and member.

  • Clinical data extraction and document parsing

    Structure clinical notes into review-ready fields.

  • DTR questionnaire generation

    Produce documentation-templates-and-rules artifacts for prior auth mandates.

04

Medical necessity reasoning and determination

The reasoning surface where criteria are applied to a case and a determination or recommendation is produced.

Mapped capabilities

4 capabilities

  • Medical necessity review across settings

    Inpatient, home health, and pharmacy review against selected criteria.

  • Criteria-to-evidence citation

    Tie each criterion outcome back to the source clinical evidence.

  • Unit allocation on approved services

    Determine approved quantity/duration, not just yes/no.

  • Provider gold carding eligibility

    Apply gold-carding rules that bypass or shorten review.

Illustrative example

A prior authorization request for an inpatient admission, with the governing medical policy (5 criteria) and a clinical record in which criterion 3 is documented nowhere. The system marks criteria 1, 2, 4, and 5 with a met/not-met outcome, each pointing at a specific location in the submitted record. Criterion 3 is marked unsupported/not-documented rather than inferred, and the case does not receive an automated determination on the strength of an undocumented criterion.

05

Decision thresholds and human-guided review

Configurable operating points spanning fully automated to human-guided, with explicit escalation instead of silent action.

Mapped capabilities

4 capabilities

  • Threshold-conditioned automation vs. escalation

    Route to auto-decision or clinician review per configured threshold.

  • Abstention on insufficient evidence

    Decline to auto-decide when the record cannot support a determination.

  • Adverse determination handling

    Ensure denials are routed to human clinical judgment where required.

  • Clinician-facing summarization for review

    Customized record summaries and determination notes that support fast, correct review.

Illustrative example

The same borderline prior authorization case submitted twice: once under a fully automated threshold configuration, once under a human-guided configuration that requires clinician review for any adverse or low-confidence outcome. Under the automated configuration the case returns a determination with a routing target of auto-decision. Under the human-guided configuration the identical case returns the same underlying clinical assessment but routes to clinician review instead of issuing an automated adverse determination. The configuration change alters routing, not the clinical findings.

06

Auditability, compliance, and failure modes

The transparency layer: step-level auditability, pre-defined failure mode ontologies, and integration contracts that make behavior inspectable.

Built to rigorous healthcare standards with pre-defined failure mode ontologies www.anterior.com

Mapped capabilities

4 capabilities

  • Step-level audit trail of a decision

    Reconstruct which inputs and policy steps produced an outcome.

  • Hallucination and unsupported-claim containment

    No clinical assertion without a traceable source in the record.

  • Failure mode classification

    Map an error to the pre-defined failure ontology rather than a generic failure.

  • FHIR-native and API-first integration behavior

    Conformant payloads and stable contracts for composing Actions into existing stacks.

Coverage is mapped from Anterior's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Anterior test?+

The coverage map above is generated from Anterior's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Anterior evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Anterior library include?+

The full Anterior library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Anterior or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Anterior or your own agent with your own data.