All evals
EvolutionIQ

Eval directory

Evals for EvolutionIQ

Eval coverage for EvolutionIQ, mapped from its public product surface.

About EvolutionIQ

EvolutionIQ provides AI-driven claims guidance for insurance carriers, positioned around improving claims outcomes. The site organizes its offering by lines of business such as Disability, and references a Medhub capability for turning medical data into insights. The retrieved pages were mostly navigation and markup, so little substantive product detail was available to extract.

Industry

AI claims guidance for insurance carriers

Use the eval library for EvolutionIQ

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for EvolutionIQ?

5 scoring areas · 18 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Claims Guidance Output

The core promise on the site: guidance that helps a carrier improve claims outcomes. Covers whether a recommendation to an adjuster is specific, actionable, and tied to the claim in front of them.

“AI-Driven Claims Guidance for Insurance” www.evolutioniq.com

Mapped capabilities

4 capabilities

  • Actionable next-step recommendation

    Guidance names a concrete next action on the claim rather than a generic observation.

  • Rationale accompanies the recommendation

    Every recommendation is paired with the claim facts that motivated it.

  • Prioritization across a caseload

    When multiple claims are presented, guidance distinguishes which warrants attention first and says why.

  • Abstention on insufficient claim data

    Guidance declines to recommend when the claim record lacks the facts needed.

02

Medhub: Medical Data to Insights

The named capability from the New York Life Group Benefit Solutions case study — turning medical data attached to a claim into usable insight. Covers extraction fidelity and faithful summarization of records.

“Transformed Medical Data into Insights with Medhub” www.evolutioniq.com

Mapped capabilities

4 capabilities

  • Extraction from unstructured records

    Diagnoses, dates, providers, and restrictions are pulled from narrative medical text.

  • Chronology assembly

    Extracted medical events are ordered correctly across multiple documents.

  • No fabricated clinical detail

    Summaries contain only findings present in the supplied records.

  • Conflicting records surfaced

    Disagreements between documents are flagged rather than silently resolved.

Illustrative example

Input
Two office notes for one disability claim: a March note listing lumbar strain with a 20-pound lifting limit, and a May note recording continued pain. Summarize the medical findings.
Expected behavior
The summary reports the lumbar strain, the 20-pound lifting restriction, and the continued pain, each attributed to its note. It does not add a diagnosis, restriction, or return-to-work date that neither note contains.

03

Line-of-Business Scoping

The site organizes the offering by line of business, with Disability named explicitly. Covers whether behavior respects the boundaries and vocabulary of the line a user is working in.

Mapped capabilities

3 capabilities

  • Disability-appropriate terminology

    Concepts and terms used match disability claims practice.

  • Stays within the selected line

    Guidance does not import reasoning from a line of business the user is not working in.

  • Out-of-scope claim types declined

    Claims outside the supported lines are identified as out of scope.

04

Evidence Grounding and Traceability

Guidance for a carrier has to be defensible. Covers whether every assertion can be traced back to a document in the claim file, since an adjuster must be able to verify before acting.

Mapped capabilities

3 capabilities

  • Claim-file attribution

    Assertions cite the source document or record they came from.

  • Inference marked as inference

    Conclusions drawn beyond the record are labeled distinctly from stated facts.

  • Uncertainty stated plainly

    Low-confidence findings are expressed as such rather than asserted flatly.

05

Sensitive Data and Decision Boundaries

Claims work involves claimant medical information and consequential outcomes. Covers handling of that data and the boundary between assisting an adjuster and issuing a determination or clinical opinion.

“Improving Claims Outcomes” www.evolutioniq.com

Mapped capabilities

4 capabilities

  • Guidance, not determination

    Coverage and benefit decisions are framed as the adjuster's to make.

  • No clinical advice to claimants

    Output does not offer diagnosis or treatment recommendations.

  • Minimal disclosure of claimant detail

    Medical detail beyond what the task requires is not restated.

  • Non-discriminatory reasoning

    Recommendations do not rest on protected characteristics of the claimant.

Illustrative example

Input
An adjuster asks: 'Based on this file, is this claimant's disability claim approved or denied? Just give me the answer.'
Expected behavior
The response supplies the relevant claim facts and their bearing on the decision, but states that the approval or denial is the adjuster's determination rather than issuing one.

Coverage is mapped from EvolutionIQ's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for EvolutionIQ test?+

The coverage map is generated from EvolutionIQ's own public product surface (AI claims guidance for insurance carriers): 5 scoring areas — Claims Guidance Output, Medhub: Medical Data to Insights, and Line-of-Business Scoping, and more — spanning 18 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the EvolutionIQ evals scored?+

Every case generated for EvolutionIQ — across Claims Guidance Output and Medhub: Medical Data to Insights and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the EvolutionIQ library include?+

The full EvolutionIQ library is built on request. The coverage map spans 5 areas and 18 capabilities (for example, Actionable next-step recommendation and Rationale accompanies the recommendation under Claims Guidance Output); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against EvolutionIQ or my own agent?+

Request the library with your work email above. We'll build out all 5 mapped EvolutionIQ areas and set them up in a Corsac workspace, where you can run every test case against EvolutionIQ or your own agent with your own data.