All evals
Isomer

Eval directory

Evals for Isomer

Eval coverage for Isomer, mapped from its public product surface.

About Isomer

Isomer is an AI intelligence layer that sits ahead of a carrier's claims management system, ingesting every inbound communication (email, letters, attachments, faxes) and mapping it to the right claim. It evaluates each item against a library of 85+ insurance-specific risk detectors — time-limited demands, spoliation notices, attorney representation, fraud patterns — and routes findings or triggers workflows. It targets carriers, MGAs, TPAs, and self-insured enterprises, with single-tenant architecture and SOC 2 Type II compliance.

Industry

insurance claims communication risk-detection AI

Use the eval library for Isomer

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Isomer?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Inbound Ingestion & Claim Association

Accepting every inbound channel described in the product — email, letters, attachments, faxes — and mapping each item to the correct claim before it reaches the CMS.

Isomer processes every inbound communication before it reaches your CMS, so nothing is lost, delayed, or reduced. www.isomer.ai

Mapped capabilities

4 capabilities

  • Multi-channel intake

    Email, scanned letters, PDF attachments, and inbound fax handled as first-class inputs, including multi-page and mixed-format documents.

  • Claim matching

    Associating an item to the right open claim from identifiers, parties, dates of loss, and thread history; behavior when no confident match exists.

  • Thread and duplicate handling

    Long back-and-forth threads, forwarded chains, and re-sent documents resolved without double-counting or losing the originating item.

  • Timeline placement

    Ordering ingested items on the claim timeline by receipt, including weekend and after-hours arrival relative to FNOL.

02

Risk Signal Detection

Evaluating each communication against the 85+ insurance-specific detectors spanning the seven risk categories referenced in the Signal Library.

Evaluate every inbound communication against 85+ insurance-specific detectors at arrival. www.isomer.ai

Mapped capabilities

4 capabilities

  • Immediate risk detectors

    Time-limited demands, spoliation and evidence-preservation notices, and critical-severity events identified at the point of receipt.

  • Representation and litigation posture

    Notices of attorney representation, bad-faith setups, and sanctions-related correspondence surfaced as escalations.

  • Emerging patterns across a claim

    Escalation that builds over a claim lifecycle, such as a routine FNOL drifting into litigation posture over weeks.

  • Portfolio-wide patterns

    Fraud rings, medical-legal referral patterns, and aggregate limit erosion visible only across claims, not within one file.

Illustrative example

Input
A 12-page PDF demand letter arrives by email from a plaintiff firm. Page 8 states a $475,000 demand with a 30-day response deadline. Earlier pages are medical narrative.
Expected behavior
Fires the time-limited demand detector despite the trigger appearing mid-document, extracts the $475,000 amount and the 30-day deadline as a dated response date, and marks the finding urgent with a citation to page 8.

03

Structured Signal Output & Evidence

What each detected signal carries downstream: source evidence, urgency, confidence, and the configured response.

Every signal carries source evidence, urgency scoring, confidence level, and a configured response. www.isomer.ai

Mapped capabilities

4 capabilities

  • Source-evidence citation

    Every signal points back to the specific document and page or passage that triggered it.

  • Urgency and confidence scoring

    Consistent urgency scoring and calibrated confidence attached to each finding.

  • Extracted claim entities

    Deadlines, demand amounts, parties, providers, and severity indicators pulled into structured fields.

  • Cross-claim entity linking

    Linking a provider, firm, or claimant appearing on one claim to their occurrences on others.

04

Routing, Actions & Workflows

Turning findings into delivery and downstream execution across the customer's existing stack.

Mapped capabilities

4 capabilities

  • Destination routing

    Sending findings to the right team via Slack, email, or SMS according to configuration.

  • CMS write-back

    Writing structured signal data back into the claims management system without corrupting or duplicating records.

  • Agentic workflow triggering

    Invoking a managed Isomer Actions workflow when a signal's configured response calls for it.

  • Stack connectivity

    Operating through Microsoft 365, Google Workspace, email forwarding, or API intake as documented.

05

AI Governance & Fail-Safe Behavior

How the system behaves under uncertainty and how its decisions are made auditable in a regulated workflow.

Mapped capabilities

4 capabilities

  • Uncertainty escalation

    Low-confidence or ambiguous items routed to human review rather than silently approved or auto-actioned.

  • Processing audit trail

    Logging what was extracted, which signals were evaluated, and which actions were taken for each document.

  • Regulatory framing

    Answers consistent with the stated alignment to the NAIC Model Bulletin on AI Systems and the NIST AI Risk Management Framework.

  • Decision explainability

    Explaining why a signal fired or did not fire in terms an adjuster or examiner can review.

Illustrative example

Input
An inbound fax of medical records names a claimant matching two open claims with different dates of loss and carries no claim number or policy number.
Expected behavior
Does not attach the document to either claim on its own. Routes it to human review as an unresolved match, presents both candidate claims with the conflicting dates of loss, and logs the item as received so it is not dropped.

06

Tenancy, Security & Data Handling

The isolation, credential, and retention guarantees stated on the Trust page, as they surface in product behavior and customer-facing answers.

We never use customer data to train our foundational models. www.isomer.ai

Mapped capabilities

4 capabilities

  • Single-tenant isolation

    No cross-customer data access; dedicated pipelines from ingestion through extraction.

  • Credential and access model

    OAuth 2.0 token-based connections with automatic renewal, no stored passwords, IP allowlisting for API access.

  • Retention and training policy

    Zero-day retention default for foundation model interactions and no use of customer data to train foundational models.

  • PHI and compliance posture

    BAA-covered workflows involving protected health information and SOC 2 Type II control claims stated accurately.

Coverage is mapped from Isomer's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Isomer test?+

The coverage map is generated from Isomer's own public product surface (insurance claims communication risk-detection AI): 6 scoring areas — Inbound Ingestion & Claim Association, Risk Signal Detection, and Structured Signal Output & Evidence, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Isomer evals scored?+

Every case generated for Isomer — across Inbound Ingestion & Claim Association and Risk Signal Detection and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Isomer library include?+

The full Isomer library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Multi-channel intake and Claim matching under Inbound Ingestion & Claim Association); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Isomer or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Isomer areas and set them up in a Corsac workspace, where you can run every test case against Isomer or your own agent with your own data.