All evals
Sixfold

Eval directory

Evals for Sixfold

Eval coverage for Sixfold, mapped from its public product surface.

About Sixfold

Sixfold is an AI system for insurance underwriters that ingests a carrier's underwriting guidelines and risk appetite and automatically assesses incoming submissions against them. It covers commercial P&C, specialty, and life, health and disability lines, reading documents such as loss runs, SOVs, APS records, EHRs, and lab results to surface cited risk insights and next best actions. It runs in isolated, single-tenant environments and positions itself as a decision-support layer rather than an underwriting workbench.

Industry

insurance underwriting AI

Headquarters

New York City (NYC office, hybrid)

Use the eval library for Sixfold

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Sixfold?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Appetite & Guideline Alignment

Whether the system faithfully applies a specific carrier's ingested underwriting guidelines and risk appetite, rather than generic industry judgment, when assessing an incoming submission.

Mapped capabilities

4 capabilities

  • Guideline ingestion fidelity

    Findings reflect the carrier's own stated rules, thresholds and exclusions as written.

  • In-appetite vs. out-of-appetite calls

    Correct classification of submissions against appetite, including borderline and mixed cases.

  • Carrier-specific over generic reasoning

    Does not substitute general underwriting heuristics where the carrier's manual differs.

  • Line-of-business coverage

    Consistent behavior across commercial P&C, specialty, and life, health and disability.

Illustrative example

Input
Carrier guidelines exclude frame-construction habitational risks above four stories. A submission arrives for a five-story frame apartment building with clean loss history and strong broker relationship.
Expected behavior
Flags the risk as outside the carrier's stated appetite on the construction and height exclusion, citing that guideline, and does not let the clean loss runs or broker relationship override the written rule.

02

Submission Document Comprehension (P&C)

Extraction and interpretation of the complex documents that arrive with a commercial submission, where structure is inconsistent and detail is buried.

“Reads complex documents such as loss runs, SOVs and more” www.sixfold.ai

Mapped capabilities

4 capabilities

  • Loss run interpretation

    Reads claim history, frequency and severity patterns out of varied loss run formats.

  • SOV extraction

    Pulls location, construction and value detail from large statements of values.

  • Unstructured and mixed-format handling

    Behavior on scanned, malformed or partially illegible submission documents.

  • External and third-party enrichment

    Use of online research and third-party data alongside submitted materials.

03

Medical Evidence Synthesis (Life & Health)

Turning hundreds of pages of clinical evidence into a connected, condition-level picture an underwriter can act on.

“Sixfold reviews hundreds of pages across APS, EHRs, lab results, and medical exams” www.sixfold.ai

Mapped capabilities

4 capabilities

  • Condition-level assembly

    Groups medications, labs, procedures and diagnoses under the relevant condition.

  • High-risk diagnosis flagging

    Surfaces diagnoses the carrier's underwriting manual treats as high risk.

  • Severity and progression

    Conveys history, severity and progression rather than a flat summary.

  • Application-versus-record discrepancies

    Detects mismatches between application data and medical records, including newly arriving facts.

Illustrative example

Input
A life application states no tobacco use. The attached APS includes a primary care note recording current daily smoking and a nicotine-positive lab result.
Expected behavior
Reports the discrepancy between the application's tobacco declaration and the medical evidence, cites both the physician note and the lab result as sources, and routes it as a point requiring underwriter follow-up rather than resolving it silently.

04

Citation & Traceability

Every surfaced insight is documented back to where the information came from and how it was reached, which is the basis for underwriter trust and audit.

Mapped capabilities

3 capabilities

  • Source attribution accuracy

    Citations point to the passage that actually supports the claim.

  • Unsupported claim suppression

    Avoids asserting findings the underlying documents do not support.

  • Reasoning transparency

    Shows how an insight follows from the cited evidence and the applied guideline.

05

Next Best Action & Bounded Authority

Recommending, and where authorized taking, the appropriate next step on a submission while leaving the underwriter accountable for the decision.

Mapped capabilities

4 capabilities

  • Next best action selection

    Chooses among referral to higher authority, broker follow-up, or proceed.

  • Authority boundaries

    Acts only within granted authority and defers otherwise.

  • Decision-support posture

    Shapes the path to a decision without issuing the bind/decline decision itself.

  • Broker and portfolio fit context

    Weighs broker context and book fit alongside the standalone risk.

06

Privacy, Isolation & Compliance Posture

The deployment and data-handling commitments a carrier's security and compliance reviewers evaluate before adoption.

“All Sixfold customers operate within isolated environments, and end-user data is never persisted in the LLM-powered Gen AI layer” www.sixfold.ai

Mapped capabilities

3 capabilities

  • Single-tenant isolation

    Data stays within the carrier's isolated instance.

  • Non-persistence in the Gen AI layer

    End-user data is not persisted in the LLM-powered layer or used to train models.

  • Certification claims

    Accurate representation of SOC 2 Type II and HIPAA posture when asked.

Coverage is mapped from Sixfold's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sixfold test?+

The coverage map is generated from Sixfold's own public product surface (insurance underwriting AI): 6 scoring areas — Appetite & Guideline Alignment, Submission Document Comprehension (P&C), and Medical Evidence Synthesis (Life & Health), and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sixfold evals scored?+

Every case generated for Sixfold — across Appetite & Guideline Alignment and Submission Document Comprehension (P&C) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sixfold library include?+

The full Sixfold library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, Guideline ingestion fidelity and In-appetite vs. out-of-appetite calls under Appetite & Guideline Alignment); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sixfold or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Sixfold areas and set them up in a Corsac workspace, where you can run every test case against Sixfold or your own agent with your own data.