All evals
S

Eval directory

Evals for Sixfold

Mapped eval coverage for Sixfold — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Sixfold

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Sixfold?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Underwriting Document Comprehension

Reading and extracting decision-relevant facts from the complex, unstructured artifacts that arrive with a submission or application, including loss runs, SOVs, APS, EHRs, lab results, and medical exams spanning hundreds of pages.

Sixfold reviews hundreds of pages across APS, EHRs, lab results, and medical exams www.sixfold.ai

Mapped capabilities

4 capabilities

  • Loss run and SOV extraction

    Pull structured exposure, location, and loss-history fields from tabular and scanned commercial submission documents.

  • Medical evidence parsing

    Extract diagnoses, medications, labs, and procedures from APS, EHRs, and lab reports across long unstructured records.

  • Long-document coverage

    Retain and surface hard-to-spot details buried deep in hundreds of pages rather than only front-matter or summaries.

  • Third-party and online research enrichment

    Incorporate external research and third-party data sources alongside submission-supplied documents.

02

Appetite and Guideline Application

Turning a specific carrier's underwriting manuals and stated risk appetite into consistent, automatic assessment of each submission, rather than generic model judgment.

Flags high-risk diagnoses according to your underwriting manuals www.sixfold.ai

Mapped capabilities

4 capabilities

  • Guideline ingestion fidelity

    Represent carrier-specific manual rules, thresholds, and exclusions as ingested rather than substituting industry defaults.

  • In-appetite vs out-of-appetite scoring

    Classify submissions against the carrier's appetite and explain which criteria drove the result.

  • High-risk condition flagging

    Flag diagnoses and exposures the carrier's own manual designates as high risk, in life/health and P&C alike.

  • Cross-carrier consistency

    Apply the same submission differently when appetite differs, and identically when appetite is unchanged.

Illustrative example

Carrier appetite excerpt states: 'Commercial property — decline any account with total insured value above $75M at a single location, or with two or more fire losses exceeding $250K in the trailing five years.' Attached SOV lists one location at $92M TIV; attached loss run shows a single $310K fire loss in 2023 and a $40K water loss in 2021. Ask: assess this submission against appetite and state the disposition. Classify the submission as out of appetite, citing the $92M single-location TIV against the stated $75M limit. Do not cite the loss history as a decline trigger, since only one fire loss exceeds $250K and the rule requires two or more. Cite the SOV and loss run as the sources for each figure, and recommend referral or decline consistent with the carrier's stated rule rather than a generic industry threshold.

03

Insight Synthesis and Traceability

Assembling the connected clinical or risk picture and attaching clear provenance so an underwriter can verify where every insight came from and how it was derived.

Mapped capabilities

4 capabilities

  • Condition-level clustering

    Group medications, labs, procedures, and diagnoses under the condition they belong to.

  • Citation to source location

    Attach a specific document and location to each surfaced insight.

  • Severity and progression narrative

    Convey history, severity, and progression rather than an undifferentiated fact list.

  • Rationale legibility

    State the reasoning path from evidence to finding in terms an underwriter can audit.

04

Discrepancy and Change Detection

Catching mismatches between what an applicant or broker submitted and what the underlying records show, and correctly handling evidence that arrives after the initial review.

Mapped capabilities

4 capabilities

  • Application vs record mismatch

    Surface contradictions between declared application data and medical or loss documentation.

  • New-fact surfacing on late evidence

    Highlight materially new information when additional documents arrive mid-case.

  • Missing or incomplete evidence handling

    Identify required documents that are absent and refrain from asserting facts they would supply.

  • Conflicting-source resolution

    Signal unresolved conflicts between sources instead of silently choosing one.

Illustrative example

Life application declares: 'No history of cardiac diagnosis or treatment.' Attached APS excerpt, page 47, records an office visit dated 2024-03-12 with assessment 'paroxysmal atrial fibrillation' and a prescription for apixaban 5mg BID. Ask: review this case and report what an underwriter needs to know. Surface the contradiction explicitly: the application declares no cardiac history while the APS documents a 2024 atrial fibrillation diagnosis with anticoagulant therapy. Cite the APS page and visit date as the source. Present it as a discrepancy for underwriter attention rather than choosing one version as true or omitting the application's declaration.

05

Next Best Action and Authority Boundaries

Recommending the concrete next step on each submission and acting only within the authority the carrier has granted, consistent with a decision-support rather than decision-making posture.

Mapped capabilities

4 capabilities

  • Next-action recommendation

    Propose the appropriate step such as referral to higher authority, broker follow-up, or additional evidence request.

  • Referral threshold adherence

    Escalate cases that exceed delegated authority rather than resolving them.

  • Action-taking within granted authority

    Execute agentic steps only where authority was explicitly conferred, and stop otherwise.

  • Broker and portfolio-fit context

    Weigh broker relationship and book fit alongside the standalone risk assessment.

06

Governance, Privacy, and Compliance Posture

Behavior consistent with the stated deployment and compliance model: isolated single-tenant environments, no end-user data persisted in the GenAI layer, no customer data training the models, SOC 2 Type II, and HIPAA.

Sixtold is SOC 2 Type II certified and HIPAA compliant www.sixfold.ai

Mapped capabilities

4 capabilities

  • Tenant isolation claims

    Never surface or imply access to another carrier's data, guidelines, or cases.

  • PHI and sensitive-data handling

    Handle protected health information consistent with the stated HIPAA and non-persistence posture.

  • Accurate self-description of scope

    Represent supported lines and the non-workbench, decision-support positioning without overclaiming.

  • Human accountability framing

    Preserve that underwriters govern exceptions and remain accountable for the decision.

Coverage is mapped from Sixfold's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sixfold test?+

The coverage map above is generated from Sixfold's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sixfold evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sixfold library include?+

The full Sixfold library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sixfold or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Sixfold or your own agent with your own data.