All evals
F

Eval directory

Evals for Federato

Mapped eval coverage for Federato — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Federato

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Federato?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Submission Intake & Data Extraction

Ingesting broker submissions and turning unstructured documents and correspondence into structured, reviewable submission data at the front of the lifecycle.

Follow a submission from intake to quote in minutes, not days. www.federato.ai

Mapped capabilities

4 capabilities

  • Structured field extraction from submission documents

    Named insured, exposures, limits, and other submission fields pulled from supplied materials rather than retyped.

  • Provenance and traceability of extracted values

    Each extracted value can be traced back to where it came from in the source submission.

  • Missing or ambiguous data handling

    Gaps surfaced as unresolved rather than silently filled, so the underwriter sees what is incomplete.

  • Consolidated single-system intake

    Intake replaces multi-system handoffs, per the platform's consolidation claim.

Illustrative example

A broker email plus an attached submission packet for a mid-market commercial property risk. The packet states the named insured, primary location address, and requested limit, but the construction class field is blank and the loss run covers only two of the requested five years. The agent extracts the fields that are actually present and attaches each to its source location in the packet. Construction class and the three missing loss-run years are reported as unresolved gaps requiring broker follow-up or underwriter input. No value is supplied for construction class by inference from the address, occupancy, or comparable risks, and the submission is not presented as complete.

02

Agentic Underwriting & Quote Generation

The agentic path from extracted submission to a complete, explained, on-strategy quote staged for underwriter review.

Our agentic AI creates complete, fully explained, on strategy quotes for underwriter review in minutes www.federato.ai

Mapped capabilities

4 capabilities

  • Risk assessment on extracted submission data

    Agent evaluates the risk using the data it extracted and available signals.

  • Underwriting guideline application

    Carrier guidelines applied to the submission, with the outcome reflecting the governing rules.

  • Explained quote drafting

    Quote is complete and fully reasoned, not a bare number, so the rationale is inspectable.

  • Underwriter review handoff

    Draft output is positioned for human review rather than autonomous binding.

03

Appetite & Winnability Triage

Strategy-first prioritization that replaces first-in-first-out queueing with scoring on appetite, guidelines, and winnability.

Mapped capabilities

4 capabilities

  • Non-FIFO queue ordering

    High-value deals advance ahead of earlier-arriving lower-fit submissions.

  • Appetite scoring against carrier strategy

    Submissions graded on fit with the carrier's stated appetite.

  • Winnability signals in prioritization

    Likelihood of winning the business factored alongside appetite.

  • Triage rationale exposed to the underwriter

    Why a submission was ranked where it was is visible, not opaque.

Illustrative example

Five submissions enter the queue in a fixed order. Submissions 1 through 4 arrive first and fall outside the carrier's stated appetite on class or geography. Submission 5 arrives last, sits squarely in the defined high-appetite class, and carries strong winnability signals. The triaged queue places submission 5 above submissions 1 through 4 rather than preserving arrival order, and the displayed rationale for that ranking cites appetite fit and winnability. Arrival time is not presented as the reason submission 5 was promoted.

04

Product Configuration & Governance (Product Studio)

Defining rating, forms, and eligibility in one place, and rolling changes across states and programs with controlled, auditable versioning.

Rating, forms, and eligibility all in one place. www.federato.ai

Mapped capabilities

4 capabilities

  • Product-level rating definition

    Rates set at the product level so pricing reflects product structure and flows through.

  • Policy forms and eligibility rules

    Forms and eligibility configured alongside rating rather than scattered across tools.

  • Multi-state and multi-program rollout

    A change made once propagates across states and programs without rework.

  • Versioned rules and audit trail

    What is in force is tracked with version history for regulatory defensibility.

05

Portfolio Oversight & Control Tower

Real-time visibility for underwriting leaders into growth, exposure, performance, and whether frontline decisions stay aligned to portfolio strategy.

89% reduction in time to quote www.federato.ai

Mapped capabilities

4 capabilities

  • Live growth, exposure, and performance tracking

    Portfolio metrics available in real time rather than in periodic reports.

  • Team activity oversight

    Leaders can see underwriting activity across the team.

  • Quote-to-strategy alignment monitoring

    Quotes checked against the portfolio strategy they are meant to serve.

  • Portfolio mix reporting on appetite

    Share of bound business meeting the carrier's high-appetite definition is measurable.

06

Downstream Lifecycle & Distribution Surfaces

Post-quote lifecycle modules and the external-facing portals through which producers and policyholders interact with the platform.

Mapped capabilities

4 capabilities

  • Billing and payments

    Billing and payment handling as part of the same end-to-end platform.

  • Claims handling

    Claims module covering the loss side of the lifecycle.

  • Claims-to-portfolio feedback loop

    Claims outcomes feed back into portfolio strategy, per the Claims launch framing.

  • Producer and policyholder portals

    Distribution surfaces for brokers/producers and for insureds.

Coverage is mapped from Federato's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Federato test?+

The coverage map above is generated from Federato's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Federato evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Federato library include?+

The full Federato library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Federato or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Federato or your own agent with your own data.