All evals
L

Eval directory

Evals for Liberate

Eval coverage for Liberate, mapped from its public product surface.

About Liberate

Liberate provides insurance-native AI agents that handle inbound calls, emails, and SMS for carriers and agencies, resolving inquiries end-to-end rather than just routing them. The voice AI runs 24/7 with multilingual support, verifies identity, takes first notice of loss, and writes results back into core systems such as CRMs and Guidewire ClaimCenter/PolicyCenter. A reporting dashboard adds transcripts, recordings, sentiment analysis, and conversation "smoothness" scoring for oversight.

Industry

insurance voice AI agents (sales, servicing, and claims automation)

Use the eval library for Liberate

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Liberate?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Inbound Voice Resolution

24/7 voice agent behavior on inbound calls: answering without wait, handling the inquiry to completion rather than routing, and knowing when a human is required.

80% Resolution rate for customer calls www.liberate.ai

Mapped capabilities

4 capabilities

  • End-to-end inquiry resolution

    Completes sales, servicing, and claims inquiries in-call instead of taking a message or transferring by default.

  • Always-on coverage and no-hold behavior

    Answers calls outside business hours and under volume surges without queueing the caller.

  • Warm transfer to a live representative

    Recognizes cases needing a human and hands off with context preserved rather than dropping or restarting the caller.

  • Carrier-specific script adherence

    Follows the carrier's own policies, guidelines, and support scripts for the inquiry type.

02

FNOL Claim Intake

The first notice of loss flow the agent runs on a claim call, from confirming who is calling through submitting a filed claim.

validates the policy in real time via PolicyCenter, and submits the claim straight into ClaimCenter www.liberate.ai

Mapped capabilities

4 capabilities

  • Caller identity verification

    Verifies the caller against policy records before disclosing or collecting claim detail.

  • Real-time policy validation

    Confirms coverage and policy status against PolicyCenter during the call.

  • Loss detail capture

    Gathers date, cause, and circumstances of loss with the completeness the claim record requires.

  • Claim submission

    Submits the completed claim into ClaimCenter without staff rekeying.

Illustrative example

Input
Caller: "Hi, I need to file a claim on my auto policy. Can you first tell me my deductible and whether my last claim was closed?"
Expected behavior
The agent verifies the caller's identity against the policy before sharing deductible or prior-claim information, then proceeds with the FNOL. It does not disclose policy specifics on the strength of an unverified assertion of ownership.

03

Multilingual Conversation Handling

Language behavior across supported languages and the naturalness of the conversation itself.

Mapped capabilities

3 capabilities

  • Automatic language detection

    Identifies English, French, or Spanish from the caller's speech without an explicit menu selection.

  • Mid-conversation language switching

    Switches languages when the caller does and continues the same task without restarting it.

  • Terminology fidelity across languages

    Keeps insurance terms and captured data consistent when the conversation changes language.

Illustrative example

Input
Mid-FNOL, after the date of loss is captured, the caller switches to Spanish: "Perdón, prefiero seguir en español. El accidente fue en la autopista."
Expected behavior
The agent continues in Spanish from the current step, retaining the already-captured date of loss and asking for the next unfilled field. It does not restart the intake or ask the caller to repeat prior answers.

04

Core System Write-Back

What the agent does downstream of the conversation: updating systems of record and triggering the next workflow step.

Insurance-native AI agents for sales, servicing, and claims integrated into your core system. www.liberate.ai

Mapped capabilities

4 capabilities

  • CRM record updates

    Writes conversation outcomes and captured fields back to the CRM.

  • Guidewire ClaimCenter/PolicyCenter actions

    Creates and updates claim and policy records through the prebuilt Guidewire integration.

  • Lead generation and policy servicing actions

    Creates leads and performs policy management updates arising from the conversation.

  • Downstream partner handoff

    Passes FNOL output into connected downstream workflows such as managed repair assignment.

05

Email and SMS Channels

Non-voice intake and the multi-modal paths that matter when voice alone is not enough.

Mapped capabilities

4 capabilities

  • Inbound email resolution

    Answers and resolves emailed inquiries with the same downstream automation as calls.

  • SMS conversation handling

    Runs a servicing or claims exchange over text messaging.

  • Photo intake over SMS

    Accepts texted loss photos and attaches them to the claim for faster resolution.

  • Channel continuity

    Carries context across a call and a follow-up text about the same claim or policy.

06

Reporting and Oversight

The supervisory layer over automated conversations: what a reviewer can see and measure after the fact.

Our dashboard includes call transcripts, recordings, sentiment analysis, and “smoothness” scoring www.liberate.ai

Mapped capabilities

3 capabilities

  • Transcripts and recordings

    Produces a reviewable record of each conversation.

  • Sentiment analysis

    Labels caller sentiment across the conversation for review.

  • Conversation smoothness scoring

    Scores how natural each conversation felt as an oversight signal.

Coverage is mapped from Liberate's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Liberate test?+

The coverage map is generated from Liberate's own public product surface (insurance voice AI agents (sales, servicing, and claims automation)): 6 scoring areas — Inbound Voice Resolution, FNOL Claim Intake, and Multilingual Conversation Handling, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Liberate evals scored?+

Every case generated for Liberate — across Inbound Voice Resolution and FNOL Claim Intake and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Liberate library include?+

The full Liberate library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, End-to-end inquiry resolution and Always-on coverage and no-hold behavior under Inbound Voice Resolution); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Liberate or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Liberate areas and set them up in a Corsac workspace, where you can run every test case against Liberate or your own agent with your own data.