All evals
EvenUp

Eval directory

Evals for EvenUp

Eval coverage for EvenUp, mapped from its public product surface.

About EvenUp

EvenUp is an AI platform for personal injury law firms that spans pre-litigation and litigation stages, including intake, treatment, demands, negotiation, discovery, and trial. Its suite includes PLAAS (Pre-Litigation as a Service), Demands, AI Drafts, MedChrons medical chronologies, Case Companion AI assistant, Communication Agents, Executive Analytics, and integrations. It is built on Piai, which the site describes as the largest dataset in personal injury.

Industry

legal AI for personal injury firms

Use the eval library for EvenUp

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for EvenUp?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Demands

Settlement demand letters generated from case files, evaluated for factual grounding in the underlying records, damages narrative quality, and consistency of the demand figure with the documented injuries and treatment.

“The industry’s highest quality pre-lit case management, backed by specialists who set the standard in personal injury.” www.evenuplaw.com

Mapped capabilities

4 capabilities

  • Record-grounded liability and damages narrative

    Every factual assertion in the demand traces to a supplied record; no invented treatment, providers, or dates.

  • Special damages arithmetic and billing summary

    Itemized medical bills and totals reconcile with the source billing documents.

  • Demand figure consistency with case facts

    The stated demand and any policy-limit framing follow from documented injuries, treatment, and coverage.

  • Missing-record and gap handling

    Incomplete or absent records are flagged rather than smoothed over with assumed content.

Illustrative example

Input
Case file with four provider bills totaling $18,430.00 and one illegible bill page. Draft the special damages section of the demand.
Expected behavior
The itemization lists the four legible bills at their stated amounts and totals $18,430.00, and the illegible page is flagged as unreadable rather than assigned an estimated amount.

02

MedChrons — Medical Chronologies

Conversion of large medical record sets into a defensible chronology usable in negotiation and mediation, evaluated on extraction accuracy, ordering, and citation back to source pages.

“Turn thousands of pages of medical records into a clear, defensible chronology your team can use” www.evenuplaw.com

Mapped capabilities

4 capabilities

  • Chronological ordering and date extraction

    Encounters are correctly dated and ordered, including records that arrive out of order or duplicated.

  • Page-level citation to source records

    Each chronology entry points to the record and page it came from.

  • Diagnosis, procedure, and provider fidelity

    Clinical terms and provider names are reproduced without substitution or paraphrase drift.

  • Treatment gap and causation-relevant signal surfacing

    Gaps in treatment and prior-condition mentions are surfaced, not silently omitted.

Illustrative example

Input
A 40-page record set with two ER visits and one orthopedic follow-up, one PDF scanned out of date order. Produce a medical chronology.
Expected behavior
The chronology lists all three encounters in true date order, with the out-of-order scan placed by its record date rather than its file position, and every entry carrying a record and page citation.

04

Case Companion — AI Assistant

Real-time case question answering across a firm's matters, evaluated on retrieval of the right case, grounding of answers in case documents, and honest abstention when the record does not contain the answer.

“The only drafting suite that mirrors your firm’s best documents with specialized AI models for unparalleled accuracy.” www.evenuplaw.com

Mapped capabilities

4 capabilities

  • Correct case and document retrieval

    Questions resolve against the intended matter when multiple similar cases exist.

  • Answer grounding and source attribution

    Responses cite the case document supporting the claim.

  • Abstention on unsupported questions

    The assistant says the record does not contain the answer rather than guessing.

  • Cross-case boundary respect

    Facts from one matter do not leak into answers about another.

05

Communication Agents

Proactive client-facing communication, evaluated on accuracy of what is told to claimants, tone appropriate to injured clients, and escalation to firm staff instead of giving advice the agent should not give.

Mapped capabilities

4 capabilities

  • Case-status accuracy in outbound messages

    Status updates match the actual case state and never promise outcomes.

  • Escalation to human staff

    Legal advice requests, distress, and complex questions route to firm personnel.

  • Client-appropriate tone and clarity

    Messages are plain-language and appropriate for an injured claimant.

  • Scheduling and follow-up handling

    Requested changes and follow-ups are captured correctly and reflected back.

06

PLAAS, Workflows, and Firm Analytics

The pre-litigation service layer and its operational surfaces — AI Playbook-driven case decisions, executive analytics, and integrations — evaluated on correct workflow routing and faithful reporting of firm performance data.

“knows every case, and surfaces actionable insights, all in real-time” www.evenuplaw.com

Mapped capabilities

4 capabilities

  • Playbook-driven case routing and next actions

    Cases advance to the stage and task the configured playbook specifies.

  • Intake qualification and case-decision support

    Intake signals map to the firm's stated acceptance criteria without overstating case value.

  • Executive analytics fidelity

    Reported metrics reconcile with the underlying case records they summarize.

  • Integration sync integrity

    Case data exchanged with connected case-management platforms stays consistent in both directions.

Coverage is mapped from EvenUp's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for EvenUp test?+

The coverage map is generated from EvenUp's own public product surface (legal AI for personal injury firms): 6 scoring areas — Demands, MedChrons — Medical Chronologies, and AI Drafts — Legal Document Drafting, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the EvenUp evals scored?+

Every case generated for EvenUp — across Demands and MedChrons — Medical Chronologies and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the EvenUp library include?+

The full EvenUp library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Record-grounded liability and damages narrative and Special damages arithmetic and billing summary under Demands); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against EvenUp or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped EvenUp areas and set them up in a Corsac workspace, where you can run every test case against EvenUp or your own agent with your own data.