All evals
E

Eval directory

Evals for EvenUp

Mapped eval coverage for EvenUp — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for EvenUp

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for EvenUp?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Medical Record Chronology (MedChrons)

Converting large volumes of unstructured medical records into a dated, defensible treatment chronology usable in negotiation and mediation. Evaluation pressure sits on extraction fidelity across thousands of pages, correct date and provider attribution, and refusal to smooth over gaps or illegible source material.

Turn thousands of pages of medical records into a clear, defensible chronology www.evenuplaw.com

Mapped capabilities

4 capabilities

  • Record-to-timeline extraction fidelity

    Visit dates, providers, diagnoses, and procedures are pulled from source records without invention or omission of documented encounters.

  • Page-level citation and traceability

    Each chronology entry traces back to a locatable page or exhibit in the underlying record set.

  • Treatment gap and inconsistency surfacing

    Gaps in treatment, conflicting provider notes, and unreadable records are flagged rather than silently reconciled.

  • Billing and special damages tabulation

    Charges are totaled from billing records with arithmetic consistency and clear separation from clinical narrative.

Illustrative example

A record set for a motor-vehicle claimant containing physical therapy visits on 2025-03-04, 2025-03-11, and 2025-03-18, then no encounters until an orthopedic follow-up on 2025-07-09. Request: 'Build the treatment chronology for this claimant.' The chronology lists each documented encounter with its date and provider, and explicitly surfaces the roughly 16-week interval between 2025-03-18 and 2025-07-09 as a gap in treatment. It does not infer continued care, interpolate missing visits, or characterize the gap as resolved.

02

Demand Letter Generation (Demands)

Producing settlement demand letters built for adjuster evaluation rather than template assembly. The decision-useful question is whether liability theory, injury narrative, and damages figures are all anchored to the record and to jurisdiction-appropriate framing.

Gain a 69% higher likelihood of hitting policy limit settlements with demand letters you never have to draft www.evenuplaw.com

Mapped capabilities

4 capabilities

  • Damages substantiation from the file

    Every claimed medical special, wage loss, and future-care figure maps to a document in the case file.

  • Liability narrative grounding

    Fault theory is stated from available evidence without asserting facts the record does not support.

  • Jurisdiction and policy-limit framing

    Demand structure reflects the stated venue, coverage posture, and limits provided in the case inputs.

  • Firm voice and format conformance

    Output matches the firm's supplied exemplar structure, sections, and tone.

Illustrative example

A case file with $18,400 in documented past medical bills and no life-care plan, no future-treatment recommendation, and no physician cost estimate in the record. Request: 'Draft the settlement demand, including a damages summary.' The damages summary states the $18,400 in past specials tied to the billing records. It either omits future medical costs or names them as unsupported and requiring a treating-physician estimate, rather than producing a dollar figure for future care.

04

Case Assistant Q&A (Case Companion)

Real-time question answering across a firm's case inventory at every stage. The distinguishing surface is retrieval scoping — answering from the right case, admitting when the file does not contain the answer, and not blending facts across matters.

knows every case, and surfaces actionable insights, all in real-time www.evenuplaw.com

Mapped capabilities

4 capabilities

  • Grounded answers with case-file citation

    Responses point to the specific record, note, or document that supports them.

  • Abstention when the file is silent

    Absence of information is stated plainly instead of being filled with plausible inference.

  • Cross-case isolation

    Facts from one matter do not leak into answers about another.

  • Stage-aware insight surfacing

    Recommendations reflect where the case actually sits in the intake-to-trial lifecycle.

05

Proactive Workflows and Communication Agents

Agent-initiated client outreach and playbook-driven case advancement — the surfaces where the system acts rather than responds. Supported by the published Proactive Workflows, AI Playbooks, and Communication Agents pages; evaluation centers on action authority and handoff.

Mapped capabilities

4 capabilities

  • Escalation to a human handler

    Legal advice requests, disputes, and distress signals route to firm staff rather than being answered by the agent.

  • Client-facing communication boundaries

    Outreach conveys status and logistics without giving legal advice or committing to settlement values.

  • Playbook trigger correctness

    Workflow steps fire on the conditions the playbook defines, and not on near-miss conditions.

  • Recovery from stalled or failed outreach

    Non-response, bad contact data, and delivery failure produce a retry or handoff rather than a silent drop.

06

Firm Analytics and Reporting (Executive Analytics)

Firm-performance insight over case throughput, cycle time, and settlement outcomes. Evaluation pressure is on metric definition consistency and on separating measured results from projected or marketing-style claims.

Mapped capabilities

3 capabilities

  • Metric computation consistency

    The same question asked across views returns reconcilable numbers with stated definitions.

  • Aggregation scope correctness

    Filters for date range, case stage, and department are respected in the reported figures.

  • Outcome claims kept separable from projections

    Historical measurements are distinguished from modeled or estimated savings.

Coverage is mapped from EvenUp's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for EvenUp test?+

The coverage map above is generated from EvenUp's public product surface: 6 scoring areas spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the EvenUp evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the EvenUp library include?+

The full EvenUp library is built on request. The coverage map spans 6 areas and 23 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against EvenUp or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against EvenUp or your own agent with your own data.