All evals
Everlaw

Eval directory · Legal AI

Evals for Everlaw

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Everlaw AI products.

About Everlaw

Everlaw is a cloud-native litigation and e-discovery platform used by law firms, corporations, and government agencies to manage the full discovery lifecycle — from document review and analysis to deposition prep and trial. Its AI features accelerate review, surface key documents, and assist with case narrative and writing.

Employees

~700

Industry

Legal AI / E-Discovery

Headquarters

Oakland, CA

Use the eval library for Everlaw

All 121 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Everlaw?

4 areas · 121 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Batch Genai Actions At Scale

Mapped capabilities

16 scenarios

  • Multi-action batch job combining summarize and extract

Public sample case

Input
Batch job BJ-2291 ran summarization and extraction on 500 documents. The audit log has one merged entry per document (combining both actions) rather than two separate action-labeled entries. The compliance officer has asked the a…
Expected behavior
Agent counts action-labeled log entries per document. It detects that only one entry exists per document rather than two separate entries bearing distinct action type labels ('summarize', 'extract'). It reports this discrepancy explicitly, states it cannot certify per-action logging compliance, and…
Check
Pass / fail check

02

Coding Suggestions Tiered Ai Coding With Justifications

Mapped capabilities

16 scenarios

  • Standard-tier assignment for a single document

Public sample case

Input
A 6-page contract addendum contains the exact phrases 'as amended by Exhibit C-4', 'warranty exclusion clause 12(a)', and 'delivery deadline of October 31'. It does NOT contain the phrase 'force majeure provision' or any year-spe…
Expected behavior
Every phrase enclosed in quotation marks or explicitly attributed to the document in the justification appears verbatim in DOC-00294's source text. The phrase 'force majeure provision' does not appear in the justification. No year-specific date (e.g., 'October 31, 2019') appears when the source con…
Check
Pass / fail check

03

Predictive Coding Statistical Defensibility

Mapped capabilities

38 scenarios

  • Seed-set document selection

Public sample case

Input
The agent is asked to add a specific parent email (EVR-00770) to the seed set as a relevant exemplar. EVR-00770 has 11 attachment children, including three Excel spreadsheets categorized as financial data and one compressed archi…
Expected behavior
Before committing any seed addition, the agent detects that EVR-00770 has 11 attachment children and that the family-cascade rule is active. It presents the user with the specific cascade consequence—naming the child count, document types (3 Excel, 1 archive, 7 other), and the label that would be a…
Check
Pass / fail check

04

Writing Assistant Storybuilder Draft Quality

Mapped capabilities

51 scenarios

  • Initial legal argument draft from key documents

Frequently asked questions

What do the Corsac evals for Everlaw test?+

Each eval pack tests Everlaw's public product surface — including Batch Genai Actions At Scale, Coding Suggestions Tiered Ai Coding With Justifications, and Predictive Coding Statistical Defensibility — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Everlaw evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 121 Everlaw cases — from Writing Assistant Storybuilder Draft Quality (51 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Everlaw library.

How many test cases does the Everlaw library include?+

The Everlaw eval library includes 121 graded test cases across 4 eval packs, the largest being Writing Assistant Storybuilder Draft Quality with 51 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Everlaw or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 4 Everlaw packs — Batch Genai Actions At Scale and Coding Suggestions Tiered Ai Coding With Justifications and the rest — against Everlaw or your own agent with your own data.