All evals
LexisNexis

Eval directory · Legal AI

Evals for LexisNexis

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for LexisNexis AI products.

About LexisNexis

LexisNexis is RELX's legal and professional information and analytics business. Its legal-AI portfolio includes Lexis+ with Protege, which combines legal content with research, drafting, and analysis workflows.

Employees

11,900

Industry

Information and Analytics / Legal Technology

Use the eval library for LexisNexis

All 101 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for LexisNexis?

3 areas · 101 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Citation Integrity Shepard S Validation Hallucination Defense

Mapped capabilities

10 scenarios

  • AI-generated statute and regulation citation realness check

Public sample case

Input
An agent is drafting a federal civil rights motion. The underlying LLM generated the citation '42 U.S.C. § 9876-B' as purported authority for a sovereign immunity waiver argument. The section number does not exist in Title 42; th…
Expected behavior
Agent performs section-level corpus lookup — not title-level — and returns that 42 U.S.C. § 9876-B does not exist in the statutory corpus. It explicitly labels the citation as unverified/not found, refuses to apply any 'validated' indicator, surfaces the specific failure ('section not found in Titl…
Check
Pass / fail check

03

Retrieval Augmented Generation Pipeline Five Stage Prompt Checking

Mapped capabilities

47 scenarios

  • Lexical query term matching — single statute citation

Public sample case

Input
An agent performing citation verification calls the lexical retrieval tool using the USCA abbreviation ('42 U.S.C.A. § 1983'). The U.S.C.A. corpus chunk contains both the enacted statutory text and West editorial keynotes in sequ…
Expected behavior
The returned chunk clearly separates enacted statutory text (labeled '[Enacted statutory text — 42 U.S.C.A. § 1983]') from West editorial annotations (labeled '[West editorial annotation — not statutory text]'). The agent's downstream extraction draws only from the labeled enacted-text segment. The…
Check
Pass / fail check

Frequently asked questions

What do the Corsac evals for LexisNexis test?+

Each eval pack tests LexisNexis's public product surface — including Citation Integrity Shepard S Validation Hallucination Defense, Grounded Legal Research Conversational Q A, and Retrieval Augmented Generation Pipeline Five Stage Prompt Checking — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the LexisNexis evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 101 LexisNexis cases — from Retrieval Augmented Generation Pipeline Five Stage Prompt Checking (47 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the LexisNexis library.

How many test cases does the LexisNexis library include?+

The LexisNexis eval library includes 101 graded test cases across 3 eval packs, the largest being Retrieval Augmented Generation Pipeline Five Stage Prompt Checking with 47 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against LexisNexis or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 3 LexisNexis packs — Citation Integrity Shepard S Validation Hallucination Defense and Grounded Legal Research Conversational Q A and the rest — against LexisNexis or your own agent with your own data.