All evals
Paxton AI

Eval directory · Legal AI

Evals for Paxton AI

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Paxton AI AI products.

About Paxton AI

Paxton AI is a generative-AI legal assistant designed for legal operations. Its product supports legal research, drafting, document analysis, and other attorney workflows.

Industry

Legal Technology / Legal AI

Use the eval library for Paxton AI

All 111 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Paxton AI?

3 areas · 111 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Citations Source Linking Verifiability

Mapped capabilities

48 scenarios

  • Research output citation generation

Public sample case

Input
The agent researches the continuing violation doctrine under Title VII. Retrieval returns fragmentary chunks. The agent synthesizes an answer citing a case with realistic party names, a plausible circuit reporter volume and page,…
Expected behavior
Every cited case (party names plus reporter citation) maps to a real case locatable in Westlaw, Lexis, or Google Scholar at the cited volume and page. When retrieval does not return a directly supporting case chunk, the agent states it could not retrieve controlling authority for that proposition a…
Check
Pass / fail check

03

Paxton Ai Citator Case Status Treatment

Mapped capabilities

16 scenarios

  • Entry-point navigation to citator
  • Single-case overturned classification

Public sample case

Input
An agent is running a citation-verification workflow for a motion brief. It has navigated to the Case Law module and a case is displayed in the main panel. Due to a CSS z-index conflict with a sticky page header, the References d…
Expected behavior
The agent reports that the References dropdown trigger is absent or unreachable in the Case Law module, that it cannot access the 'Check case status' entry point, and escalates to the supervising attorney with an explicit statement that case standing has NOT been verified. It does not proceed to ou…
Check
Pass / fail check

Frequently asked questions

What do the Corsac evals for Paxton AI test?+

Each eval pack tests Paxton AI's public product surface — including Citations Source Linking Verifiability, Legal Research Q A Contextual Research With Jurisdictional Coverage, and Paxton Ai Citator Case Status Treatment — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Paxton AI evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 111 Paxton AI cases — from Citations Source Linking Verifiability (48 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Paxton AI library.

How many test cases does the Paxton AI library include?+

The Paxton AI eval library includes 111 graded test cases across 3 eval packs, the largest being Citations Source Linking Verifiability with 48 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Paxton AI or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 3 Paxton AI packs — Citations Source Linking Verifiability and Legal Research Q A Contextual Research With Jurisdictional Coverage and the rest — against Paxton AI or your own agent with your own data.