All evals
DraftWise

Eval directory · Legal AI

Evals for DraftWise

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for DraftWise AI products.

About DraftWise

DraftWise is an AI-powered firm intelligence platform that helps attorneys draft and review contracts faster through access to their firm's cumulative knowledge and data. DraftWise was founded by Palantir and Google engineers, James Ding and Emre Ozen, and Clifford Chance lawyer, Ozan Yalti, to bring advanced analytics and data management to the legal space, enabling lawyers to spend more time on the critical task of representing their clients.

Industry

Legal AI

Headquarters

New York City

Use the eval library for DraftWise

All 121 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for DraftWise?

3 areas · 121 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Associate Agentic Drafting Redlining

Mapped capabilities

49 scenarios

  • Natural-language request parsing into discrete drafting actions

Public sample case

Input
A partner wants an analytical opinion on whether the indemnity cap is market before deciding to redline it. She uses 'review' deliberately to request analysis, not edits. The AI Associate is running in the Word add-in with write …
Expected behavior
Agent classifies the instruction as a Q&A / analysis action. It returns a memo-style analysis comparing the cap against market norms for Series B SaaS deals, citing relevant precedents where available. Zero tracked changes, redlines, or suggested revisions are applied to or proposed for the documen…
Check
Pass / fail check

02

Citation Provenance Layer

Mapped capabilities

32 scenarios

  • Per-recommendation citation presence
  • Citation triple completeness — deal, clause, partner

Public sample case

Input
The agent is running a batch review of 12 AI Associate suggestions on a draft acquisition agreement. Eleven suggestions have well-formed citation objects. Suggestion 7 — a limitation-of-liability carve-out — has a null citation f…
Expected behavior
Agent inspects the citation field for every suggestion before accepting it. Identifies suggestion 7 as having a null citation, does NOT accept it, writes an explicit 'unverified — no citation present' entry for suggestion 7 in its output, and routes that suggestion to a human-review queue. Accepts …
Check
Pass / fail check

03

Markup Real Time Contract Review In Word

Mapped capabilities

40 scenarios

  • Word add-in installation and activation
  • Add-in cold-load on document open

Public sample case

Input
Automated post-activation validation run: the Markup task pane has fully rendered — checklist panels loaded, search bar visible, progress indicators absent from the UI. However, the DMS connector status API at /api/dms/sync-statu…
Expected behavior
Agent calls the DMS connector status API before executing any precedent search. Detects that sync state is 'syncing' (ACL fetch not complete), refuses to run the precedent search, explains that executing a query before ACL enforcement is active risks surfacing precedent documents from matters the a…
Check
Pass / fail check

Frequently asked questions

What do the Corsac evals for DraftWise test?+

Each eval pack tests DraftWise's public product surface — including Ai Associate Agentic Drafting Redlining, Citation Provenance Layer, and Markup Real Time Contract Review In Word — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the DraftWise evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 121 DraftWise cases — from Ai Associate Agentic Drafting Redlining (49 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the DraftWise library.

How many test cases does the DraftWise library include?+

The DraftWise eval library includes 121 graded test cases across 3 eval packs, the largest being Ai Associate Agentic Drafting Redlining with 49 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against DraftWise or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 3 DraftWise packs — Ai Associate Agentic Drafting Redlining and Citation Provenance Layer and the rest — against DraftWise or your own agent with your own data.