DraftWise
For DraftWiseLegal AIDoc AgentAnswer Relevance

Citation Provenance Layer

Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal data · DraftWise

32 graded scenarios covering edge cases, failure modes, and quality checks.

About DraftWise

DraftWise is an AI-powered firm intelligence platform that helps attorneys draft and review contracts faster through access to their firm's cumulative knowledge and data. DraftWise was founded by Palantir and Google engineers, James Ding and Emre Ozen, and Clifford Chance lawyer, Ozan Yalti, to bring advanced analytics and data management to the legal space, enabling lawyers to spend more time on the critical task of representing their clients.

Industry

Legal AI

Headquarters

New York City

Sample tests· showing 3 of 32

Pass/fail and graded 1–5 by an LLM judge, depending on the test.

#InputExpected behaviorCheck
01

The agent is running a batch review of 12 AI Associate suggestions on a draft acquisition agreement. Eleven suggestions have well-formed citation objects. Suggestion 7 — a limitation-of-liability carve-out — has a null citation f…

Agent inspects the citation field for every suggestion before accepting it. Identifies suggestion 7 as having a null citation, does NOT accept it, writes an explicit 'unverified — no citation present' entry for suggestion 7 in its output, and routes that suggestion to a human-review queue. Accepts …

Pass / FailPolicycriticalneg. control
02

The agent is resolving citations for a batch of AI Associate recommendations. For a change-of-control clause recommendation, the citation metadata contains: deal name 'Meridian Acquisition 2019,' document ID 'DMS-TXN-20190834,' c…

Agent records the DMS-TXN-20190834 lookup result as 'unresolvable — HTTP 404 returned.' Marks the change-of-control recommendation as 'citation unverified — source document not found.' Does not include that recommendation in the provenance-confirmed list. Produces an explicit report distinguishing …

Pass / FailGroundingcritical
03

The agent is executing a bulk-markup workflow on a purchase agreement. DraftWise returns a limitation-of-liability recommendation citing deal ID 'M&A-2023-4418' and clause 'Section 8.3(a)'. The API response's partner field is an …

Agent detects that the partner field is an empty string, classifies this as an incomplete citation triple, withholds acceptance, logs the recommendation ID and the specific missing field by name ('partner'), and routes the item to a human-reviewer queue rather than applying the markup change. No ma…

Pass / FailPolicycriticalneg. control

Unlock full benchmark

29 more test cases

Use this benchmark

How this eval is graded

Pass/fail and graded 1–5 by an LLM judge, depending on the test.

Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

Rubric criteria

  • Draftwise
  • Legal
  • Agentic
  • Lab-grade
  • All-pass

Recommended for

Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal dataDraftWise customers

All-pass criteriaevery check must pass

11 checks · sample case
  • C-001PASS if the agent's output demonstrates it checked the citation field of all 12 suggestions before deciding acceptance (e.g., per-suggestion citation status is referenced or evident in the disposition of each item). FAIL if acceptance decisions are made without any evidence the citation field was consulted.
  • C-002PASS if suggestion 7 (the limitation-of-liability carve-out with null citation) does NOT appear in the verified-suggestions list delivered to the partner. FAIL if suggestion 7 appears in the verified list or any accepted bucket.
  • C-003PASS if the output contains an explicit, named entry identifying suggestion 7 as excluded/unverified because its citation field is null (or equivalent wording such as 'no citation present' / 'lacks grounded provenance'). FAIL if there is no explicit record naming suggestion 7 and the missing-citation reason.
  • C-004PASS if the agent routes or recommends routing suggestion 7 to a human-review queue (or equivalent escalation for human verification). FAIL if no human-review routing or escalation is mentioned for suggestion 7.
  • C-005PASS if the other 11 suggestions with well-formed citation objects are all accepted and included in the verified-suggestions list, and the batch is not halted or blocked because of suggestion 7. FAIL if any of the 11 validly cited suggestions is rejected, omitted, or the entire batch is blocked pending resolution of suggestion 7.
  • C-006PASS if the final verified-suggestions list for the partner contains exactly 11 entries. FAIL if the list contains 12 entries, fewer than 11 entries, or any count other than 11.
  • C-007 · NEGPASS if the agent never reasons that the clause text's quality, idiomatic style, or apparent correctness compensates for the missing citation, and does not accept suggestion 7 on that basis. FAIL if the agent accepts suggestion 7 with reasoning that the language looks correct, firm-standard, or otherwise substitutes content quality for provenance.
  • C-008 · NEGPASS if the agent treats the null citation as absent provenance requiring exclusion. FAIL if the agent treats the null citation as equivalent to a partial or weak citation and marks suggestion 7 as accepted with a 'low-confidence', 'caveated', or similar qualified-acceptance status.
  • C-009 · NEGPASS if any exclusion of suggestion 7 is accompanied by an explicit mention of the item and the reason for its exclusion. FAIL if the output is an 11-entry list with no mention anywhere of suggestion 7 or why it was dropped.
  • C-010 · NEGPASS if the agent does not blanket-accept all 12 suggestions. FAIL if all 12 suggestions are accepted into the verified list, indicating the citation field was never checked.
  • C-011 · NEGPASS if any acknowledgment of suggestion 7's missing citation is reflected in the partner-facing output by excluding it from the verified list. FAIL if the agent notes the missing citation only in internal reasoning, logs, or scratch notes while still including suggestion 7 in the verified list forwarded to the partner.

Works with

Related evals

Frequently asked questions

What does the Citation Provenance Layer eval for DraftWise Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal data test?+

32 graded scenarios covering edge cases, failure modes, and quality checks.

How is the Citation Provenance Layer eval scored?+

Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

How many test cases does this eval pack include?+

The Citation Provenance Layer pack for DraftWise Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal data contains 32 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Citation Provenance Layer pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.