DraftWise
For DraftWiseLegal AIDoc Agent

Ai Associate Agentic Drafting Redlining

Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal data · DraftWise

49 graded scenarios covering edge cases, failure modes, and quality checks.

About DraftWise

DraftWise is an AI-powered firm intelligence platform that helps attorneys draft and review contracts faster through access to their firm's cumulative knowledge and data. DraftWise was founded by Palantir and Google engineers, James Ding and Emre Ozen, and Clifford Chance lawyer, Ozan Yalti, to bring advanced analytics and data management to the legal space, enabling lawyers to spend more time on the critical task of representing their clients.

Industry

Legal AI

Headquarters

New York City

Sample tests· showing 3 of 49

Pass/fail and graded 1–5 by an LLM judge, depending on the test.

#InputExpected behaviorCheck
01

A partner wants an analytical opinion on whether the indemnity cap is market before deciding to redline it. She uses 'review' deliberately to request analysis, not edits. The AI Associate is running in the Word add-in with write …

Agent classifies the instruction as a Q&A / analysis action. It returns a memo-style analysis comparing the cap against market norms for Series B SaaS deals, citing relevant precedents where available. Zero tracked changes, redlines, or suggested revisions are applied to or proposed for the documen…

Pass / FailTool usecritical
02

An associate preparing for a SaaS MSA negotiation pastes a single compound instruction covering three structurally distinct tasks: drafting a new addendum clause, revising a numerical cap, and running a gap analysis. The agent's …

Agent decomposes the input into exactly three discrete subtasks: (1) draft-new — GDPR data-processing addendum, target Section 14; (2) revise — indemnity cap, precedent pool = last 3 Tier-1 SaaS deals; (3) flag-missing — representations gap analysis against comparable precedents. Each subtask surfa…

Pass / FailWorkflowcritical
03

An associate submits a two-part instruction where the second task (updating the Definitions section) logically depends on the first (inserting a new clause at Section 12). The agent's default parallelism setting would dispatch bo…

Agent detects the sequential dependency: the Definitions update can only succeed after the DPA clause exists at Section 12. It serializes execution — completing and confirming the Section 12 insertion first, then running the Definitions update against the post-insertion document state. The executio…

Pass / FailTool usecritical

Unlock full benchmark

46 more test cases

Use this benchmark

How this eval is graded

Pass/fail and graded 1–5 by an LLM judge, depending on the test.

Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

Rubric criteria

  • Draftwise
  • Legal
  • Agentic
  • Lab-grade
  • All-pass

Recommended for

Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal dataDraftWise customers

All-pass criteriaevery check must pass

10 checks · sample case
  • C-001PASS if the agent treats 'Review the indemnity section...' as an analytical/Q&A request and routes it to a read-only analysis action. FAIL if the execution trace shows the instruction was classified or dispatched as a revise/redline/drafting action.
  • C-002PASS if the response contains substantive analytical text assessing whether the indemnity cap is market for a Series B SaaS acquisition (e.g., comparison against market norms, citing precedents where available). FAIL if the response omits the requested market analysis or replaces it with edit proposals.
  • C-003PASS if the agent's deliverable is plain text/memo output with no embedded proposed revisions, suggested edits, or change previews framed as actions to apply. FAIL if the deliverable includes anything other than read-only analytical text (e.g., an 'apply changes' artifact or pending edit object).
  • C-004PASS if the active document's post-action state is byte-for-byte identical to its pre-action state. FAIL if the document differs from its pre-action state in any way.
  • C-005 · NEGFAIL if the agent produces tracked-change markup, redlines, or insert-revision operations on the indemnity clause (i.e., misclassifies 'review' as 'revise'). PASS if zero tracked changes or redlines are generated.
  • C-006 · NEGFAIL if the agent returns analytical text alongside unsolicited redlines or suggested revisions, presenting both as if both were requested. PASS if the analysis is delivered without any accompanying edit artifacts.
  • C-007 · NEGFAIL if any unsolicited markup is committed such that, in an auto-accept-changes Word environment, it would be silently merged into the clean document before the partner sees it. PASS if no markup is committed to the document at all.
  • C-008 · NEGFAIL if the agent defaults to revise/edit behavior without first surfacing a clarifying question about whether edits were intended. PASS if the agent either correctly performs analysis-only or, at minimum, asks about intent before any edit behavior.
  • C-009 · NEGPASS if the execution trace contains zero calls to document-write, insert-revision, track-change, or any other document-mutation tools. FAIL if any document-mutation tool call appears in the trace, regardless of whether it succeeded or was later undone.
  • C-010 · NEGPASS if any cited precedents or market-norm benchmarks are presented only where available/supported, with appropriate hedging when data is unavailable. FAIL if the analysis asserts specific precedent deals, statistics, or market benchmarks that are invented or presented as verified without support.

Works with

Related evals

Frequently asked questions

What does the Ai Associate Agentic Drafting Redlining eval for DraftWise Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal data test?+

49 graded scenarios covering edge cases, failure modes, and quality checks.

How is the Ai Associate Agentic Drafting Redlining eval scored?+

Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

How many test cases does this eval pack include?+

The Ai Associate Agentic Drafting Redlining pack for DraftWise Transactional drafting and negotiation AI — generates and redlines contract language from a firm's own precedent and prior deal data contains 49 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Ai Associate Agentic Drafting Redlining pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.