Harvey
For HarveyLegal AIDoc Agent

Contract Review Negotiation Intelligence

Agentic legal work platform — Assistant, Vault, and Workflows for research, drafting, document review, and multi-step matter execution across 24+ practice areas · Harvey

10 graded scenarios covering edge cases, failure modes, and quality checks.

About Harvey

Harvey is an AI platform purpose-built for legal professionals, trusted by leading law firms and legal departments. It applies large language models to contract analysis, due diligence, legal research, and document drafting — all with law-firm-grade accuracy and confidentiality.

Employees

~200

Industry

Legal AI

Headquarters

San Francisco, CA

Website

harvey.ai

Sample tests· showing 3 of 10

Pass/fail checks, each adjudicated by an LLM judge.

#InputExpected behaviorCheck
01

Multi-jurisdiction comparison must not apply Texas non-compete rules to California agreements.

Side-by-side comparison applies Texas law analysis only to Texas-governed docs and explains California voidness policy for CA agreements with citations per row.

Pass / FailFactualityhigh
02

Liquidated damages enforceability varies by jurisdiction and reasonableness of estimate.

Harvey cites LD clause text, identifies jurisdiction from agreement, applies reasonableness framing, and flags penalty risk factors without computing damages amount absent inputs.

Pass / FailFactualitymedium
03

Anti-assignment clauses sometimes exempt change-of-control; agent must read exceptions.

Flags contracts where assignment restriction applies to change-of-control absent exception; cites exception text when present.

Pass / FailGroundinghigh

Unlock full benchmark

7 more test cases

Use this benchmark

How this eval is graded

Pass/fail checks, each adjudicated by an LLM judge.

Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

Rubric criteria

  • Harvey
  • Legal
  • Agentic
  • Lab-grade
  • All-pass

Recommended for

Agentic legal work platform — Assistant, Vault, and Workflows for research, drafting, document review, and multi-step matter execution across 24+ practice areasHarvey customers

All-pass criteriaevery check must pass

9 checks · sample case
  • C-001 · NEGPASS if every Texas-law enforceability conclusion is attributed solely to agreements whose governing-law clause specifies Texas; FAIL if any Texas non-compete enforceability rule or standard is applied to or cited in the analysis of a California-governed agreement.
  • C-002PASS if the response correctly characterizes California non-compete provisions as void under Cal. Bus. & Prof. Code § 16600 (or equivalent authority) and explains the policy rationale for that voidness; FAIL if California non-compete law is omitted or described without reference to the statutory voidness default.
  • C-003 · NEGPASS if the response provides separate, jurisdiction-specific enforceability assessments for Texas and California agreements; FAIL if the response assigns a single combined or blended enforceability score, rating, or conclusion that applies to both jurisdictions simultaneously.
  • C-004 · NEGPASS if the comparison explicitly identifies the governing law for each agreement or comparison row (e.g., a dedicated 'Governing Law' column or per-row label); FAIL if governing law attribution is absent from the comparison output.
  • C-005PASS if the output presents a structured side-by-side or tabular format that juxtaposes Texas and California agreement provisions row by row or column by column; FAIL if the response delivers only a narrative summary with no comparative structure distinguishing the two jurisdictions.
  • C-006PASS if each jurisdiction row or section in the comparison contains at least one specific legal citation (statute, regulation, or case law); FAIL if any jurisdiction row presents an enforceability conclusion without a supporting citation.
  • C-007 · NEGPASS if the response does not describe or imply that non-compete agreements are broadly or routinely enforceable under California law; FAIL if the response states, implies, or suggests that California courts enforce non-compete clauses without prominently acknowledging the statutory voidness default.
  • C-008 · NEGPASS if enforceability conclusions are stated in jurisdiction-specific terms tied to Texas or California law; FAIL if the response uses generic parametric language such as 'generally enforceable,' 'typically enforced,' or 'commonly upheld' without qualifying the statement to a specific jurisdiction.
  • C-009PASS if the Texas enforceability analysis references at least one Texas-specific enforceability requirement (e.g., reasonableness of duration, geographic scope, or legitimate business interest standard); FAIL if the Texas analysis contains no reference to Texas-specific enforceability elements or cites only generic common-law standards without Texas attribution.

Works with

Related evals

Frequently asked questions

What does the Contract Review Negotiation Intelligence eval for Harvey Agentic legal work platform — Assistant, Vault, and Workflows for research, drafting, document review, and multi-step matter execution across 24+ practice areas test?+

10 graded scenarios covering edge cases, failure modes, and quality checks.

How is the Contract Review Negotiation Intelligence eval scored?+

Pass/fail checks, each adjudicated by an LLM judge. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

How many test cases does this eval pack include?+

The Contract Review Negotiation Intelligence pack for Harvey Agentic legal work platform — Assistant, Vault, and Workflows for research, drafting, document review, and multi-step matter execution across 24+ practice areas contains 10 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Contract Review Negotiation Intelligence pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.