All evals
Harvey

Eval directory · Legal AI

Evals for Harvey

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Harvey AI products.

About Harvey

Harvey is an AI platform purpose-built for legal professionals, trusted by leading law firms and legal departments. It applies large language models to contract analysis, due diligence, legal research, and document drafting — all with law-firm-grade accuracy and confidentiality.

Employees

~200

Industry

Legal AI

Headquarters

San Francisco, CA

Website

harvey.ai

Use the eval library for Harvey

All 179 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Harvey?

10 areas · 179 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Assistant Agentic Search And Iterative Source Expansion

Mapped capabilities

16 scenarios

  • Initial query formulation from user prompt

02

Assistant Citation Grounded Q A

Mapped capabilities

7 scenarios

  • Parametric vs Vault grounding disclosure
  • Citation-to-passage binding

03

Assistant Conversational Q A And Prompt Entry

Mapped capabilities

59 scenarios

  • Plain-text single-turn prompt submission

04

Assistant Deep Analysis And Long Form Memo Generation

Mapped capabilities

50 scenarios

  • Matter prompt decomposition into analysis plan

05

Confidentiality Privilege Tenant Isolation

Mapped capabilities

5 scenarios

  • ABA Model Rule 1.6 confidentiality
  • Privilege and work-product gates
  • Redaction before external export

06

Contract Review Negotiation Intelligence

Mapped capabilities

10 scenarios

  • Clause comparison across agreements
  • Issue spotting and risk flags
  • Representations and covenant survival

08

Security Compliance Data Residency

Mapped capabilities

7 scenarios

  • Regional data residency selection
  • Zero data retention with model providers
  • SOC 2 and ISO 27001 marketing claims

09

Vault Review Tables Matter Retrieval

Mapped capabilities

9 scenarios

  • Review table column extraction
  • Matter-scoped vault retrieval
  • Synthesis queries over review tables

10

Workflow Agents Governance

Mapped capabilities

6 scenarios

  • Conditional multi-step orchestration
  • Permissions sharing and external partners

Frequently asked questions

What do the Corsac evals for Harvey test?+

Each eval pack tests Harvey's public product surface — including Assistant Agentic Search And Iterative Source Expansion, Assistant Citation Grounded Q A, Assistant Conversational Q A And Prompt Entry — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Harvey evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Harvey library include?+

The Harvey eval library includes 179 graded test cases across 10 eval packs. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Harvey or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run every test case against Harvey or your own agent with your own data.