
Assistant Deep Analysis And Long Form Memo Generation
Agentic legal work platform — Assistant, Vault, and Workflows for research, drafting, document review, and multi-step matter execution across 24+ practice areas · Harvey
50 graded scenarios covering edge cases, failure modes, and quality checks.
About Harvey
Harvey is an AI platform purpose-built for legal professionals, trusted by leading law firms and legal departments. It applies large language models to contract analysis, due diligence, legal research, and document drafting — all with law-firm-grade accuracy and confidentiality.
Sample tests· showing 3 of 50
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | A partner submits a single matter prompt that explicitly lists eight distinct legal questions for a complex employment dispute. The agent must decompose the prompt into an analysis plan before initiating any vault retrieval. This… | The decomposition plan contains exactly eight labeled sub-issues, one corresponding to each stated question. No issue is silently merged with another or omitted. The plan is output as a visible, inspectable artifact before any retrieval operation begins. The count of plan entries is either stated e… | Pass / FailWorkflowcritical |
| 02 | An associate submits a contract review prompt where the governing law clause in the attached master services agreement has been redacted by the client before upload. The agent must decompose the matter into an analysis plan. The … | The plan explicitly flags that the governing law clause is unavailable and identifies all jurisdiction-dependent issues (at minimum issues 1, 2, and 4) as having an unresolved dependency. The plan does not name any specific assumed jurisdiction (e.g., New York, Delaware, California) and does not pr… | Pass / FailPolicycriticalneg. control |
| 03 | A matter prompt explicitly names nine distinct issues for a commercial real estate lease dispute. This eval tests whether an internal issue-count limit silently truncates late-listed issues (e.g., issues 6–9) without warning. The… | The plan lists all nine issues as separate, labeled entries. The count is either stated explicitly or verifiable from the plan's structure. Issues 6 through 9 are present with the same level of specificity as issues 1 through 5. If any internal system constraint would limit the number of issues, th… | Pass / FailFactualitycritical |
How this eval is graded
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
Rubric criteria
- Harvey
- Legal
- Agentic
- Lab-grade
- All-pass
Recommended for
All-pass criteriaevery check must pass
11 checks · sample case- C-001PASS if the decomposition plan contains exactly eight labeled sub-issue entries, one corresponding to each of the eight questions stated in the user prompt. FAIL if the plan contains fewer than eight entries for any reason.
- C-002 · NEGPASS if the complete decomposition plan is output as a visible, inspectable artifact before any document retrieval, vault query, or attachment read operation is initiated. FAIL if any retrieval operation begins before the plan is presented to the user.
- C-003 · NEGPASS if enforceability of the non-compete under Delaware law (issue 1) and enforceability under California law (issue 2) are represented as two distinctly labeled, separate entries in the plan. FAIL if they are merged into a single 'non-compete enforceability' or jurisdictionally undifferentiated entry.
- C-004 · NEGPASS if the likelihood of obtaining a TRO based on the employee declaration (issue 5) appears as its own labeled entry in the plan. FAIL if this issue is absent from the plan, including when the omission occurs because the issue is attachment-dependent rather than a pure legal question.
- C-005 · NEGPASS if the choice-of-law analysis for the non-solicitation clause (issue 8) appears as a distinct labeled entry separate from the independent enforceability of the non-solicitation clause (issue 3). FAIL if issue 8 is folded into issue 3 without explicit disclosure of the merger.
- C-006 · NEGPASS if issues listed toward the end of the user prompt — damages methodology (issue 6), arbitration clause (issue 7), and choice-of-law for non-solicitation (issue 8) — are all present in the plan. FAIL if one or more of these late-listed issues is absent while earlier issues are included.
- C-007PASS if the independent enforceability of the non-solicitation clause contingent on the non-compete being voided (issue 3) appears as its own labeled entry distinct from either non-compete enforceability entry. FAIL if it is absent or subsumed under issue 1 or issue 2.
- C-008PASS if the scope of the trade secrets claim under the Defend Trade Secrets Act (issue 4) appears as a distinct labeled entry. FAIL if it is absent or merged with another issue without disclosure.
- C-009PASS if the damages calculation methodology for lost client revenue (issue 6) appears as a distinct labeled entry in the plan. FAIL if it is absent or merged with another issue.
- C-010PASS if the question of whether the arbitration clause compels arbitration of this dispute (issue 7) appears as a distinct labeled entry in the plan. FAIL if it is absent or merged with another issue.
- C-011PASS if the number of plan entries is either stated explicitly as a count or can be unambiguously determined by counting visually discrete, labeled items in a structured list. FAIL if the plan is rendered as undifferentiated prose that makes counting plan entries ambiguous without inference.
Works with
Related evals
Professional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
6 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
65 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
46 graded scenarios covering edge cases, failure modes, and quality checks.
ViewFrequently asked questions
What does the Assistant Deep Analysis And Long Form Memo Generation eval for Harvey Agentic legal work platform — Assistant, Vault, and Workflows for research, drafting, document review, and multi-step matter execution across 24+ practice areas test?+
50 graded scenarios covering edge cases, failure modes, and quality checks.
How is the Assistant Deep Analysis And Long Form Memo Generation eval scored?+
Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
How many test cases does this eval pack include?+
The Assistant Deep Analysis And Long Form Memo Generation pack for Harvey Agentic legal work platform — Assistant, Vault, and Workflows for research, drafting, document review, and multi-step matter execution across 24+ practice areas contains 50 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Assistant Deep Analysis And Long Form Memo Generation pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.