
Agent Orchestration Plan Execute Review Deliver Loop
Collaborative agentic legal workspace (research, drafting, contract review, due-diligence; multi-jurisdiction) · Legora
44 graded scenarios covering edge cases, failure modes, and quality checks.
About Legora
Legora (formerly Leya) is a collaborative AI platform for lawyers that brings research, drafting, and review into a single agentic workspace, used by law firms and in-house teams across Europe and North America. Its workflows run multi-step legal work — analyzing contracts, drafting documents, and reviewing across large document sets — grounded in a firm's own materials.
Sample tests· showing 3 of 44
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | An associate uploads 40 target-company contracts from a cross-border acquisition whose governing-law clauses span Delaware, Ontario, and England & Wales. No governing law has been pre-tagged in the DMS. The associate instructs th… | The agent produces a plan in which a dedicated 'Identify governing law per contract' step appears before any step labeled or described as applying jurisdiction-specific change-of-control analysis. Each jurisdiction-specific analysis step references the governing-law step as an explicit hard prerequ… | Pass / FailWorkflowcritical |
| 02 | A multi-step compliance memo plan is executing for an Australian client. The firm has two legal research connectors: a US-primary connector defaulting to US federal/state law, and an AU/UK connector covering Commonwealth and UK j… | Agent selects the AU/UK legal research connector exclusively. Before invocation, it writes a routing rationale to the audit trail that names the jurisdiction tag 'AU' and the specific legislation (Fair Work Act 2009, Commonwealth) as the basis for the connector choice. If the AU/UK connector is una… | Pass / FailPolicycritical |
| 03 | In-house counsel at a multinational uploads 15 supplier agreements explicitly tagged in their metadata with governing law: 5 under Delaware law, 5 under English law, 5 under Singapore law. The tags are visible in the document met… | The plan contains three separately labeled jurisdiction-specific research steps — one each for Delaware, England & Wales, and Singapore — each of which explicitly names the governing-law subset it applies to (five contracts each). Each research step is followed by a contract-analysis step scoped to… | Pass / FailWorkflowcritical |
How this eval is graded
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
Rubric criteria
- Legora
- Legal
- Agentic
- Lab-grade
- All-pass
Recommended for
All-pass criteriaevery check must pass
9 checks · sample case- C-001PASS if the plan contains a discrete, explicitly named step dedicated solely to identifying the governing law of each contract (e.g., 'Identify governing law per contract' or equivalent). FAIL if no such dedicated step appears anywhere in the plan.
- C-002PASS if the governing-law identification step is assigned a sequence index strictly lower than every jurisdiction-specific change-of-control analysis step in the plan. FAIL if any jurisdiction-specific analysis step appears at the same or a lower sequence index than the governing-law identification step.
- C-003PASS if every jurisdiction-specific change-of-control analysis step in the plan explicitly references the governing-law identification step as a named dependency or hard prerequisite in the plan's dependency graph. FAIL if any such analysis step omits or does not name this dependency.
- C-004PASS if the plan includes separate or clearly differentiated analysis steps for each relevant jurisdiction class (e.g., Delaware, Ontario, England & Wales or materially equivalent groupings). FAIL if the plan provides only a single undifferentiated change-of-control analysis step that does not distinguish between governing jurisdictions.
- C-005PASS if the governing-law identification step is explicitly scoped to all 40 contracts or the full contract set without carve-outs. FAIL if the step is scoped to a subset of contracts or does not account for the entire document population described in the task.
- C-006 · NEGPASS if no version of the plan places the governing-law identification step and any jurisdiction-specific threshold analysis step within the same concurrent execution group or parallel batch. FAIL if the plan permits governing-law identification and change-of-control analysis to run simultaneously for any contract or contract batch.
- C-007 · NEGPASS if the plan does not contain a single undifferentiated step (e.g., 'Analyze all contracts for change-of-control triggers') that appears before or merges with the governing-law identification step. FAIL if such a merged or pre-positioned step exists and would allow substantive analysis to begin before governing law is determined.
- C-008 · NEGPASS if no analysis step in the plan assumes or applies a uniform governing law to all 40 contracts prior to the completion of the governing-law identification step. FAIL if any step explicitly or implicitly treats all contracts as subject to a single jurisdiction's standards before governing-law determination is complete.
- C-009 · NEGPASS if the governing-law identification step is designated as mandatory, required, or blocking in the plan with no language suggesting it can be skipped or treated as best-effort. FAIL if the step is labeled optional, advisory, recommended, or otherwise non-blocking, or if it lacks a dependency edge that would prevent downstream steps from executing without it.
Works with
Related evals
Professional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
6 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
65 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
46 graded scenarios covering edge cases, failure modes, and quality checks.
ViewFrequently asked questions
What does the Agent Orchestration Plan Execute Review Deliver Loop eval for Legora Collaborative agentic legal workspace (research, drafting, contract review, due-diligence; multi-jurisdiction) test?+
44 graded scenarios covering edge cases, failure modes, and quality checks.
How is the Agent Orchestration Plan Execute Review Deliver Loop eval scored?+
Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
How many test cases does this eval pack include?+
The Agent Orchestration Plan Execute Review Deliver Loop pack for Legora Collaborative agentic legal workspace (research, drafting, contract review, due-diligence; multi-jurisdiction) contains 44 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Agent Orchestration Plan Execute Review Deliver Loop pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.