
Eval directory
Evals for Legora
4 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Legora AI products.
About Legora
Legora (formerly Leya) is a collaborative AI platform for lawyers that brings research, drafting, and review into a single agentic workspace, used by law firms and in-house teams across Europe and North America. Its workflows run multi-step legal work — analyzing contracts, drafting documents, and reviewing across large document sets — grounded in a firm's own materials.
How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
4/4 packs
Scoring metrics
4/4 packs
Judge rubrics
4/4 packs
Use-case maps
0/4 packs
Pack context
4/4 packs
Available eval packs for Legora
4 packs ready to run.
Agent Orchestration Plan Execute Review Deliver Loop
44 graded scenarios covering edge cases, failure modes, and quality checks.
Authentication Sso Session Management
39 graded scenarios covering edge cases, failure modes, and quality checks.
Skills System
18 graded scenarios covering edge cases, failure modes, and quality checks.
Workflows Workflow Builder
9 graded scenarios covering edge cases, failure modes, and quality checks.
Why eval Legora AI
Legora's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Legora measures four dimensions teams care about most when deploying legal ai agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Legora's public product surface and runnable in Corsac with your own data.