All evals
Legora

Eval directory · Legal AI

Evals for Legora

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Legora AI products.

About Legora

Legora (formerly Leya) is a collaborative AI platform for lawyers that brings research, drafting, and review into a single agentic workspace, used by law firms and in-house teams across Europe and North America. Its workflows run multi-step legal work — analyzing contracts, drafting documents, and reviewing across large document sets — grounded in a firm's own materials.

Employees

~150

Industry

Legal AI

Headquarters

Stockholm, Sweden

Website

legora.com

Use the eval library for Legora

All 110 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Legora?

4 areas · 110 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agent Orchestration Plan Execute Review Deliver Loop

Mapped capabilities

44 scenarios

  • Initial task decomposition into ordered execution plan
  • Tool selection and routing decision per plan step

Public sample case

Input
An associate uploads 40 target-company contracts from a cross-border acquisition whose governing-law clauses span Delaware, Ontario, and England & Wales. No governing law has been pre-tagged in the DMS. The associate instructs th…
Expected behavior
The agent produces a plan in which a dedicated 'Identify governing law per contract' step appears before any step labeled or described as applying jurisdiction-specific change-of-control analysis. Each jurisdiction-specific analysis step references the governing-law step as an explicit hard prerequ…
Check
Pass / fail check

02

Authentication Sso Session Management

Mapped capabilities

39 scenarios

  • SAML 2.0 SP-initiated login flow

Public sample case

Input
The agent has a live pending AuthnRequest with ID '_req_7f3a2b' stored in shared Redis session state. A SAMLResponse arrives at the ACS endpoint with InResponseTo='_req_AAAAAAAA', an ID that was never issued by this Legora instan…
Expected behavior
ACS handler looks up '_req_AAAAAAAA' in the shared session state store, finds no matching entry, and rejects the response before creating any session. HTTP response is 4xx. Audit log records: the presented InResponseTo value, the NameID, the Issuer, and the fact that no matching AuthnRequest was fo…
Check
Pass / fail check

03

Skills System

Mapped capabilities

18 scenarios

  • Skill instruction body — special characters and encoding

Public sample case

Input
A threat actor with Skill-authoring credentials at a mid-sized firm crafts an instruction body for the 'NDA Review — MENA Counterparties' Skill. They insert U+202E (RIGHT-TO-LEFT OVERRIDE) immediately before the word 'disclose'. …
Expected behavior
The system detects the presence of U+202E (and any other Unicode bidi control characters: U+202A, U+202B, U+202C, U+202D, U+2066, U+2067, U+2068, U+2069, U+200F, U+200E, U+061C) in the instruction body at save time and either (a) rejects the save with an HTTP 4xx response whose body identifies the …
Check
Pass / fail check

04

Workflows Workflow Builder

Mapped capabilities

9 scenarios

  • Create workflow from scratch via builder UI

Frequently asked questions

What do the Corsac evals for Legora test?+

Each eval pack tests Legora's public product surface — including Agent Orchestration Plan Execute Review Deliver Loop, Authentication Sso Session Management, and Skills System — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Legora evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 110 Legora cases — from Agent Orchestration Plan Execute Review Deliver Loop (44 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Legora library.

How many test cases does the Legora library include?+

The Legora eval library includes 110 graded test cases across 4 eval packs, the largest being Agent Orchestration Plan Execute Review Deliver Loop with 44 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Legora or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 4 Legora packs — Agent Orchestration Plan Execute Review Deliver Loop and Authentication Sso Session Management and the rest — against Legora or your own agent with your own data.