All evals
Luminance

Eval directory · Legal AI

Evals for Luminance

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Luminance AI products.

About Luminance

Luminance is a legal-technology company offering a Legal-Grade AI platform for contract work. Its product covers drafting, negotiation, analysis, compliance, investigations, and collaboration.

Industry

Legal Technology / Contract Automation

Use the eval library for Luminance

All 116 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Luminance?

3 areas · 116 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Assisted Negotiation In Microsoft Word Negotiate Negotiation Ai

Mapped capabilities

32 scenarios

  • Add-in installation and initial authentication
  • Add-in session persistence across Word restarts

Public sample case

Input
An autonomous negotiation agent is two hours into reviewing a 60-page MSA. The user's VPN briefly drops and recovers, but the silent token-refresh call to the Luminance auth endpoint times out during the outage window. The access…
Expected behavior
Upon receiving an empty or 401-equivalent response from the Luminance backend on the clause 39 call, the agent immediately pauses the redlining loop, does NOT emit any clause edits for clause 39 onward, does NOT transmit any pending or in-progress redlines to any counterparty channel, and surfaces …
Check
Pass / fail check

02

Autonomous Negotiation Formerly Autopilot

Mapped capabilities

46 scenarios

  • Contract ingestion — upload path

Public sample case

Input
A Contracts Manager uploads a 42-page DOCX software license agreement that has 23 open tracked changes — counterparty deletions of the limitation-of-liability clause (Section 14.2) and additions of a most-favored-nation pricing c…
Expected behavior
Before creating any negotiation record or parsing clause structure, the agent presents a blocking disambiguation prompt that names both possible bases (accepted/clean version vs. current-markup/redlined version), identifies at least the LoL clause and the MFN clause as materially differing between …
Check
Pass / fail check

03

Contract Drafting Draft

Mapped capabilities

38 scenarios

  • Template selection from catalog

Public sample case

Input
Legal ops deprecates 'Mutual NDA v2.1' after a deficiency in its limitation-of-liability clause is identified, replacing it with 'Mutual NDA v3.0'. v2.1 is marked 'deprecated:true' in catalog metadata but not hard-deleted (retain…
Expected behavior
The agent applies a hard exclusion filter on deprecated:true entries before any scoring step, selects 'Mutual NDA v3.0' for all 34 renewal drafts, and logs the template name and version identifier on each draft record. It does not use v2.1 for any draft, and the batch path applies the same deprecat…
Check
Pass / fail check

Frequently asked questions

What do the Corsac evals for Luminance test?+

Each eval pack tests Luminance's public product surface — including Ai Assisted Negotiation In Microsoft Word Negotiate Negotiation Ai, Autonomous Negotiation Formerly Autopilot, and Contract Drafting Draft — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Luminance evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 116 Luminance cases — from Autonomous Negotiation Formerly Autopilot (46 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Luminance library.

How many test cases does the Luminance library include?+

The Luminance eval library includes 116 graded test cases across 3 eval packs, the largest being Autonomous Negotiation Formerly Autopilot with 46 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Luminance or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 3 Luminance packs — Ai Assisted Negotiation In Microsoft Word Negotiate Negotiation Ai and Autonomous Negotiation Formerly Autopilot and the rest — against Luminance or your own agent with your own data.