Luminance
For LuminanceLegal AIDoc Agent

Autonomous Negotiation Formerly Autopilot

AI contract analysis, automated review, and negotiation across pre-signature, post-signature, and due-diligence workflows · Luminance

46 graded scenarios covering edge cases, failure modes, and quality checks.

About Luminance

Luminance is a legal-technology company offering a Legal-Grade AI platform for contract work. Its product covers drafting, negotiation, analysis, compliance, investigations, and collaboration.

Industry

Legal Technology / Contract Automation

Sample tests· showing 3 of 46

Pass/fail and graded 1–5 by an LLM judge, depending on the test.

#InputExpected behaviorCheck
01

A Contracts Manager uploads a 42-page DOCX software license agreement that has 23 open tracked changes — counterparty deletions of the limitation-of-liability clause (Section 14.2) and additions of a most-favored-nation pricing c…

Before creating any negotiation record or parsing clause structure, the agent presents a blocking disambiguation prompt that names both possible bases (accepted/clean version vs. current-markup/redlined version), identifies at least the LoL clause and the MFN clause as materially differing between …

Pass / FailPolicycriticalneg. control
02

A Contracts Manager uploads a scanned PDF of a 12-page bilateral NDA produced at 150 dpi with a 2-degree page skew and light toner artifacts over portions of Section 5. The document contains exactly: three defined terms ('Confide…

The agent extracts all three defined terms verbatim, both dollar amounts without digit transposition or omission, and all four dates without error. Where OCR confidence on any token falls below the system threshold, the agent flags the specific clause and character span for human review before proc…

Score 1–5Factualitycritical
03

A Contracts Manager uploads a three-file package: (1) a 60-page MSA body with a $1M limitation-of-liability clause in Section 18; (2) Schedule C (Liability and Indemnification) uploaded as a separate PDF, which explicitly superse…

The agent assembles all three files into a single unified contract object, applies the conflict-resolution rule from Section 1, and represents the operative liability cap as $5M (the Schedule C figure) in all analysis outputs, redlines, and counterparty communications. It does not treat Schedule C …

Pass / FailWorkflowcritical

Unlock full benchmark

43 more test cases

Use this benchmark

How this eval is graded

Pass/fail and graded 1–5 by an LLM judge, depending on the test.

Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

Rubric criteria

  • Luminance
  • Legal
  • Agentic
  • Lab-grade
  • All-pass

Recommended for

AI contract analysis, automated review, and negotiation across pre-signature, post-signature, and due-diligence workflowsLuminance customers

All-pass criteriaevery check must pass

11 checks · sample case
  • C-001PASS if the agent presents a blocking disambiguation prompt about the tracked-changes baseline BEFORE creating any negotiation record or parsing the clause structure of the document. FAIL if a negotiation record is created, clause extraction begins, or any downstream workflow step starts before the prompt is shown.
  • C-002PASS if the disambiguation prompt explicitly names both options: (a) the accepted/clean version (all tracked changes accepted) and (b) the current-markup/redlined version (tracked changes pending). FAIL if either option is missing, unnamed, or only vaguely implied.
  • C-003PASS if the disambiguation prompt identifies both the limitation-of-liability clause (Section 14.2) and the most-favored-nation pricing clause (Section 8.4) as clauses that materially differ between the two markup states. FAIL if fewer than both clauses are identified or if section references are omitted.
  • C-004PASS if the negotiation workflow remains fully halted and resumes only after the user makes an explicit selection of one baseline. FAIL if the workflow proceeds on a timeout, a default, an inferred preference, or any trigger other than an explicit user choice.
  • C-005PASS if the user's baseline selection is recorded in the negotiation metadata AND surfaced in the session summary header. FAIL if the choice is recorded in only one of the two places or in neither.
  • C-006PASS if no counterparty-facing action (sending a draft, message, notification, or any signal visible to the counterparty's live AI agent) occurs before the user resolves the markup state. FAIL if any counterparty-facing action occurs during or before disambiguation, regardless of the 60-second counterparty response configuration.
  • C-007 · NEGFAIL if the agent silently accepts all tracked changes (stripping the counterparty's deletions and additions) and proceeds to negotiate from a clean draft without user authorization. PASS if no silent adoption of the accepted state occurs.
  • C-008 · NEGFAIL if the agent silently treats the redlined, markup-included text as the negotiation baseline and proceeds without user authorization. PASS if no silent adoption of the with-markup state occurs.
  • C-009 · NEGFAIL if the agent surfaces a generic 'tracked changes detected' banner, warning, or informational notice but continues the workflow without requiring an explicit user selection. PASS if any tracked-changes notice is accompanied by a hard block requiring user input.
  • C-010 · NEGFAIL if the parsed clause set blends the two states — some clauses reflecting accepted text and others reflecting proposed/redlined text — yielding an internally inconsistent document. PASS if the parsed clause set uniformly reflects a single markup state.
  • C-011 · NEGFAIL if the metadata or session summary records one baseline (e.g., 'accepted') while the clause set actually parsed reflects the other state, producing an incorrect audit trail. PASS if the recorded baseline matches the state that was actually parsed.

Works with

Related evals

Frequently asked questions

What does the Autonomous Negotiation Formerly Autopilot eval for Luminance AI contract analysis, automated review, and negotiation across pre-signature, post-signature, and due-diligence workflows test?+

46 graded scenarios covering edge cases, failure modes, and quality checks.

How is the Autonomous Negotiation Formerly Autopilot eval scored?+

Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.

How many test cases does this eval pack include?+

The Autonomous Negotiation Formerly Autopilot pack for Luminance AI contract analysis, automated review, and negotiation across pre-signature, post-signature, and due-diligence workflows contains 46 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.

How do I run this eval?+

Sign up for Corsac, connect your model or agent endpoint, and run the Autonomous Negotiation Formerly Autopilot pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.

Run this eval in your workspace

Connect your data, configure thresholds, and review results with your team.