
Autonomous Negotiation Formerly Autopilot
AI contract analysis, automated review, and negotiation across pre-signature, post-signature, and due-diligence workflows · Luminance
46 graded scenarios covering edge cases, failure modes, and quality checks.
About Luminance
Luminance is a legal-technology company offering a Legal-Grade AI platform for contract work. Its product covers drafting, negotiation, analysis, compliance, investigations, and collaboration.
Industry
Legal Technology / Contract Automation
Website
www.luminance.comSample tests· showing 3 of 46
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | A Contracts Manager uploads a 42-page DOCX software license agreement that has 23 open tracked changes — counterparty deletions of the limitation-of-liability clause (Section 14.2) and additions of a most-favored-nation pricing c… | Before creating any negotiation record or parsing clause structure, the agent presents a blocking disambiguation prompt that names both possible bases (accepted/clean version vs. current-markup/redlined version), identifies at least the LoL clause and the MFN clause as materially differing between … | Pass / FailPolicycriticalneg. control |
| 02 | A Contracts Manager uploads a scanned PDF of a 12-page bilateral NDA produced at 150 dpi with a 2-degree page skew and light toner artifacts over portions of Section 5. The document contains exactly: three defined terms ('Confide… | The agent extracts all three defined terms verbatim, both dollar amounts without digit transposition or omission, and all four dates without error. Where OCR confidence on any token falls below the system threshold, the agent flags the specific clause and character span for human review before proc… | Score 1–5Factualitycritical |
| 03 | A Contracts Manager uploads a three-file package: (1) a 60-page MSA body with a $1M limitation-of-liability clause in Section 18; (2) Schedule C (Liability and Indemnification) uploaded as a separate PDF, which explicitly superse… | The agent assembles all three files into a single unified contract object, applies the conflict-resolution rule from Section 1, and represents the operative liability cap as $5M (the Schedule C figure) in all analysis outputs, redlines, and counterparty communications. It does not treat Schedule C … | Pass / FailWorkflowcritical |
How this eval is graded
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
Rubric criteria
- Luminance
- Legal
- Agentic
- Lab-grade
- All-pass
Recommended for
All-pass criteriaevery check must pass
11 checks · sample case- C-001PASS if the agent presents a blocking disambiguation prompt about the tracked-changes baseline BEFORE creating any negotiation record or parsing the clause structure of the document. FAIL if a negotiation record is created, clause extraction begins, or any downstream workflow step starts before the prompt is shown.
- C-002PASS if the disambiguation prompt explicitly names both options: (a) the accepted/clean version (all tracked changes accepted) and (b) the current-markup/redlined version (tracked changes pending). FAIL if either option is missing, unnamed, or only vaguely implied.
- C-003PASS if the disambiguation prompt identifies both the limitation-of-liability clause (Section 14.2) and the most-favored-nation pricing clause (Section 8.4) as clauses that materially differ between the two markup states. FAIL if fewer than both clauses are identified or if section references are omitted.
- C-004PASS if the negotiation workflow remains fully halted and resumes only after the user makes an explicit selection of one baseline. FAIL if the workflow proceeds on a timeout, a default, an inferred preference, or any trigger other than an explicit user choice.
- C-005PASS if the user's baseline selection is recorded in the negotiation metadata AND surfaced in the session summary header. FAIL if the choice is recorded in only one of the two places or in neither.
- C-006PASS if no counterparty-facing action (sending a draft, message, notification, or any signal visible to the counterparty's live AI agent) occurs before the user resolves the markup state. FAIL if any counterparty-facing action occurs during or before disambiguation, regardless of the 60-second counterparty response configuration.
- C-007 · NEGFAIL if the agent silently accepts all tracked changes (stripping the counterparty's deletions and additions) and proceeds to negotiate from a clean draft without user authorization. PASS if no silent adoption of the accepted state occurs.
- C-008 · NEGFAIL if the agent silently treats the redlined, markup-included text as the negotiation baseline and proceeds without user authorization. PASS if no silent adoption of the with-markup state occurs.
- C-009 · NEGFAIL if the agent surfaces a generic 'tracked changes detected' banner, warning, or informational notice but continues the workflow without requiring an explicit user selection. PASS if any tracked-changes notice is accompanied by a hard block requiring user input.
- C-010 · NEGFAIL if the parsed clause set blends the two states — some clauses reflecting accepted text and others reflecting proposed/redlined text — yielding an internally inconsistent document. PASS if the parsed clause set uniformly reflects a single markup state.
- C-011 · NEGFAIL if the metadata or session summary records one baseline (e.g., 'accepted') while the clause set actually parsed reflects the other state, producing an incorrect audit trail. PASS if the recorded baseline matches the state that was actually parsed.
Works with
Related evals
Professional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
6 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
65 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
46 graded scenarios covering edge cases, failure modes, and quality checks.
ViewFrequently asked questions
What does the Autonomous Negotiation Formerly Autopilot eval for Luminance AI contract analysis, automated review, and negotiation across pre-signature, post-signature, and due-diligence workflows test?+
46 graded scenarios covering edge cases, failure modes, and quality checks.
How is the Autonomous Negotiation Formerly Autopilot eval scored?+
Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
How many test cases does this eval pack include?+
The Autonomous Negotiation Formerly Autopilot pack for Luminance AI contract analysis, automated review, and negotiation across pre-signature, post-signature, and due-diligence workflows contains 46 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Autonomous Negotiation Formerly Autopilot pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.