All evals
D

Eval directory

Evals for Dioptra

Eval coverage for Dioptra, mapped from its public product surface.

About Dioptra

Dioptra is an AI agent for legal contract review that generates redlines directly in Microsoft Word from custom playbooks, standard forms, past agreements, or prompts. It also distills playbooks from executed agreements, extracts terms at scale, and produces issues lists, risk summaries, and a clause library. It integrates with existing CLM, CRM, and P2P workflows and is sold in Starter, Teams, and Enterprise tiers; Icertis has acquired the company.

Industry

AI contract review and redlining

Headquarters

New York

Use the eval library for Dioptra

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Dioptra?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Redline Generation

Producing precise, minimal tracked-change redlines in Microsoft Word from a playbook, standard form, past agreement, example redline, or a plain prompt — on both the company's own paper and counterparty paper.

Generate precise redlines in Microsoft Word based on your custom playbook. www.dioptra.ai

Mapped capabilities

4 capabilities

  • Playbook-driven redlining

    Applies the correct playbook position to the matching clause, including fallback positions and non-negotiables.

  • Counterparty paper handling

    Redlines unfamiliar third-party structures where clause names, order, and defined terms differ from the standard form.

  • Minimal-edit discipline

    Leverages existing contract language rather than rip-and-replace; edits stay scoped to the identified issue.

  • Cross-reference and definition integrity

    Keeps defined terms, section cross-references, and dependent clauses consistent after an edit.

Illustrative example

Input
Counterparty MSA caps liability at fees paid in the prior 3 months. Playbook requires 12 months, with unlimited liability carved out for confidentiality breach. Redline the clause.
Expected behavior
The agent edits the existing cap language to a 12-month measure and adds the confidentiality carve-out, keeping the counterparty's surrounding sentence structure and defined terms intact instead of replacing the whole clause.

02

Playbook Distillation

Deriving AI-ready, custom playbooks from executed agreements, standard forms, and example redlines, so positions reflect what the organization has actually accepted.

Distill fully custom AI-ready playbooks from your executed agreements, standard forms and more. www.dioptra.ai

Mapped capabilities

4 capabilities

  • Position extraction from executed agreements

    Identifies preferred and fallback positions from what was signed rather than what was proposed.

  • Standard form ingestion

    Turns a template or standard form into structured playbook rules.

  • Conflict and inconsistency surfacing

    Flags where source agreements imply contradictory positions instead of silently choosing one.

  • Playbook traceability

    Ties each distilled rule back to the source document or clause it came from.

03

Term Search & Extraction

Searching and extracting terms across a repository of executed contracts at scale, for reporting, obligation tracking, and portfolio-level questions.

Mapped capabilities

4 capabilities

  • Targeted term retrieval

    Returns the governing clause for a requested term across heterogeneous contract formats.

  • Bulk extraction consistency

    Produces comparably structured output across many documents in one run.

  • Absence and non-answer handling

    Reports that a term is not present rather than fabricating a value.

  • Citation to source

    Points to the document and section supporting each extracted value.

Illustrative example

Input
Extract the assignment-on-change-of-control provision from 40 executed vendor agreements, 6 of which contain no assignment clause at all.
Expected behavior
The agent returns the governing clause and a source citation for the 34 agreements that have one, and marks the remaining 6 as not present rather than inferring a default or borrowing language from a similar contract.

04

Review Deliverables & Clause Library

Downstream artifacts that communicate a review: issues lists, risk summaries, and one-click clause insertions from a maintained clause library in Word.

Mapped capabilities

4 capabilities

  • Issues list generation

    Converts redlines and open points into a reviewable list of issues tied to specific clauses.

  • Risk summary framing

    Summarizes deviations and exposure at a level suited to a business or executive reader.

  • Clause library insertion

    Inserts the correct library clause into the document with formatting and numbering intact.

  • Deliverable–document consistency

    Keeps the issues list and risk summary aligned with the redlines actually present in the file.

05

Workflow & Integration Surfaces

How the agent fits existing legal workflows: the Microsoft Word add-in, prompting, the AI assistant chat, and integrations with CLM, CRM, and P2P systems.

Mapped capabilities

4 capabilities

  • Word add-in behavior

    Operates inside the document surface, preserving tracked changes, comments, and formatting.

  • Prompt-driven review

    Honors ad hoc natural-language instructions, including instructions that override a playbook default.

  • Assistant Q&A grounding

    Answers contract questions from the loaded documents rather than general legal knowledge.

  • CLM/CRM/P2P handoff

    Passes reviewed documents and extracted terms into connected systems of record.

06

Trust, Data & Policy Behavior

Commitments stated in the public materials that shape acceptable agent behavior: no training on customer data, SOC 2 Type II posture, tenancy tiers, and the boundary between drafting assistance and legal advice.

Unlimited Contract Reviews Out of the box Market Playbooks Out of the box Prompt Library No Training on your Data www.dioptra.ai

Mapped capabilities

4 capabilities

  • No-training-on-customer-data adherence

    Customer contract content is not reused outside the customer's own context.

  • Repository and tenancy boundaries

    Bring-Your-Own-Repository and team libraries stay scoped to the entitled tenant and seats.

  • Scope-of-advice boundary

    Frames output as reviewable work product and escalates judgment calls to counsel.

  • Confidence and escalation signaling

    Marks low-confidence or ambiguous positions instead of presenting them as settled.

Coverage is mapped from Dioptra's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Dioptra test?+

The coverage map is generated from Dioptra's own public product surface (AI contract review and redlining): 6 scoring areas — Redline Generation, Playbook Distillation, and Term Search & Extraction, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Dioptra evals scored?+

Every case generated for Dioptra — across Redline Generation and Playbook Distillation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Dioptra library include?+

The full Dioptra library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Playbook-driven redlining and Counterparty paper handling under Redline Generation); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Dioptra or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Dioptra areas and set them up in a Corsac workspace, where you can run every test case against Dioptra or your own agent with your own data.