All evals
Ironclad

Eval directory · Legal AI

Evals for Ironclad

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Ironclad AI products.

About Ironclad

Ironclad is a digital contracting platform that connects people, processes, and data to manage business contracts. Its product scope includes contract lifecycle management and related workflow automation.

Employees

925

Industry

Contract Lifecycle Management Software

Headquarters

San Francisco, CA

Use the eval library for Ironclad

All 107 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Ironclad?

3 areas · 107 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Intake Third Party Paper Ingestion

Mapped capabilities

29 scenarios

  • Scanned/image-only PDF OCR extraction

Public sample case

Input
A 40-page vendor NDA arrives as a fax-originated TIFF-to-PDF at approximately 80 DPI. The scan renders the counterparty name 'Acme Corp.' as 'Acm€ Corp.' due to a character substitution at that resolution. The Intake Agent is ope…
Expected behavior
The agent detects low image resolution (at or below 150 DPI), assigns a low OCR confidence score to character-level fields, withholds or flags every extracted field that falls below a defined confidence threshold, and surfaces a human-review prompt that names which fields are uncertain and why — sp…
Check
Pass / fail check

02

Playbook Guided Ai Redlining Jurist Redlining Agent

Mapped capabilities

39 scenarios

  • Playbook upload — .docx format ingestion

Public sample case

Input
The org's NDA playbook stores its 'Walk-Away Terms' (unlimited liability, no IP assignment) exclusively in Word floating text boxes and footnotes — a common authoring pattern in law firm templates. The OOXML parser processes the …
Expected behavior
Agent does not treat a numeric rule count alone as proof of completeness. Before activating, it requests or displays the list of extracted clause categories and compares them against the known clause taxonomy (or the table of contents of the uploaded document). If high-risk categories such as limit…
Check
Pass / fail check

03

Risk Review Clause Extraction Property Population

Mapped capabilities

39 scenarios

  • Pre-built clause type identification across the 190+ library

Public sample case

Input
A counterparty SaaS agreement's Section 9 ('Limited Warranty and Disclaimer') is a four-sentence paragraph. The final sentence, in all-caps, reads: 'IN NO EVENT SHALL EITHER PARTY'S AGGREGATE LIABILITY ARISING OUT OF OR RELATED T…
Expected behavior
Agent identifies the final sentence of Section 9 as a Limitation of Liability clause, extracts the fee-based cap formula, reports the clause as present, and does not route the contract as low-risk or auto-approve without human review of the actual cap amount. The routing decision is made only after…
Check
Pass / fail check

Frequently asked questions

What do the Corsac evals for Ironclad test?+

Each eval pack tests Ironclad's public product surface — including Intake Third Party Paper Ingestion, Playbook Guided Ai Redlining Jurist Redlining Agent, and Risk Review Clause Extraction Property Population — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Ironclad evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 107 Ironclad cases — from Playbook Guided Ai Redlining Jurist Redlining Agent (39 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Ironclad library.

How many test cases does the Ironclad library include?+

The Ironclad eval library includes 107 graded test cases across 3 eval packs, the largest being Playbook Guided Ai Redlining Jurist Redlining Agent with 39 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Ironclad or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 3 Ironclad packs — Intake Third Party Paper Ingestion and Playbook Guided Ai Redlining Jurist Redlining Agent and the rest — against Ironclad or your own agent with your own data.