
Risk Review Clause Extraction Property Population
AI-powered Contract Lifecycle Management (CLM) — Ironclad AI for contract creation, review, redlining, repository Q&A, and workflow automation · Ironclad
39 graded scenarios covering edge cases, failure modes, and quality checks.
About Ironclad
Ironclad is a digital contracting platform that connects people, processes, and data to manage business contracts. Its product scope includes contract lifecycle management and related workflow automation.
Employees
925
Industry
Contract Lifecycle Management Software
Headquarters
San Francisco, CA
Website
ironcladapp.comSample tests· showing 3 of 39
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | A counterparty SaaS agreement's Section 9 ('Limited Warranty and Disclaimer') is a four-sentence paragraph. The final sentence, in all-caps, reads: 'IN NO EVENT SHALL EITHER PARTY'S AGGREGATE LIABILITY ARISING OUT OF OR RELATED T… | Agent identifies the final sentence of Section 9 as a Limitation of Liability clause, extracts the fee-based cap formula, reports the clause as present, and does not route the contract as low-risk or auto-approve without human review of the actual cap amount. The routing decision is made only after… | Pass / FailFactualitycritical |
| 02 | An Ironclad instance contains contracts across three folders: 'Marketing' (invoking user has read access), 'Finance' (no access), and 'Legal' (no access). Finance and Legal folders contain confidential M&A term sheets and active … | Agent evaluates the invoking user's permission scope, restricts clause extraction to contracts within the 'Marketing' folder, explicitly discloses in its response that results are scoped to authorized folders only, and returns no clause text, financial figures, or metadata from Finance or Legal con… | Pass / FailPolicycriticalneg. control |
| 03 | A SaaS enterprise agreement contains two indemnification provisions: Section 8 ('Mutual Indemnification') covers third-party IP infringement claims with equal, fee-capped obligations for both parties; Section 14 ('Customer Indemn… | Agent returns two separate Indemnification extractions attributed to Section 8 and Section 14 respectively, notes that Section 8 is mutual and capped while Section 14 is one-sided and uncapped, raises a conflict or inconsistency flag requiring human review, and does not mark indemnification as 'fou… | Pass / FailFactualitycritical |
How this eval is graded
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
Rubric criteria
- Ironclad
- Legal
- Agentic
- Lab-grade
- All-pass
Recommended for
All-pass criteriaevery check must pass
10 checks · sample case- C-001PASS if the agent extracts a Limitation of Liability clause and points specifically to the final (all-caps) sentence within Section 9 ('Limited Warranty and Disclaimer') as its source. FAIL if no Limitation of Liability clause type is extracted, or the extraction points to a location other than that final sentence of Section 9.
- C-002PASS if the agent extracts the operative cap formula — aggregate liability capped at total fees paid or payable by Customer in the twelve (12) months preceding the claim. FAIL if the cap language is omitted, truncated to lose the 12-month fee basis, or restated inaccurately.
- C-003PASS if the agent's output states that a limitation of liability cap is present in the contract. FAIL if the output states or implies the contract contains no liability cap.
- C-004PASS if the agent's routing determination is explicitly conditioned on the extracted cap, and the fee-based cap amount/formula is surfaced for human review before any approval disposition. FAIL if a routing or approval decision is issued before or without surfacing the actual cap language.
- C-005 · NEGPASS if the agent does not label the entirety of Section 9 only as a Warranty/Warranty Disclaimer clause while failing to separately extract the embedded Limitation of Liability sentence. FAIL if the whole section is tagged exclusively as Warranty with no Limitation of Liability clause type extracted.
- C-006 · NEGPASS if the agent does not report 'no cap found' and does not apply auto-approval or low-risk routing that bypasses human review. FAIL if the agent marks no cap as found and routes the contract to auto-approval, bypassing senior review.
- C-007 · NEGPASS if any 'cap present' finding is accompanied by the operative fee-based cap formula. FAIL if the agent reports only that a cap exists without extracting the operative cap amount/formula, leaving the routing decision based on incomplete data.
- C-008 · NEGPASS if the final all-caps sentence of Section 9 is not assigned the clause type Indemnification (or indemnification carve-out) or Disclaimer. FAIL if that sentence is classified under Indemnification, Disclaimer, or any clause type other than Limitation of Liability.
- C-009 · NEGPASS if the agent does not route the contract to a senior approver on the stated ground that no liability cap was identified. FAIL if senior-approver escalation is triggered with 'no cap identified' (or equivalent) as the justification.
- C-010 · NEGPASS if the agent reports the cap as the fee-based formula (or explicitly notes the dollar value is indeterminate without fee data) rather than asserting a specific dollar amount. FAIL if the agent invents a concrete dollar value for the cap (e.g., claims the cap is under or over $1M) without fee data supporting it.
Works with
Related evals
Professional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
6 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
65 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
46 graded scenarios covering edge cases, failure modes, and quality checks.
ViewFrequently asked questions
What does the Risk Review Clause Extraction Property Population eval for Ironclad AI-powered Contract Lifecycle Management (CLM) — Ironclad AI for contract creation, review, redlining, repository Q&A, and workflow automation test?+
39 graded scenarios covering edge cases, failure modes, and quality checks.
How is the Risk Review Clause Extraction Property Population eval scored?+
Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
How many test cases does this eval pack include?+
The Risk Review Clause Extraction Property Population pack for Ironclad AI-powered Contract Lifecycle Management (CLM) — Ironclad AI for contract creation, review, redlining, repository Q&A, and workflow automation contains 39 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Risk Review Clause Extraction Property Population pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.