All evals
IV

Eval directory

Evals for Ivo

Eval coverage for Ivo, mapped from its public product surface.

About Ivo

Ivo is an enterprise contract intelligence platform that reviews, redlines, and extracts insights from contracts. It spans three products: Review (playbook-grounded redlining inside Microsoft Word and Google Docs), Intelligence (an AI-native contract repository with clustering, relationship mapping, and plain-English AI columns), and Assistant (an agentic AI legal associate answering natural-language questions grounded in your contracts and trusted legal databases). It positions itself as not a CLM, emphasizing insight extraction with no metatagging or long implementation.

Industry

AI contract review and contract intelligence for in-house legal teams

Website

www.ivo.ai

Use the eval library for Ivo

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Ivo?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Playbook-Grounded Redlining

Ivo Review's core claim: surgical redlines produced against the customer's playbook positions, delivered inside Microsoft Word and Google Docs. Covers whether proposed edits trace to a stated playbook position, whether jurisdiction-specific rules are applied after governing-law detection, and whether prompt-based review instructions are honored alongside standing playbook rules.

Ivo works as an add-in inside both Microsoft Word and Google Docs www.ivo.ai

Mapped capabilities

4 capabilities

  • Playbook position adherence in proposed redlines

    Redline suggestions reflect the customer's stated fallback and non-negotiable positions rather than generic market defaults.

  • Governing-law detection and region-specific rule application

    Ivo detects the agreement's governing law and applies the matching regional rule set to its positions.

  • Prompt-based review instructions alongside standing playbook

    Ad-hoc natural-language review instructions are honored without silently overriding playbook non-negotiables.

  • Playbook Builder drafting from executed agreements

    Positions drafted from a customer's own executed contracts are surfaced for human review before being saved.

02

Benchmarking Against Negotiation History

Ivo compares each deal against previously negotiated agreements and narrows the comparison set by contract type, governing law, party role, and industry. This area tests whether the comparison set is validly constrained and whether counterparty positions are correctly placed relative to the customer's own standards and outliers.

Ivo maintains industry-recognized certifications including SOC 2 Type II and ISO 27001. www.ivo.ai

Mapped capabilities

4 capabilities

  • Comparable-set narrowing by type, law, role, and industry

    Benchmark sets exclude agreements that fail any of the four stated comparability dimensions.

  • Standard-versus-outlier positioning of counterparty terms

    Counterparty clause positions are placed against the customer's historical distribution, not an external norm.

  • Deal context propagation into recommendations

    Stated deal value, urgency, and strategic weight shift recommended positions consistently across the review.

  • Behavior when comparable history is thin or absent

    Benchmark claims are withheld or qualified when too few comparable agreements exist.

03

Repository Intelligence and Clustering

Ivo Intelligence's promise of complete contract visibility without metatagging or predefined fields. Covers detection of contract relationships (amendments, restatements, addenda, superseding agreements), consolidation of base agreements with amendments into a governing composite, and plain-English AI columns extracted across the full library.

Mapped capabilities

4 capabilities

  • Contract relationship detection across amendments and restatements

    Amendment, addendum, restatement, and superseding relationships are identified and correctly directed.

  • Composite agreement assembly and governing-version selection

    Base plus amendments resolve to the currently operative terms rather than superseded language.

  • AI columns from plain-English prompts across the library

    Prompt-defined columns extract the requested term consistently across heterogeneous contract formats.

  • Clause deviation and outlier flagging across agreements

    Added, removed, and modified clauses are detected as deviations from the portfolio pattern.

Illustrative example

Input
An MSA caps liability at 12 months of fees. Amendment 1 raises the cap to 24 months. Amendment 2 restates the cap at 18 months. What is the current liability cap?
Expected behavior
The response states the operative cap is 18 months of fees, attributing it to Amendment 2 as the latest controlling instrument. It does not present the 12-month or 24-month figures as currently in force, though it may note them as superseded.

04

Agentic Assistant Task Execution

Ivo Assistant is positioned as an agentic legal associate that completes tasks from natural-language prompts in both Word and the Repository. This area covers multi-step task execution, cross-document comparison where language and formatting differ, and generation of editable Word reports and memos.

Mapped capabilities

4 capabilities

  • Multi-step task completion from a single plain-language prompt

    Compound instructions are carried through to completion without dropping requested sub-tasks.

  • Cross-document legal position comparison despite wording variance

    Substantively equivalent clauses are matched across differing language and formatting.

  • Editable report and memo generation in Word format

    Requested output format, structure, and scope are honored in the produced document.

  • Complex analytical queries beyond keyword or yes/no search

    Questions requiring legal-nuance judgment across many contracts return structured, answerable results.

05

Grounding, Citation, and Traceability

Ivo claims every Assistant response links to the exact clause or source and displays a reasoning trail, drawing on customer contracts, playbooks, and trusted legal databases. This area tests whether claims are anchored to real retrievable sources, whether the reasoning trail matches the answer, and whether the system abstains when the corpus does not support an answer.

Every response links to the exact clause or source and displays a clear reasoning trail. www.ivo.ai

Mapped capabilities

4 capabilities

  • Clause-level citation accuracy to the cited source

    Each linked citation resolves to text that actually supports the assertion made.

  • Reasoning trail consistency with the delivered answer

    The displayed reasoning trail reflects the sources and steps that produced the answer.

  • Abstention when the corpus does not support an answer

    Unsupported questions yield an explicit gap statement rather than a fabricated citation.

  • Separation of customer-contract evidence from legal-database sources

    Answers distinguish what came from the customer's agreements versus trusted external databases.

Illustrative example

Input
Which of our vendor agreements include a most-favored-nation pricing clause? (Repository contains no agreement with an MFN or equivalent pricing-parity provision.)
Expected behavior
The response states that no agreements in scope contain an MFN or equivalent pricing-parity clause. It cites no contract as containing such a provision and does not substitute a related term such as a price-increase cap or benchmarking clause as if it were MFN.

06

Access Scoping and Deployment Posture

Custom rooms isolate projects, acquisitions, and business units, and room-scoped queries restrict analysis to associated documents. Ivo also makes specific posture claims — SOC 2 Type II, ISO 27001, no metatagging, usable within a week of implementation, and explicitly not a CLM. This area covers boundary enforcement and accurate self-description.

Ivo does not require a long implementation period or extensive metatagging. www.ivo.ai

Mapped capabilities

4 capabilities

  • Room-scoped query boundary enforcement

    Analyses restricted to a room draw on no documents outside that room.

  • Isolation across acquisitions and business units

    Documents segregated by project or business unit do not leak across room boundaries.

  • Accurate self-description of scope and positioning

    Statements about CLM boundaries, metatagging, and implementation time match Ivo's published claims.

  • Handling of certification and trust-documentation requests

    Requests for SOC 2 or ISO 27001 artifacts route to the stated trust channel rather than asserting details.

Coverage is mapped from Ivo's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Ivo test?+

The coverage map is generated from Ivo's own public product surface (AI contract review and contract intelligence for in-house legal teams): 6 scoring areas — Playbook-Grounded Redlining, Benchmarking Against Negotiation History, and Repository Intelligence and Clustering, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Ivo evals scored?+

Every case generated for Ivo — across Playbook-Grounded Redlining and Benchmarking Against Negotiation History and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Ivo library include?+

The full Ivo library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Playbook position adherence in proposed redlines and Governing-law detection and region-specific rule application under Playbook-Grounded Redlining); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Ivo or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Ivo areas and set them up in a Corsac workspace, where you can run every test case against Ivo or your own agent with your own data.