All evals
D

Eval directory

Evals for Docsum

Eval coverage for Docsum, mapped from its public product surface.

About Docsum

Docsum is an AI contract repository, review, and negotiation platform for legal and sales teams. It centralizes agreements with configurable fields, playbooks, renewal alerts, and agentic search across the full repository, plus integrations with tools like Docusign, Google Drive, Slack, and Word. It is sold in Core, Advanced, and Enterprise tiers with a demo-first onboarding process and a 14-day free trial.

Industry

AI contract review and CLM (legal contract repository)

Use the eval library for Docsum

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Docsum?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Contract Repository & Data Model

The centralized agreement database and the structure imposed on it: standard and configurable fields, document types, and how related agreements connect. Grounded in the Core/Advanced tier descriptions, the document relationships release, and the bulk editing and automatic document grouping changelog entries.

AI contract repository and negotiation for legal and sales teams www.docsum.ai

Mapped capabilities

4 capabilities

  • Field extraction and document typing

    Populating standard and configurable fields from uploaded agreements and assigning the correct document type.

  • Document relationships and composite terms

    Linking MSAs, amendments, SOWs, and renewals, and resolving the effective term when a later document modifies an earlier one.

  • Bulk editing and automatic grouping

    Applying edits across many records and grouping related uploads without corrupting unrelated fields.

  • Ingestion paths and legacy migration

    Upload via email integration and migration of legacy agreement sets into the structured repository.

02

Agentic Repository Search & Citations

Plain-English questions asked across the full contract repository, with inline and deep-linked citations back to source clauses. Grounded in the Agentic Repository Search release and the inline citation, view-scoped chat, chat attachment, and live web research entries.

Agentic Repository Search, so you can ask questions of your entire repository in plain English www.docsum.ai

Mapped capabilities

4 capabilities

  • Repository-wide question answering

    Answering multi-contract questions in plain English across the agreement set rather than a single open document.

  • Inline and deep-linked citations

    Attaching clickable citations that resolve to the specific source clause supporting each claim.

  • View-scoped chat

    Honoring the scope of the current view so answers draw only from the intended subset of agreements.

  • Chat attachments and live web research

    Incorporating an attached document or live web research into an answer and distinguishing those sources from repository content.

Illustrative example

Input
Which of our vendor agreements renew automatically before March 31 and have no termination-for-convenience right? List each one.
Expected behavior
Returns only agreements that satisfy both conditions, with an inline citation deep-linking to the auto-renewal and termination language in each. If the repository does not support a determination for some agreement, it says so rather than inferring.

03

AI Contract Review & Negotiation

Reviewing agreements against playbooks and producing negotiation output. Grounded in the product's core positioning, configurable playbooks, AI-generated playbooks with alignment scoring, AI contract editing, and AI signing detection.

Mapped capabilities

4 capabilities

  • Playbook alignment scoring

    Comparing contract language to playbook positions and reporting where terms deviate.

  • AI-generated playbooks

    Deriving a playbook from existing agreements or stated standards for review use.

  • AI contract editing and redlines

    Proposing edits that move non-conforming clauses toward the playbook position.

  • Signing detection

    Identifying whether an agreement is executed and reflecting that status on the record.

Illustrative example

Input
Review this MSA against our playbook. The playbook caps liability at 12 months of fees; the MSA's limitation of liability section is unlimited for all claims.
Expected behavior
Flags the limitation of liability clause as out of alignment with the playbook cap, reflects that deviation in the alignment score, and proposes a redline setting the cap at 12 months of fees rather than accepting the clause as written.

05

Alerts, Renewals & Analytics

Time- and value-driven notifications over the repository plus reporting on it. Grounded in renewal alerts in Core, value-based alerts in the April 2026 release, the Advanced analytics dashboard, and Enterprise custom data reporting.

Mapped capabilities

4 capabilities

  • Renewal and expiry alerts

    Firing alerts on the correct dates derived from stored and amended contract terms.

  • Value-based alerts

    Triggering on contract value thresholds rather than dates alone.

  • Analytics dashboard

    Aggregate views over the agreement set that reconcile with the underlying records.

  • Custom reporting

    Enterprise-tier custom data reporting over configured fields and document types.

06

Integrations, Entitlements & Trust Boundaries

How Docsum connects to surrounding tools and where its limits sit. Grounded in the Docusign, Google Drive, and Slack integrations, the Word Add-in with unified login and deep-linked citations, tier contract volume limits, human-in-the-loop validation, and the stated SOC 2 Type II, CASA, and California Bar AI guidance posture.

California Bar AI Guidance Compliant www.docsum.ai

Mapped capabilities

4 capabilities

  • Docusign, Drive, and Slack integrations

    Syncing agreements and notifications across connected systems without divergence.

  • Word Add-in parity

    Delivering file uploads, web research, unified login, and deep-linked citations inside Word as on the web.

  • Tier entitlements and contract volume limits

    Enforcing Core and Advanced contract caps and feature gating across tiers.

  • Human-in-the-loop validation and legal-advice boundaries

    Routing to human validation where offered and staying within the product's stated compliance and guidance posture.

Coverage is mapped from Docsum's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Docsum test?+

The coverage map is generated from Docsum's own public product surface (AI contract review and CLM (legal contract repository)): 6 scoring areas — Contract Repository & Data Model, Agentic Repository Search & Citations, and AI Contract Review & Negotiation, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Docsum evals scored?+

Every case generated for Docsum — across Contract Repository & Data Model and Agentic Repository Search & Citations and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Docsum library include?+

The full Docsum library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Field extraction and document typing and Document relationships and composite terms under Contract Repository & Data Model); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Docsum or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Docsum areas and set them up in a Corsac workspace, where you can run every test case against Docsum or your own agent with your own data.