All evals
Spellbook

Eval directory · Legal AI

Evals for Spellbook

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Spellbook AI products.

About Spellbook

Spellbook is a legal-specific AI assistant for contract review, drafting, redlining, and document questions. Its workflows are delivered through a Microsoft Word add-in.

Industry

Legal Technology / Contract AI

Use the eval library for Spellbook

All 116 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Spellbook?

3 areas · 116 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ask Citation Grounded Q A

Mapped capabilities

21 scenarios

  • Document-clause retrieval accuracy

Public sample case

Input
An agent is operating Ask on a SaaS Master Services Agreement that has exactly 11 numbered sections (§1 through §11). The limitation of liability clause is at §9.4. The document has no §12 at all. The agent invokes Ask to ground …
Expected behavior
Ask returns the verbatim text of §9.4, cites §9.4 by its correct identifier, and does not mention §12.3 or any other non-existent section. The citation is auditable: a grader can open the document, navigate to §9.4, and confirm the returned text matches character-for-character (modulo whitespace no…
Check
Pass / fail check

02

Draft Clause And Document Generation In Word

Mapped capabilities

44 scenarios

  • Single-clause generation from free-text prompt
  • Multi-clause generation from a single prompt

Public sample case

Input
A Word document based on a firm template has locked content controls marking unfilled placeholders — e.g., the text 'shall be governed by [GOVERNING LAW PLACEHOLDER]' in Section 18, where '[GOVERNING LAW PLACEHOLDER]' is a locked…
Expected behavior
The agent detects via Office.js that the current selection anchor is inside a locked content control. It refuses to insert any text and surfaces a clear, actionable error to the user — e.g., 'Cannot insert here: cursor is inside a protected field. Click into a blank paragraph between sections and t…
Check
Pass / fail check

03

Review Redline Risk Spotting And Tracked Changes Generation

Mapped capabilities

51 scenarios

  • Adverse-provision flagging — standard risk categories

Public sample case

Input
Section 12 contains an aggregate LOL cap: 'Vendor's liability shall not exceed $50,000 in aggregate for all claims.' Section 9 (indemnification) reads: 'Client shall indemnify, defend, and hold harmless Vendor from and against an…
Expected behavior
Agent: (1) flags the $50,000 LOL cap as a risk item; (2) separately flags Section 9 as a first-party indemnification—covering Vendor's own direct losses rather than third-party claims asserted against Vendor—citing the absence of any 'third-party claim' qualifier; (3) explicitly states that the fir…
Check
Pass / fail check

Frequently asked questions

What do the Corsac evals for Spellbook test?+

Each eval pack tests Spellbook's public product surface — including Ask Citation Grounded Q A, Draft Clause And Document Generation In Word, and Review Redline Risk Spotting And Tracked Changes Generation — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Spellbook evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 116 Spellbook cases — from Review Redline Risk Spotting And Tracked Changes Generation (51 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Spellbook library.

How many test cases does the Spellbook library include?+

The Spellbook eval library includes 116 graded test cases across 3 eval packs, the largest being Review Redline Risk Spotting And Tracked Changes Generation with 51 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Spellbook or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 3 Spellbook packs — Ask Citation Grounded Q A and Draft Clause And Document Generation In Word and the rest — against Spellbook or your own agent with your own data.