
Ask Citation Grounded Q A
Generative AI for transactional lawyers in Microsoft Word — contract drafting, review, redlining, and the agentic Spellbook Associate workflow · Spellbook
21 graded scenarios covering edge cases, failure modes, and quality checks.
About Spellbook
Spellbook is a legal-specific AI assistant for contract review, drafting, redlining, and document questions. Its workflows are delivered through a Microsoft Word add-in.
Industry
Legal Technology / Contract AI
Website
www.spellbook.legalSample tests· showing 3 of 21
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | An agent is operating Ask on a SaaS Master Services Agreement that has exactly 11 numbered sections (§1 through §11). The limitation of liability clause is at §9.4. The document has no §12 at all. The agent invokes Ask to ground … | Ask returns the verbatim text of §9.4, cites §9.4 by its correct identifier, and does not mention §12.3 or any other non-existent section. The citation is auditable: a grader can open the document, navigate to §9.4, and confirm the returned text matches character-for-character (modulo whitespace no… | Pass / FailFactualitycritical |
| 02 | An agent is building a negotiation memo from a contract under active redlining. §9.1 contains a tracked change: the original text 'ten million dollars ($10,000,000)' is struck through (tracked deletion) and the counterparty's pro… | Ask returns both the deleted amount ($10,000,000) and the proposed inserted amount ($2,000,000), clearly labels which is the original/deleted text and which is the proposed/pending revision, cites §9.1, and explicitly notes that the cap is subject to a pending tracked change that has not been accep… | Pass / FailFactualityhigh |
| 03 | An agent is performing a gap-analysis review pass on a 22-page technology services agreement. The document has been fully indexed. The agreement has no force majeure provision — it is not present in any section, exhibit, or sched… | Ask explicitly states that no force majeure provision was found in the document. The response does not provide any clause text, does not cite any section number, and does not describe what a force majeure clause 'typically' says in a way that could be mistaken for a document-specific finding. Optio… | Pass / FailFactualitycriticalneg. control |
How this eval is graded
Pass/fail and graded 1–5 by an LLM judge, depending on the test.
Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
Rubric criteria
- Spellbook
- Legal
- Agentic
- Lab-grade
- All-pass
Recommended for
All-pass criteriaevery check must pass
9 checks · sample case- C-001PASS if the response identifies §9.4 as the location of the limitation of liability clause and attributes the liability cap to that section. FAIL if the response attributes the liability cap to any section other than §9.4 or gives no section attribution.
- C-002PASS if the liability cap language returned is the verbatim text of §9.4 as it appears in the document, character-for-character modulo whitespace normalization. FAIL if the returned text is paraphrased, truncated in a way that changes meaning, or cannot be located verbatim at §9.4 in the document.
- C-003PASS if the citation uses the document's actual section identifier format (e.g., '§9.4') such that a grader can open the document, navigate directly to that section header, and verify the quote. FAIL if the citation is vague (e.g., 'the liability section', 'near the end of the agreement') or uses an identifier format that does not correspond to a locatable header in the document.
- C-004 · NEGPASS if the response never mentions §12.3 as a source, citation, or reference. FAIL if the response mentions §12.3 in any capacity that implies it exists in this agreement, including phrasing like 'Per §12.3, liability is capped at...'.
- C-005 · NEGPASS if every section identifier mentioned in the response exists in the document (the document contains only §1 through §11 and their actual subsections). FAIL if the response cites any section number outside §1–§11 or any subsection identifier that does not appear in the document, even if plausible-sounding.
- C-006 · NEGPASS if every monetary amount, multiplier (e.g., '12 months of fees'), or cap formula stated in the response appears in the actual text of §9.4. FAIL if the response states any dollar figure, multiplier, or cap formula that is not present in §9.4.
- C-007 · NEGPASS if all language presented as quoted or sourced from the agreement is actually present in §9.4. FAIL if the response presents fabricated clause language, carve-outs, exceptions, or defined terms as if they appear in §9.4 when they do not.
- C-008 · NEGPASS if the substantive answer about the liability cap is tied to a specific section citation. FAIL if the response describes a liability cap in generic terms (e.g., 'liability is typically capped at fees paid' or an uncited summary) without citing any section of this document.
- C-009 · NEGPASS if the response contains no confidently-stated claims about the document that the document does not support (this criterion grades tone-plus-content: any unsupported document claim delivered as fact). FAIL if the response asserts any unsupported or non-existent document content in confident, unhedged prose.
Works with
Related evals
Professional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
6 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
65 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
46 graded scenarios covering edge cases, failure modes, and quality checks.
ViewFrequently asked questions
What does the Ask Citation Grounded Q A eval for Spellbook Generative AI for transactional lawyers in Microsoft Word — contract drafting, review, redlining, and the agentic Spellbook Associate workflow test?+
21 graded scenarios covering edge cases, failure modes, and quality checks.
How is the Ask Citation Grounded Q A eval scored?+
Pass/fail and graded 1–5 by an LLM judge, depending on the test. The judge rubric: Grade the agent's response against EACH criterion in expected.criteria independently (PASS/FAIL per criterion, using each criterion's match_criteria). The case passes only if EVERY criterion passes (all-pass) — partial completion fails. For negative criteria (is_negative=true), PASS means the agent did NOT exhibit the described behavior.
How many test cases does this eval pack include?+
The Ask Citation Grounded Q A pack for Spellbook Generative AI for transactional lawyers in Microsoft Word — contract drafting, review, redlining, and the agentic Spellbook Associate workflow contains 21 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Ask Citation Grounded Q A pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.