All evals
W

Eval directory

Evals for Wordsmith

Eval coverage for Wordsmith, mapped from its public product surface.

About Wordsmith

Wordsmith is an AI platform for in-house legal teams that acts as a "legal front door," capturing, triaging, resolving and recording legal requests from across the business. It combines configurable AI agents, contract review against a company playbook, bulk contract reporting, multi-jurisdiction legal research, document blueprints, and a general legal assistant. Agents work inside existing tools like Slack, Teams and email, escalating to lawyers only when a matter needs human judgement.

Industry

in-house legal AI platform

Use the eval library for Wordsmith

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Wordsmith?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

03

Playbook-driven contract review

Inbound counterparty paper checked against the company playbook: deviations identified, risks explained against policy, redlines drafted, and revised versions reviewed for what changed.

Wordsmith reads the counterparty paper, identifies deviations from your standard terms, and automatically drafts the redlines for your review www.wordsmith.ai

Mapped capabilities

4 capabilities

  • Deviation detection against playbook

    Finding terms that depart from standard and fallback positions in counterparty paper.

  • Redline drafting and counter-positions

    Producing edits that reflect the company's stated fallback rather than generic market language.

  • Risk explanation grounded in policy

    Stating why a clause violates policy, drawing on prior negotiated positions.

  • Turn-by-turn diff review

    Focusing review on what changed since the previous version without re-opening agreed terms.

Illustrative example

Input
Counterparty MSA with unlimited liability for data breach. Playbook: liability capped at 12 months' fees, no carve-outs above 2x, escalate anything uncapped to counsel.
Expected behavior
The clause is flagged as a deviation from the 12-month cap, the response cites the playbook rule it breaches, and the matter is escalated to counsel rather than redlined and returned as routine.

04

Bulk contract reporting and extraction

Extraction of dates, clauses, financial terms and compliance requirements across large contract sets, delivered as a reviewable grid with verifiable citations and downstream actions.

Every answer links to its source with precise citations: exact clause, page number, and full context. www.wordsmith.ai

Mapped capabilities

4 capabilities

  • Field extraction across a corpus

    Pulling requested fields consistently across many agreements of varying format.

  • Citation to exact clause and page

    Linking each extracted value to the specific clause and page it came from.

  • Coverage and gap identification

    Identifying agreements missing a required provision rather than silently omitting them.

  • Handoff into drafting workflows

    Triggering blueprint or drafting actions from report results.

Illustrative example

Input
Extract the auto-renewal notice period from a 40-page vendor agreement where the term is stated in Section 9.2 and referenced again in an amendment.
Expected behavior
The extracted notice period matches the operative amended term, and the cell carries a citation to the clause and page it was taken from, with the superseded original not reported as current.

Coverage is mapped from Wordsmith's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Wordsmith test?+

The coverage map is generated from Wordsmith's own public product surface (in-house legal AI platform): 6 scoring areas — Legal front door: intake, triage and routing, Configurable AI legal agents, and Playbook-driven contract review, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Wordsmith evals scored?+

Every case generated for Wordsmith — across Legal front door: intake, triage and routing and Configurable AI legal agents and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Wordsmith library include?+

The full Wordsmith library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Request capture from chat and email and Triage classification and prioritisation under Legal front door: intake, triage and routing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Wordsmith or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Wordsmith areas and set them up in a Corsac workspace, where you can run every test case against Wordsmith or your own agent with your own data.