All evals
CoCounsel (Thomson Reuters)

Eval directory · Legal AI

Evals for CoCounsel (Thomson Reuters)

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for CoCounsel (Thomson Reuters) AI products.

About CoCounsel (Thomson Reuters)

CoCounsel Legal is Thomson Reuters' AI legal product for research, analysis, and drafting. It is grounded in Westlaw and Practical Law content.

Industry

Legal Technology / Legal AI

Use the eval library for CoCounsel (Thomson Reuters)

All 117 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for CoCounsel (Thomson Reuters)?

3 areas · 117 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Cocounsel Deep Research Westlaw Grounded Retrieval

Mapped capabilities

6 scenarios

  • Research plan formulation from single query

Public sample case

Input
An associate is advising a California-resident software engineer whose former New York-headquartered employer is asserting breach of a signed non-disclosure and non-solicitation agreement. Both states are named in a single query …
Expected behavior
The research plan explicitly identifies both California and New York as implicated jurisdictions. It sequences a choice-of-law analysis step — examining which state's law governs under the contract's choice-of-law clause and conflict-of-laws principles — as step 1 or step 2, before any substantive …
Check
Pass / fail check

02

Cocounsel Guided Agentic Workflows

Mapped capabilities

65 scenarios

  • Workflow plan generation and display

Public sample case

Input
The underlying LLM returns a single streamed response containing both the structured plan JSON and the text of step 1's draft output (a facts-and-parties section). The UI layer and agent runtime receive this combined response. Th…
Expected behavior
The system parses the streaming response and displays only the structured plan to the user. The step-1 draft embedded in the LLM response is held in a buffer and not rendered to the user and not forwarded as input to step 2. A distinct, explicit user action (button click or equivalent confirmation)…
Check
Pass / fail check

03

Cocounsel Skills Library Skill Invocation

Mapped capabilities

46 scenarios

  • Skill catalog render completeness

Public sample case

Input
The 'Draft discovery request' skill has just been made visible in the catalog but the execution backend returns a 404 on the first real invocation due to a deployment timing gap. The agent has already collected matter parameters …
Expected behavior
Upon receiving a 404 from the 'Draft discovery request' backend, the agent halts the workflow and reports: (1) the specific skill that failed and the error type, (2) that this is a platform-side issue rather than a subscription or user error, (3) that matter parameters were collected but no interro…
Check
Pass / fail check

Frequently asked questions

What do the Corsac evals for CoCounsel (Thomson Reuters) test?+

Each eval pack tests CoCounsel (Thomson Reuters)'s public product surface — including Cocounsel Deep Research Westlaw Grounded Retrieval, Cocounsel Guided Agentic Workflows, and Cocounsel Skills Library Skill Invocation — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the CoCounsel (Thomson Reuters) evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 117 CoCounsel (Thomson Reuters) cases — from Cocounsel Guided Agentic Workflows (65 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the CoCounsel (Thomson Reuters) library.

How many test cases does the CoCounsel (Thomson Reuters) library include?+

The CoCounsel (Thomson Reuters) eval library includes 117 graded test cases across 3 eval packs, the largest being Cocounsel Guided Agentic Workflows with 65 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against CoCounsel (Thomson Reuters) or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 3 CoCounsel (Thomson Reuters) packs — Cocounsel Deep Research Westlaw Grounded Retrieval and Cocounsel Guided Agentic Workflows and the rest — against CoCounsel (Thomson Reuters) or your own agent with your own data.