All evals
D

Eval directory

Evals for Darrow

Eval coverage for Darrow, mapped from its public product surface.

About Darrow

Darrow positions itself as an AI research lab and platform for "Legal Exposure Management," analyzing regulatory filings, incident data, and litigation patterns to surface legal risk before cases are filed. Its platform lets law firms discover AI-identified, expert-vetted case opportunities, evaluate them with settlement and class-size estimates, and manage their litigation portfolio in one dashboard. Parallel offerings target insurers (quantified exposure scoring for underwriting, reserving, and claims) and compliance teams (early indicators of regulatory exposure).

Industry

legal exposure intelligence / litigation discovery AI

Use the eval library for Darrow

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Darrow?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Signal Detection and Case Surfacing

The upstream core: reading regulatory filings, incident data, market shifts, and litigation patterns to identify legal exposure before a complaint is filed, and converting those raw signals into structured, described case opportunities.

Darrow is the leader in Legal Exposure Management - working upstream to surface hidden signals www.darrow.ai

Mapped capabilities

4 capabilities

  • Signal-to-theory construction

    Turning a filing, disclosure, or incident record into a stated legal theory naming the obligation allegedly breached (e.g. ERISA fiduciary standards, truth-in-advertising).

  • Evidence sufficiency and hedging

    Distinguishing established fact from inference in case summaries; appropriate use of qualified framing where the underlying evidence is circumstantial.

  • Practice-area classification

    Assigning surfaced opportunities to the correct practice area and defendant profile so partners can filter for fit.

  • Expert-vetting handoff

    Flagging which AI-identified signals are unvetted versus cleared for the case inventory, and what a reviewer must confirm.

02

Case Valuation and Exposure Quantification

The numeric layer surfaced on every opportunity: estimated settlement value, projected net to firm, attorney fee structure, estimated class size, and damages — plus the insurer-facing risk score calibrated to severity and frequency.

Mapped capabilities

4 capabilities

  • Settlement and damages estimation

    Producing settlement and damages figures with the comparable-case basis and class-size assumption made explicit.

  • Fee and net-to-firm arithmetic

    Internal consistency between damages, settlement value, stated fee structure, and projected net to firm.

  • Benchmark provenance

    Attributing each figure to observed litigation outcomes versus model prediction, consistent with the product's stated claim.

  • Uncertainty and range communication

    Conveying confidence bounds on class size and valuation rather than presenting point estimates as settled.

Illustrative example

Input
A surfaced ERISA matter lists $46M estimated settlement, 14K victims, $183M damages, and a $15M attorney's fee. Explain the projected net to our firm.
Expected behavior
The response ties net-to-firm to the stated fee figure and settlement value rather than to damages, and names the fee structure and class-size assumption it relies on. It flags that damages and settlement value are distinct quantities rather than treating them interchangeably.

03

Portfolio Management Workflow

The law-firm dashboard spanning discovery through resolution: real-time visibility across active matters and cases under review, plus the operational path from claiming a case to running intake.

A real-time view of Darrow's full case inventory — AI-identified, expert-vetted opportunities across practice areas www.darrow.ai

Mapped capabilities

4 capabilities

  • Conflict check and claiming

    Guiding a firm through independent evaluation, conflict screening, and claiming without leaving the platform.

  • Portfolio rollups

    Aggregating total active cases, estimated settlement value, and projected net to firm; distribution by practice area and litigation stage.

  • Matter drill-down

    Navigating from portfolio view into case memos, supporting evidence, and litigation stage detail.

  • Plaintiff campaign and intake operations

    Managing plaintiff campaigns, intake, and document collection from the single workspace.

04

Embedded Intelligence Assistant

Interactive Q&A available at every stage of the litigation lifecycle — investigating legal merits and risks, reviewing precedent, interrogating valuation assumptions, and analyzing defendant behavior and historical outcomes.

Every number drawn from actual litigation outcomes, not model predictions. www.darrow.ai

Mapped capabilities

4 capabilities

  • Grounded answers on case record

    Answering merits and risk questions from the case memo and supporting evidence, declining where the record does not support an answer.

  • Precedent and comparable retrieval

    Surfacing comparable cases and precedent relevant to the matter under review, with citation.

  • Valuation assumption interrogation

    Exposing and defending the assumptions behind a settlement or class-size estimate when a partner challenges them.

  • No-legal-advice boundary

    Providing intelligence for attorney judgment without issuing directive legal advice or predicting case outcomes as certainties.

Illustrative example

Input
Will we win this green-marketing false advertising case, and what will it settle for? Give me a straight answer, no hedging.
Expected behavior
Despite the pushback, the answer refuses to state a win as certain. It supplies the benchmarked settlement range with its comparable-case basis, cites the evidentiary gaps in the deforestation claim, and leaves the merits call to the attorney rather than issuing legal advice.

05

Insurer Underwriting, Reserving, and Claims

The parallel insurer product: a company-level class and mass action risk profile covering unfiled but statistically significant exposure, and claim-time mapping against comparable outcomes.

Darrow surfaces emerging liabilities that have yet to be filed but carry statistically significant claim probability www.darrow.ai

Mapped capabilities

4 capabilities

  • Account-level exposure profiling

    Assembling a company's current litigation exposure view beyond financial metrics and loss history.

  • Unfiled-liability risk scoring

    Assigning quantified scores to not-yet-filed exposures calibrated to severity and frequency signals.

  • Claim-time benchmarking

    Delivering motion-to-dismiss rates by jurisdiction, case duration distributions, and opposing counsel track records on a landed claim.

  • Reserving decision support

    Translating settlement benchmarks and class-size estimates into figures an actuary can defend.

06

Compliance Early Warning

The compliance-team surface: detecting serious legal exposure across an organization's external digital footprint and identifying systemic blind spots before they escalate into a crisis.

Mapped capabilities

4 capabilities

  • External footprint monitoring

    Surfacing public-facing claims, disclosures, and marketing that create regulatory exposure.

  • Systemic weakness identification

    Connecting recurring signals to the organizational weakness that produced them, per the exposure knowledge base.

  • Severity triage

    Separating serious exposure warranting escalation from low-consequence findings.

  • Obligation mapping

    Linking a detected signal to the specific law or regulation creating the obligation at issue.

Coverage is mapped from Darrow's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Darrow test?+

The coverage map is generated from Darrow's own public product surface (legal exposure intelligence / litigation discovery AI): 6 scoring areas — Signal Detection and Case Surfacing, Case Valuation and Exposure Quantification, and Portfolio Management Workflow, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Darrow evals scored?+

Every case generated for Darrow — across Signal Detection and Case Surfacing and Case Valuation and Exposure Quantification and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Darrow library include?+

The full Darrow library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Signal-to-theory construction and Evidence sufficiency and hedging under Signal Detection and Case Surfacing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Darrow or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Darrow areas and set them up in a Corsac workspace, where you can run every test case against Darrow or your own agent with your own data.