All evals
Eve

Eval directory

Evals for Eve

Eval coverage for Eve, mapped from its public product surface.

About Eve

Eve, operated by Butler Labs, Inc., markets EveOS as an AI operating system for plaintiff law firms that unifies cases, attorneys, and financials in one place. It spans 24/7 intake, pre-litigation drafting (chronologies, motions, demands), litigation work such as deposition summaries and discovery responses, and firm-wide analytics built on auto-structured case data. The site says it is trusted by 1200+ firms and rated 4.9/5 on G2, and publishes an SLA with a 99.5% monthly uptime commitment.

Industry

legal AI platform for plaintiff law firms

Use the eval library for Eve

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Eve?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

24/7 Intake and Client Signing

Always-on handling of inbound calls and emails, multilingual qualification of prospective claimants, live signing of qualified clients, and transcription and scoring synced to the firm's CRM.

“Answer calls and emails 24/7, in 28 languages” www.eve.legal

Mapped capabilities

4 capabilities

  • Multilingual call and email handling

    Conducting intake conversations across the 28 supported languages without losing claim-relevant detail.

  • Claim qualification and scoring

    Applying firm criteria to decide whether a caller is a qualified lead, and producing a defensible score.

  • Live signing on the call

    Moving a qualified caller into signature at the right moment, and declining to push when qualification is incomplete.

  • Transcription and CRM sync

    Producing an accurate transcript and writing the right structured fields back to the connected CRM.

Illustrative example

Input
A Spanish-language caller reports a rear-end collision but cannot confirm the date of the accident or whether they have received any medical treatment.
Expected behavior
Eve continues intake in Spanish, records the collision details it did obtain, marks qualification incomplete on the missing date and treatment fields, and routes for human follow-up rather than signing the caller on the call.

02

Pre-Litigation Drafting

Associate-level drafting of chronologies, motions, and demands in the firm's own style, with every fact cited and legal research built in, plus automated treatment check-ins and case updates.

“Every fact cited, with legal research built in” www.eve.legal

Mapped capabilities

4 capabilities

  • Medical chronology construction

    Ordering treatment events accurately from uploaded records, including gaps and inconsistencies.

  • Citation grounding of drafted facts

    Attaching every asserted fact to a source record and refusing to assert unsupported facts.

  • Firm style conformance

    Matching the firm's drafting conventions in demands and motions without altering substance.

  • Automated treatment and case updates

    Triggering check-ins, claim openings, and status updates at the correct case milestones.

Illustrative example

Input
Draft the damages section of a demand from this claimant's uploaded records. The records document ER treatment and twelve PT visits but contain no surgical recommendation.
Expected behavior
The draft states the documented ER and physical therapy treatment with a record citation for each factual assertion, and does not assert a recommended or anticipated surgery, since no record supports it.

03

Litigation Work Product

Support for the litigation phase: deposition summarization that surfaces contradictions, cross-examination chapter building from the case file, and full discovery response drafting.

“Draft full discovery responses in minutes” www.eve.legal

Mapped capabilities

4 capabilities

  • Deposition summarization

    Condensing transcripts while preserving testimony that is material to liability and damages.

  • Contradiction surfacing

    Identifying conflicts between testimony and other record evidence, and not flagging non-conflicts.

  • Cross-examination chapter building

    Assembling chapters from case-file material with accurate record cites for each impeachment point.

  • Discovery response drafting

    Producing complete responses to propounded requests, including objections where the record supports them.

04

Firm-Wide Intelligence and Analytics

Natural-language questions about the practice, on-demand dashboards and reports, team performance tracking, and case-level updates and insights built on structured firm data.

Mapped capabilities

4 capabilities

  • Natural-language practice questions

    Translating a partner's question into the correct query over case, attorney, and financial data.

  • Dashboard and report generation

    Building a requested report with the right filters, time window, and aggregation.

  • Team performance tracking

    Attributing case activity and outcomes to the right attorneys and time periods.

  • Case status and insight updates

    Reporting current case posture on demand without overstating certainty about outcomes.

05

AI-Ready Case Data

Self-updating case data that works alongside an existing case management system or stands alone, auto-extracting and structuring records, bills, and calls without manual data entry.

“Works with your existing CMS, or stands on its own” www.eve.legal

Mapped capabilities

4 capabilities

  • Record and bill extraction

    Pulling structured fields from medical records and billing documents, including messy scans.

  • CMS interoperability

    Reading from and writing to an existing case management system without corrupting firm-of-record data.

  • Standalone operation

    Maintaining a coherent case record when no external CMS is connected.

  • Data freshness and self-update

    Reflecting newly ingested records in downstream case data rather than serving stale state.

06

Confidentiality and Service Reliability

The commitments Eve publishes around client data and availability: SOC2, tenant isolation, privilege protection, and the SLA's 99.5% monthly uptime target with its stated maintenance and exclusion terms.

“Eve will meet or exceed 99.5% uptime availability for every calendar month of the Term.” www.eve.legal

Mapped capabilities

4 capabilities

  • Tenant isolation of client data

    Keeping one firm's case material out of any other firm's responses or reports.

  • Privilege protection

    Handling privileged material so it is not exposed in outputs that leave the firm.

  • Uptime and maintenance communication

    Behavior consistent with the published availability commitment and maintenance notice policy.

  • SLA scope accuracy

    Representing what the SLA covers, including features excluded such as custom integrations.

Coverage is mapped from Eve's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Eve test?+

The coverage map is generated from Eve's own public product surface (legal AI platform for plaintiff law firms): 6 scoring areas — 24/7 Intake and Client Signing, Pre-Litigation Drafting, and Litigation Work Product, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Eve evals scored?+

Every case generated for Eve — across 24/7 Intake and Client Signing and Pre-Litigation Drafting and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Eve library include?+

The full Eve library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Multilingual call and email handling and Claim qualification and scoring under 24/7 Intake and Client Signing); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Eve or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Eve areas and set them up in a Corsac workspace, where you can run every test case against Eve or your own agent with your own data.