All evals
Manifest OS

Eval directory · Legal AI

Evals for Manifest OS

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Manifest OS AI products.

About Manifest OS

Manifest OS provides an AI-native operating model for its law-firm partners, combining a unified brand, centralized back-office services, and software for client communication, legal research, drafting, billing, and reporting.

Industry

Legal Technology / Law-Firm Operating Platform

Use the eval library for Manifest OS

All 132 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Manifest OS?

5 areas · 132 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Ai Case Evaluation

Mapped capabilities

15 scenarios

  • Case intake trigger — evaluation initiated from matter creation

Public sample case

Input
A matter record for an H-1B sponsorship case (matter_id: m-7741, petitioner: Arjun Sharma) was persisted to the database after a successful intake form submission at 11:04 UTC. A transient queue-broker error silently dropped the …
Expected behavior
Agent queries both the matter database record and the evaluation job queue for m-7741. Confirms the matter record exists. Confirms no corresponding evaluation job exists in the queue and no dead-letter entry is present. Raises an explicit, actionable alert naming the matter ID and the gap. Recommen…
Check
Pass / fail check

02

Ai Drafter Document Generation

Mapped capabilities

13 scenarios

  • Document type selection

Public sample case

Input
The internal template library contains an I-129 form whose edition-date metadata field reads '01/17/17'. USCIS has released at least one newer edition and has explicitly withdrawn acceptance of prior editions within its published…
Expected behavior
Agent retrieves the I-129 template and reads its edition-date metadata field before committing to selection. Agent compares the stored edition date against a platform-maintained authoritative source (USCIS form-acceptance registry, compliance metadata feed, or equivalent). If the edition does not m…
Check
Pass / fail check

03

Evidence Document Collection

Mapped capabilities

32 scenarios

  • Evidence-playbook selection and binding to matter

Public sample case

Input
A client intake form was submitted with the visa category field reading 'work visa for performing artist.' The agent's classification module assigns H-1B as the target visa type with 0.72 confidence. An O-1B classification would …
Expected behavior
The agent halts before binding. It flags the sub-threshold confidence score, identifies the textual signal 'performing artist' as a potential O-1B indicator, creates an attorney-review task explicitly describing the classification ambiguity (H-1B vs O-1B), and records the ambiguity in the matter au…
Check
Pass / fail check

04

Matter Intake Client Onboarding

Mapped capabilities

56 scenarios

  • Individual portal — new matter session initiation

05

Workflow Orchestration Agent Loop Control

Mapped capabilities

16 scenarios

  • Execution phase task dispatch

Frequently asked questions

What do the Corsac evals for Manifest OS test?+

Each eval pack tests Manifest OS's public product surface — including Ai Case Evaluation, Ai Drafter Document Generation, and Evidence Document Collection — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Manifest OS evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 132 Manifest OS cases — from Matter Intake Client Onboarding (56 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Manifest OS library.

How many test cases does the Manifest OS library include?+

The Manifest OS eval library includes 132 graded test cases across 5 eval packs, the largest being Matter Intake Client Onboarding with 56 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Manifest OS or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 5 Manifest OS packs — Ai Case Evaluation and Ai Drafter Document Generation and the rest — against Manifest OS or your own agent with your own data.