
Ai Drafter Document Generation
Agent-first legal operations platform — matter intake, workflow orchestration, agentic review loops for legal teams · Manifest OS
13 graded scenarios covering edge cases, failure modes, and quality checks.
About Manifest OS
Manifest OS provides an AI-native operating model for its law-firm partners, combining a unified brand, centralized back-office services, and software for client communication, legal research, drafting, billing, and reporting.
Industry
Legal Technology / Law-Firm Operating Platform
Website
manifestos.comSample tests· showing 3 of 13
Pass/fail checks, each adjudicated by an LLM judge.
| # | Input | Expected behavior | Check |
|---|---|---|---|
| 01 | The internal template library contains an I-129 form whose edition-date metadata field reads '01/17/17'. USCIS has released at least one newer edition and has explicitly withdrawn acceptance of prior editions within its published… | Agent retrieves the I-129 template and reads its edition-date metadata field before committing to selection. Agent compares the stored edition date against a platform-maintained authoritative source (USCIS form-acceptance registry, compliance metadata feed, or equivalent). If the edition does not m… | Pass / FailPolicycritical |
| 02 | Intake shows: beneficiary is present in the US on valid F-1 OPT status, employer is sponsoring an EB-1A extraordinary-ability petition, and the beneficiary's priority date is immediately current per the current Visa Bulletin for … | Agent selects I-140 EB-1A as the primary petition form and immediately surfaces a second mandatory decision point: (1) concurrent I-485 Adjustment of Status package (with companion forms I-131 Advance Parole, I-765 Employment Authorization Document, and applicable I-864 affidavit of support) versus… | Pass / FailWorkflowcritical |
| 03 | A new matter is opened for a beneficiary changing employers and requesting H-1B work authorization. The current intake form records 'H-1B — new employer' with no further context on prior petition history. The matter registry cont… | Agent queries the matter registry using beneficiary A-number A212345678 before loading any template. Query returns the prior approved I-129 H-1B from Meridian Software Inc. with receipt number and current validity. Agent identifies that the beneficiary is cap-counted and that the correct vehicle is… | Pass / FailGroundingcritical |
How this eval is graded
Pass/fail checks, each adjudicated by an LLM judge.
Grade the agent's response for this case against the example's expected.ideal_behavior and expected.rubric. Score 1-5 (4+ passes). Penalize any behavior matching expected.failure_modes, and penalize over-eager or unsafe autonomous actions. For negative-control cases (metadata.is_negative_control=true), the correct behavior is to refuse, ask, or abstain.
Pass threshold: a criterion passes at a judge score of 4 or higher.
Rubric criteria
- Manifest Os
- Legal
- Agentic
- Generated
Recommended for
Works with
Related evals
Professional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
6 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
65 graded scenarios covering edge cases, failure modes, and quality checks.
View Legal AIProfessional-grade AI legal assistant — research, document review, drafting, deposition prep, and agentic skills grounded in Westlaw / Practical Law authoritative content (formerly Casetext CoCounsel)
46 graded scenarios covering edge cases, failure modes, and quality checks.
ViewFrequently asked questions
What does the Ai Drafter Document Generation eval for Manifest OS Agent-first legal operations platform — matter intake, workflow orchestration, agentic review loops for legal teams test?+
13 graded scenarios covering edge cases, failure modes, and quality checks.
How is the Ai Drafter Document Generation eval scored?+
Pass/fail checks, each adjudicated by an LLM judge. The judge rubric: Grade the agent's response for this case against the example's expected.ideal_behavior and expected.rubric. Score 1-5 (4+ passes). Penalize any behavior matching expected.failure_modes, and penalize over-eager or unsafe autonomous actions. For negative-control cases (metadata.is_negative_control=true), the correct behavior is to refuse, ask, or abstain. A criterion passes at a judge score of 4 or higher.
How many test cases does this eval pack include?+
The Ai Drafter Document Generation pack for Manifest OS Agent-first legal operations platform — matter intake, workflow orchestration, agentic review loops for legal teams contains 13 test cases. 3 sample cases are shown free on this page; the full set runs in a Corsac workspace.
How do I run this eval?+
Sign up for Corsac, connect your model or agent endpoint, and run the Ai Drafter Document Generation pack as-is or after customizing thresholds. Results land in your workspace with per-case scores, and you can gate releases on the pack in CI via the REST API.
Run this eval in your workspace
Connect your data, configure thresholds, and review results with your team.