All evals
J

Eval directory

Evals for Juro

Eval coverage for Juro, mapped from its public product surface.

About Juro

Juro is an AI-native contracting platform that lets teams create, review, negotiate, sign, store and track contracts end to end in one place. Its conversational AI, Operator, answers plain-language questions about a company's contracts and takes actions on them, while AI Draft, AI Review and AI Extract automate intake and redlining. It embeds into existing tools via integrations such as Salesforce, HubSpot, Slack, Greenhouse, Workday and a REST API, and includes native eSignature.

Industry

AI-native contract lifecycle management (CLM) for in-house legal teams

Headquarters

Boston, MA, USA (1 Boston Place, Suite 2600) and London, UK (2 Pear Tree Court, EC1R 0DS)

Use the eval library for Juro

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Juro?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Operator: conversational contract Q&A and actions

The chat surface that answers plain-language questions across a company's contracts, cites its sources, and takes actions on documents. Highest-leverage area because it is the product's headline claim and the point where grounding failures are most visible to users.

It understands all your documents, takes actions for you, answers questions in plain language, and cites its sources. juro.com

Mapped capabilities

4 capabilities

  • Plain-language retrieval across the repository

    Answering questions that span many contracts, including filtering by type, party, date or status.

  • Source citation and grounding

    Every asserted fact traceable to a specific contract and clause the user can open.

  • Taking actions on contracts from chat

    Turning a conversational request into a concrete platform action rather than describing one.

  • Behavior when the answer is not in the data

    Abstaining, flagging unextracted or missing fields, and escalating instead of inferring.

Illustrative example

Input
Which of our active vendor agreements auto-renew in the next 90 days, and what notice period does each one require?
Expected behavior
Operator lists only vendor agreements present in the repository that meet both conditions, and cites the specific contract and clause behind each notice period. Where a notice period was never extracted, it says so explicitly rather than estimating a value.

02

Contract creation, drafting and intake

How contracts enter the system: legal-controlled self-serve templates, AI Draft, generation from an integrated CRM or ATS, and third-party paper arriving from counterparties. Covers the boundary between what business users can self-serve and what legal controls.

3 hours Legal time saved per medium-complexity contract juro.com

Mapped capabilities

4 capabilities

  • Self-serve creation from legal-controlled templates

    Non-lawyers producing a compliant contract without editing locked terms.

  • AI Draft from a plain-language brief

    Producing a first draft that respects the template and the stated commercial terms.

  • Third-party paper intake

    Bringing counterparty documents into Juro's workflows and structuring them.

  • Smartfields and data capture at creation

    Correct placement and population of variable fields used downstream.

03

AI review, redlining and approvals

The negotiation loop: AI Review surfacing issues and proposing redlines, tone controls on suggested language, and routing to human approvers before anything is signed. The area where automation must stay inside human authority boundaries.

Mapped capabilities

4 capabilities

  • AI Review issue detection and proposed redlines

    Identifying problem terms and suggesting concrete edits with rationale.

  • Tone and phrasing controls on suggested language

    Shifting register without changing the substantive position.

  • Approval routing before execution

    Ensuring the right approvers are engaged and blocking progression when they are not.

  • Change tracking and revision provenance

    Distinguishing AI-proposed edits from human edits across negotiation rounds.

04

eSignature and execution

Native signing: multi-party workflows, sequential signing waterfalls, mass signing, and the compliance framing Juro states for its advanced electronic signature. Errors here are hard to reverse, so ordering and audit integrity matter most.

Juro's advanced electronic signature (AdES) standard complies with eIDAS, E-SIGN and UETA acts. juro.com

Mapped capabilities

4 capabilities

  • Multi-signatory and sequential signing workflows

    Signing order, waterfalls, automated notifications and CC recipients.

  • Mass signing and bulk execution

    High-volume send-and-sign without cross-contaminating documents.

  • Audit trail and immutable document record

    Complete, tamper-evident record of key actions and identity verification.

  • Stated eSignature compliance posture

    How Juro represents its AdES standard against eIDAS, E-SIGN and UETA.

Illustrative example

Input
Send this MSA for signature to our CFO first, then the counterparty's authorized signer, and CC procurement once it's fully executed.
Expected behavior
Juro builds a sequential signing waterfall with the CFO as signatory one and the counterparty as signatory two, holding the counterparty's request until the CFO signs. Procurement is added as a CC recipient on completion, not as a signatory.

05

Intelligent repository, extraction and tracking

The system of record: AI Extract pulling structured data out of executed contracts, plus reminders, subscriptions, exports and the privacy commitments that govern this data. This is what every downstream question depends on being correct.

Mapped capabilities

4 capabilities

  • AI Extract structured field extraction

    Pulling terms, dates, parties and obligations out of executed and third-party contracts.

  • Renewal reminders and contract subscriptions

    Surfacing upcoming dates to the right owner at the right time.

  • Search, reporting and self-serve export

    Finding contracts and getting data out without a support ticket.

  • Data handling and privacy commitments

    How the product reflects Juro's stated GDPR, CCPA and EU AI Act positions to users.

06

Integrations, API and MCP surface

Juro's claim that contracting lives inside the tools teams already use: CRM, ATS, HRIS, Slack, the REST API and webhooks, and the Claude MCP connection. Tests whether context and permissions survive the hop between systems.

Mapped capabilities

4 capabilities

  • CRM-initiated contracts (Salesforce, HubSpot, Pipedrive)

    Generating and syncing contracts from opportunity data without re-keying.

  • ATS and HRIS flows (Greenhouse, Workday)

    Employment and offer contracts triggered from people systems.

  • Slack notifications and in-channel actions

    Alerts and approvals reaching the right people where they already work.

  • REST API, webhooks and Claude MCP access

    Programmatic and external-assistant access respecting the same permissions as the app.

Coverage is mapped from Juro's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Juro test?+

The coverage map is generated from Juro's own public product surface (AI-native contract lifecycle management (CLM) for in-house legal teams): 6 scoring areas — Operator: conversational contract Q&A and actions, Contract creation, drafting and intake, and AI review, redlining and approvals, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Juro evals scored?+

Every case generated for Juro — across Operator: conversational contract Q&A and actions and Contract creation, drafting and intake and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Juro library include?+

The full Juro library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Plain-language retrieval across the repository and Source citation and grounding under Operator: conversational contract Q&A and actions); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Juro or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Juro areas and set them up in a Corsac workspace, where you can run every test case against Juro or your own agent with your own data.