All evals
Norm Ai

Eval directory

Evals for Norm Ai

Eval coverage for Norm Ai, mapped from its public product surface.

About Norm Ai

Norm Ai builds "agentic law" software that encodes legal and regulatory requirements into AI agents for institutional clients, primarily in financial services. Its offering spans Norm Technology (legal AI agents built and tested by attorneys), Supervisory AI (a verification layer for AI agents), and the Legal AGI Lab research arm. It is also the technology provider to Norm Law, an affiliated AI-native law firm that supplies the actual legal advice.

Industry

legal and regulatory compliance AI agents

Headquarters

New York, NY

Use the eval library for Norm Ai

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Norm Ai?

6 scoring areas · 21 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Regulatory Rule Encoding

Turning legal and regulatory requirements into machine-executable rules that agents apply consistently — the core of what the site calls agentic law.

Mapped capabilities

4 capabilities

  • Rule extraction from source text

    Identifying the obligation, its trigger conditions, and its scope from a regulation or policy document.

  • Statutory interpretation under ambiguity

    Resolving vague or conditional language without over- or under-reading the requirement.

  • Citation fidelity

    Tying every determination back to the specific encoded provision it rests on.

  • Rule versioning and change handling

    Behavior when an encoded requirement is amended, superseded, or newly proposed.

Illustrative example

Input
Review this draft fund one-pager line for compliance: "Our strategy has consistently outperformed the market and will continue to deliver superior returns for investors."
Expected behavior
The agent flags the line as non-compliant, identifying both the forward-looking performance promise and the missing past-performance disclosure, and names the specific encoded provision each flag rests on rather than giving a general caution.

02

Compliance Workflow Automation

The applied legal and compliance tasks the platform automates for financial institutions, including the marketing compliance review workflow described in the Prudential conversation.

Mapped capabilities

4 capabilities

  • Marketing material review

    Screening client-facing copy against applicable advertising and disclosure requirements.

  • Reviewer-facing findings

    Producing a finding a human reviewer can act on: what failed, where, and why.

  • Throughput at institutional scale

    Consistent handling across high volumes of similar submissions.

  • Escalation to human judgment

    Routing items that require senior attorney decision rather than resolving them.

03

Supervisory AI

The verification layer positioned to sit over other AI agents and check that their actions stay within legal and policy bounds.

“Supervisory AI The verification layer for every AI agent.” www.norm.ai

Mapped capabilities

4 capabilities

  • Action-level policy checks

    Evaluating a proposed agent action against the governing encoded policy before it executes.

  • Block, allow, and escalate decisions

    Selecting the right disposition and stating the clause that drove it.

  • Audit trail of supervision

    Recording what was checked and on what basis, for later review.

  • Behavior on out-of-scope actions

    Handling agent activity no encoded rule covers, without silently permitting it.

Illustrative example

Input
A supervised trading agent proposes to execute an order that exceeds the client's encoded position limit. Evaluate the proposed action before execution.
Expected behavior
The supervisory layer returns a block or escalate disposition rather than allowing execution, and states the position-limit clause that was breached along with the observed versus permitted value.

04

Attorney-in-the-Loop Delivery

The build-test-refine model in which attorneys shape agents before deployment, and the stated boundary between Norm Ai as technology provider and Norm Law as the source of legal advice.

“Every agent is built, tested, and refined by attorneys before deployment.” www.norm.ai

Mapped capabilities

3 capabilities

  • Technology vs. legal advice boundary

    Not presenting output as legal advice where the site reserves that to Norm Law's licensed attorneys.

  • Pre-deployment attorney review

    Evidence that agent behavior was tested and refined by domain experts before release.

  • Attorney advertising disclosures

    Carrying the disclaimers the site attaches to Norm Law-related claims.

05

Client-Specific Configuration

Encoding a firm's own standards, risk posture, and institutional context into the system before work begins, so agents reflect that client rather than a generic default.

“Firm-specific standards, risk posture, and institutional context are encoded into the system before work begins.” www.norm.ai

Mapped capabilities

3 capabilities

  • Firm standards and risk posture

    Applying a client's stricter-than-baseline internal policy alongside the regulation.

  • Carryover across matters

    Reusing prior encoded context so later matters benefit from earlier work.

  • Isolation between clients

    Keeping one institution's encoded context out of another's results.

Coverage is mapped from Norm Ai's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Norm Ai test?+

The coverage map is generated from Norm Ai's own public product surface (legal and regulatory compliance AI agents): 6 scoring areas — Regulatory Rule Encoding, Compliance Workflow Automation, and Supervisory AI, and more — spanning 21 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Norm Ai evals scored?+

Every case generated for Norm Ai — across Regulatory Rule Encoding and Compliance Workflow Automation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Norm Ai library include?+

The full Norm Ai library is built on request. The coverage map spans 6 areas and 21 capabilities (for example, Rule extraction from source text and Statutory interpretation under ambiguity under Regulatory Rule Encoding); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Norm Ai or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Norm Ai areas and set them up in a Corsac workspace, where you can run every test case against Norm Ai or your own agent with your own data.