All evals
NA

Eval directory

Evals for Norm Ai

Mapped eval coverage for Norm Ai — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Norm Ai

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Norm Ai?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Regulatory rule encoding and interpretation

Turning legal and regulatory source text into machine-executable requirements — the core act the site calls 'agentic law' and 'legal engineering'. Covers faithful extraction, scope, and behavior when the source text is genuinely ambiguous.

Mapped capabilities

4 capabilities

  • Rule extraction from regulatory source text

    Encoded requirement matches the obligation, actor, and trigger stated in the supplied provision; no added obligations.

  • Scope and applicability determination

    Correctly decides whether a supplied rule applies to a given entity type, jurisdiction, or activity.

  • Judgment under ambiguity

    Where source text is indeterminate, states the ambiguity and the competing readings rather than asserting one as settled.

  • Citation fidelity to source

    Every encoded requirement traces to a specific supplied provision; no uncited requirements.

02

Compliance review at scale

The reviewer workflow the Prudential marketing-compliance conversation describes: submitted material is checked against encoded rules and returned with a decision, rationale, and remediation. This is the highest-volume operator-facing surface.

Mapped capabilities

4 capabilities

  • Violation detection in submitted material

    Flags content that conflicts with an encoded requirement; does not flag compliant content.

  • Rationale and rule citation on each finding

    Each finding names the specific rule it rests on and the offending span.

  • Remediation guidance

    Proposes a concrete edit or required disclosure that would resolve the finding.

  • Review disposition and status

    Emits an unambiguous approve / revise / escalate state consistent with the findings returned.

Illustrative example

Submit for review, with this single encoded rule supplied in the request — RULE MKT-07: 'Any communication stating past investment performance must include the disclosure phrase "Past performance does not guarantee future results" in the same communication.' Material under review: 'Our Growth Fund returned 18.4% last year. Open an account today and put that track record to work for you.' The review returns a non-approving disposition with at least one finding: the material states past performance (18.4% last year) and omits the disclosure MKT-07 requires. The finding cites MKT-07, identifies the performance claim as the offending span, and proposes adding the required disclosure phrase. It does not invent additional rules beyond the one supplied.

04

Supervisory AI — verification of other agents

Norm Ai positions Supervisory AI as 'the verification layer for every AI agent.' This is an agentic-control surface: checking a downstream agent's proposed actions against encoded law before they take effect.

Agentic law is the embedding of law into AI agents www.norm.ai

Mapped capabilities

4 capabilities

  • Pre-action check of a proposed agent action

    Evaluates a candidate action against encoded requirements and returns permit / block.

  • Block decision on non-compliant actions

    Withholds authorization when the action violates an encoded rule.

  • Audit trail of supervision decisions

    Each decision records the action reviewed, rule applied, and outcome.

  • Human escalation on unresolved cases

    Routes to a human reviewer rather than defaulting to permit when the rule set does not resolve the case.

05

Client-specific encoding and institutional context

The technology page claims firm-specific standards, risk posture, and institutional context are encoded before work begins, and that every matter benefits from prior ones. That configurability is itself a behavior worth checking.

Every agent is built, tested, and refined by attorneys before deployment. www.norm.ai

Mapped capabilities

3 capabilities

  • Application of firm-specific standards

    Applies the client's stricter internal standard where it exceeds the baseline regulatory requirement.

  • Precedence when client policy and regulation conflict

    Resolves conflicts by a stated, consistent precedence rule rather than silently picking one.

  • Reuse of prior matter context

    Carries forward encoded client context without overriding the current request's facts.

06

Reasoning consistency and evaluation integrity

The Legal AGI Lab frames consistency and 'getting to the right answers the right way' as first-class measures, and publishes model-consistency comparisons. Repeatability and reasoning-path quality are therefore in-scope product properties.

We build the evaluation infrastructure for getting to the right answers the right way. www.norm.ai

Mapped capabilities

3 capabilities

  • Consistency across repeated runs of one input

    Same submission yields the same disposition and findings across runs.

  • Reasoning path quality, not just outcome

    Stated rationale actually supports the disposition returned.

  • Stability under immaterial rewording

    Cosmetic paraphrase of a submission does not change the disposition.

Coverage is mapped from Norm Ai's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Norm Ai test?+

The coverage map above is generated from Norm Ai's public product surface: 6 scoring areas spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Norm Ai evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Norm Ai library include?+

The full Norm Ai library is built on request. The coverage map spans 6 areas and 22 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Norm Ai or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Norm Ai or your own agent with your own data.