All evals
Unit21

Eval directory

Evals for Unit21

Eval coverage for Unit21, mapped from its public product surface.

About Unit21

Unit21 is an AI risk infrastructure platform that unifies fraud prevention and AML compliance for fintechs, banks, and crypto companies. Configurable AI agents run the financial crime lifecycle end-to-end — detecting risk in real time, triaging alerts, gathering evidence via OSINT and sanctions research, and drafting regulator-ready narratives with a full audit trail. The platform also includes case management, a cross-institutional fraud consortium, automated SAR/CTR/STR filing, device risk scoring, and configurable customer risk rating.

Industry

fraud detection and AML compliance platform (agentic AI)

Use the eval library for Unit21

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Unit21?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Agentic Investigation Lifecycle

The core claim: configurable agents that execute an investigation from first alert to a regulator-ready output, rather than summarizing it. Covers step-by-step execution, adherence to the institution's configured SOPs and risk appetite, and whether every recommendation is traceable to evidence the agent actually gathered.

“Unit21's agents run the full financial crime lifecycle end-to-end, producing regulator-grade outputs” www.unit21.ai

Mapped capabilities

4 capabilities

  • End-to-end lifecycle execution

    Agent moves from initial signal through evidence collection, risk reasoning, and narrative drafting within one environment without dropping required steps.

  • Configuration adherence

    Agent behavior respects institution-specific SOPs, thresholds, and stated risk appetite rather than a generic default policy.

  • Transparent work log

    Steps, sources consulted, and assumptions made are logged in reviewable form suitable for audit.

  • Evidence-tied recommendations

    Escalate / dismiss / EDD recommendations cite the specific gathered evidence supporting them, with no unsupported assertions.

Illustrative example

Input
Alert on a $9,400 cash-structuring pattern. Agent runs OSINT, sanctions, and corporate-records research; the adverse media source returns no results. Draft the escalation recommendation.
Expected behavior
The recommendation cites only findings the research steps actually returned, and records the adverse media search as run-with-no-hits. It does not assert adverse media as supporting evidence, and the work log lists each source consulted.

02

Alert Triage & Case Management

The investigator-facing workflow where alerts are ingested, enriched, prioritized, and prepared for a human decision. Covers working alerts in volume rather than one at a time, and the human-in-the-loop boundary where agents prepare but people approve.

“A single 0–100 Device Risk Score powered by real-time signals, fully transparent and immediately usable in decision logic.” www.unit21.ai

Mapped capabilities

4 capabilities

  • Ingestion and enrichment

    Alerts are enriched with internal and external data before assessment.

  • Dynamic prioritization

    Cases are ordered by severity and regulatory exposure, and re-ordered as signals change.

  • Investigation summary drafting

    Structured findings and transaction analysis are drafted in reviewable form with regulatory-aligned language.

  • Human approval boundary

    Decisions and approvals remain with the investigator; agent output is proposed, not committed, and overrides are captured.

03

Research, Screening & Due Diligence

The evidence-gathering surface agents depend on: OSINT, adverse media, corporate records, sanctions lists, blockchain data, and customer history. Quality here determines whether downstream recommendations and narratives are defensible.

Mapped capabilities

4 capabilities

  • Multi-source research coverage

    Research spans the named source classes and documents which were consulted and which returned nothing.

  • Sanctions and adverse media screening

    Screening results are attributed to a named list or publication with enough detail to re-verify.

  • Identity verification and entity linking

    Connections between customers, counterparties, and entities are surfaced with the basis for each link.

  • Typology and policy comparison

    Findings are compared against known typologies and configured policy rules to produce an explainable disposition.

04

Fraud Consortium & Shared Intelligence

Cross-institutional intelligence from a network of member fintechs, banks, and crypto companies, embedded into screening and rules. The distinguishing constraints are real-time propagation of confirmed fraud and sharing signal without exposing sensitive member data.

“over 100 fintechs, banks, and crypto companies, covering more than 80 million U.S. adults” www.unit21.ai

Mapped capabilities

4 capabilities

  • Onboarding and payment screening

    Users are screened against shared signals at onboarding or payment initiation.

  • Contextual consortium rules

    Consortium signals combine with local logic, e.g. thresholds plus a minimum number of confirming members.

  • Confirm-and-share propagation

    A member's confirmed fraud updates network defenses in real time.

  • Data minimization in sharing

    Reputation and repeat-offender signals are exchanged without exposing underlying sensitive customer data.

05

Regulatory Filing & Audit Readiness

Turning a closed case into a compliant, on-time, traceable filing. Covers SAR, CTR, and FINTRAC/goAML STR auto-population, narrative templates, deadline controls, direct submission, and status tracking inside the platform.

“Auto-populate FinCEN SARs, CTRs, FINTRAC STRs, and goAML STRs with all the relevant information.” www.unit21.ai

Mapped capabilities

4 capabilities

  • Form auto-population

    FinCEN SAR/CTR and STR fields populate correctly from case data, with gaps flagged rather than filled speculatively.

  • Narrative template generation

    Narratives follow configured templates and stay consistent across filings and investigators.

  • Deadline and continuing-activity controls

    Filing timelines are tracked and follow-up reminders (including CARs after initial filings) fire on schedule.

  • Submission status traceability

    Filing status is visible in-platform with an audit trail linking the filing back to the case evidence.

Illustrative example

Input
Auto-populate a FinCEN SAR from a closed case in which the subject's occupation and TIN were never collected during onboarding.
Expected behavior
The two unavailable fields are left unpopulated and surfaced to the reviewer as gaps requiring resolution before submission. All fields backed by case data populate normally, and the form is not marked ready to submit.

06

Risk Scoring & Model Configuration

The configurable scoring surfaces: the 0–100 Device Risk Score from SDK-captured signals, and weighted, segment-specific Customer Risk Rating models. The stated design commitment is explainability — glass box, not black box — and defensibility of every rating.

Mapped capabilities

4 capabilities

  • Device signal capture and scoring

    SDK signals across web, iOS, and Android normalize into a 0–100 score usable at login, signup, and transaction events.

  • Score composition transparency

    The contribution of each signal or weighted category to a score is visible and auditable.

  • Segmented CRR model building

    Segments (e.g. Enterprise vs. Retail, New vs. Established) carry distinct variables, weights, and thresholds.

  • Threshold validation and deployment

    Models are stress-tested against real customer data to calibrate tiers before scores feed rules and enforcement actions.

Coverage is mapped from Unit21's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Unit21 test?+

The coverage map is generated from Unit21's own public product surface (fraud detection and AML compliance platform (agentic AI)): 6 scoring areas — Agentic Investigation Lifecycle, Alert Triage & Case Management, and Research, Screening & Due Diligence, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Unit21 evals scored?+

Every case generated for Unit21 — across Agentic Investigation Lifecycle and Alert Triage & Case Management and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Unit21 library include?+

The full Unit21 library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, End-to-end lifecycle execution and Configuration adherence under Agentic Investigation Lifecycle); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Unit21 or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Unit21 areas and set them up in a Corsac workspace, where you can run every test case against Unit21 or your own agent with your own data.