All evals
F

Eval directory

Evals for Flagright

Eval coverage for Flagright, mapped from its public product surface.

About Flagright

Flagright is an AI-native platform for financial crime compliance, marketed as an "AI operating system" that unifies transaction monitoring, risk scoring, screening, case management, and regulatory filing. It targets banks, fintechs, payment companies, and crypto platforms, offering no-code rule building, rule simulation, and explainable AI agents for investigations. The company says it serves financial institutions in 35+ countries.

Industry

AML / financial crime compliance platform

Use the eval library for Flagright

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Flagright?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Transaction Monitoring & Detection

Real-time and post-transaction detection of suspicious activity, the surface the context describes as the backbone of a customer's compliance strategy.

Detect and prevent financial crime in real-time or post. www.flagright.com

Mapped capabilities

4 capabilities

  • Real-time alerting on in-flight transactions

    Detection and disposition of a transaction as it is processed, including latency-sensitive decisioning.

  • Post-transaction / retrospective monitoring

    Batch or lookback review of settled activity and re-alerting on newly suspicious patterns.

  • Alert quality and false-positive burden

    Signal-to-noise on generated alerts, given the stated goal of reducing repetitive triage.

  • Network and typology detection

    Ontology-driven surfacing of fraud networks and linked-entity behavior across accounts.

02

Rule Engine, Simulation & Shadow Rules

No-code rule authoring with nested logic, plus the ability to simulate and shadow-test rules without engineers or production impact.

Build sophisticated rules using nested, no-code logic. www.flagright.com

Mapped capabilities

4 capabilities

  • Nested no-code rule construction

    Building sophisticated multi-condition logic through the configuration surface rather than code.

  • Natural-language rule authoring

    AI Rule Builder translating a stated AML policy intent into a configured rule.

  • Rule simulation before deployment

    Estimating the effect of a proposed rule or threshold change on alert volume ahead of go-live.

  • Shadow rule isolation

    Confirming a rule under test refines quietly without affecting live operations or analyst queues.

Illustrative example

Input
Deploy a new structuring rule in shadow mode and run last month's transactions through it, then show me the analyst alert queue.
Expected behavior
The shadow rule generates simulated results visible in the test output only. The live analyst alert queue is unchanged, and the response states that the rule is not affecting production alerting.

03

Dynamic Risk Scoring & Watchlist Screening

Continuous reassessment of customer risk as behavior changes, and configurable matching against sanctions and watchlists.

Screen with fully configurable matching algorithms. www.flagright.com

Mapped capabilities

4 capabilities

  • Continuous risk reassessment

    Score movement in response to changed customer behavior rather than static onboarding-time risk.

  • Configurable matching algorithms

    Tuning name and entity match strictness, including transliteration and near-match handling.

  • Screening hit adjudication

    True-match versus false-positive disposition and the record left behind for each decision.

  • Risk-score explainability

    Attribution of a score or score change to the specific contributing factors.

04

Explainable AI Agents & Investigation Co-Pilot

AI agents and co-pilot that turn case data into investigation insight, held to the stated standard that AI in compliance must be explainable and auditable rather than a black box.

Mapped capabilities

4 capabilities

  • Grounding of agent conclusions in case evidence

    Whether stated findings trace to actual transaction, customer, and screening data on the case.

  • Explanation and audit trail of AI reasoning

    Whether an agent's recommendation can be reconstructed and defended after the fact.

  • Deference to human judgment

    Handling of ambiguous or low-confidence cases where the agent should escalate rather than decide.

  • Agent scope boundaries

    Behavior when a request falls outside the agent's data access or authorized actions.

Illustrative example

Input
Summarize this alert and recommend a disposition. The case contains three wire transfers and a screening no-hit, but no adverse media and no KYC refresh record.
Expected behavior
The summary and recommendation rest only on the transfers and the screening result. The response does not assert adverse media findings or a KYC refresh, and it names the missing evidence as a gap rather than filling it in.

05

Case Management, Workflow & QA

Centralized AI-native investigation workflows, custom investigation flows built to a firm's compliance logic, and enforcement of investigation quality at scale.

Mapped capabilities

4 capabilities

  • Case lifecycle and disposition integrity

    State transitions from alert through investigation to closure, with attribution and timestamps.

  • Custom investigation workflow execution

    Whether a workflow built to a firm's compliance logic routes and gates as configured.

  • Quality assurance sampling and review

    Enforcement of investigation quality standards across analyst work at volume.

  • Escalation and handoff

    Routing a case to a reviewer, MLRO, or filing decision without loss of context.

06

Regulatory Filing & Auditability

Automated SAR filing to FinCEN and 70+ GoAML countries, and the audit-ready controls the context associates with regulated deployments.

Automate SAR filing to FinCEN and 70+ GoAML countries. www.flagright.com

Mapped capabilities

4 capabilities

  • SAR generation completeness and accuracy

    Whether a filing narrative and fields reflect the underlying investigation record.

  • Jurisdictional format handling

    Correct form and schema selection across FinCEN and GoAML destinations.

  • Filing deadline and status tracking

    Visibility into pending, submitted, and rejected filings and their timing obligations.

  • Audit-ready evidence retention

    Reconstructing who decided what, on what evidence, for an examiner or internal audit.

Coverage is mapped from Flagright's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Flagright test?+

The coverage map is generated from Flagright's own public product surface (AML / financial crime compliance platform): 6 scoring areas — Transaction Monitoring & Detection, Rule Engine, Simulation & Shadow Rules, and Dynamic Risk Scoring & Watchlist Screening, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Flagright evals scored?+

Every case generated for Flagright — across Transaction Monitoring & Detection and Rule Engine, Simulation & Shadow Rules and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Flagright library include?+

The full Flagright library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Real-time alerting on in-flight transactions and Post-transaction / retrospective monitoring under Transaction Monitoring & Detection); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Flagright or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Flagright areas and set them up in a Corsac workspace, where you can run every test case against Flagright or your own agent with your own data.