All evals
F

Eval directory

Evals for Facctum

Eval coverage for Facctum, mapped from its public product surface.

About Facctum

Facctum is a cloud-native AI compliance platform for AML, sanctions and watchlist screening aimed at banks and other financial institutions. Its modules cover watchlist management, customer screening, payment screening, transaction monitoring, alert adjudication and know your business. It markets real-time screening at enterprise data volumes with reduced false positives and faster compliance decisions.

Employees

100+

Industry

AML sanctions screening and financial crime compliance platform

Use the eval library for Facctum

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Facctum?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Watchlist Management & List Currency

Ingesting, enriching and maintaining sanctions and watchlist data so screening runs against current, validated records, and so list changes are applied and explainable.

Facctum's cloud-native platform processes millions of records per second, enabling real-time screening at enterprise scale. www.facctum.com

Mapped capabilities

4 capabilities

  • List ingestion and normalization

    Handling of incoming sanctions/watchlist records, including enrichment and validation of screening data before it reaches the matching engine.

  • Update timeliness and propagation

    Behavior when a list is updated — how quickly changes take effect and how already-screened populations are treated.

  • Impact analysis of list changes

    Surfacing which customers, payments or alerts are affected when a list entry is added, amended or removed.

  • List coverage and provenance

    Identifying which lists a decision was made against and where a matched record originated.

Illustrative example

Input
A watchlist entry that previously matched several customers is removed in the next list update. Re-screen the same customer population after the update is applied.
Expected behavior
Screening after the update produces no new matches against the removed entry, and the previously generated alerts tied to that entry are flagged as affected by the list change rather than left silently open.

02

Customer Screening & Match Precision

Screening customer records against watchlists with dynamic screening profiles, and controlling the precision/recall tradeoff the product markets as reduced false positives.

reducing false positives by up to 70% www.facctum.com

Mapped capabilities

4 capabilities

  • Name matching across variants

    Transliteration, aliases, name order, spelling variation and other name-form differences in customer records.

  • Screening profiles and thresholds

    Applying dynamic screening profiles and configured match thresholds to different customer populations.

  • False-positive reduction and AI auto-tuning

    Behavior of tuning mechanisms that suppress non-substantive matches without dropping true hits.

  • Entity disambiguation

    Distinguishing between distinct parties who share identifying attributes such as name or date of birth.

Illustrative example

Input
Screen a customer record for a name given in Latin transliteration whose sanctions-list entry stores the same individual under a different accepted transliteration and reversed name order.
Expected behavior
The record is returned as a match to the sanctioned individual, not cleared. The response identifies the matched list entry and the reason for the match, so an investigator can adjudicate it without re-deriving the name comparison.

03

Payment Screening in Real Time

Screening payments inflight against sanctions and watchlist data, where latency budgets and hold/release decisions are operationally binding.

Screen customers, payments, transactions, and alerts in real time www.facctum.com

Mapped capabilities

4 capabilities

  • Inflight screening and latency behavior

    Screening a payment message in real time and returning a decision within the payment flow.

  • Hold, release and block decisioning

    Which payments are stopped, which pass, and how the outcome is recorded.

  • Payment message field handling

    Extracting and screening the parties and fields carried in a payment instruction.

  • Throughput under enterprise volume

    Behavior when screening high record volumes rather than isolated payments.

04

Transaction Monitoring & Know Your Business

Monitoring transaction activity for financial crime risk and screening business counterparties, as distinct modules from customer and payment screening.

Superior end-to-end screening for the modern financial institution. www.facctum.com

Mapped capabilities

4 capabilities

  • Transaction risk detection

    Identifying activity that warrants review under AML monitoring rules or thresholds.

  • Monitoring rules and threshold configuration

    Applying configured AML thresholds and rules to observed transaction activity.

  • Business entity screening (KYB)

    Screening corporate counterparties and their associated parties rather than individual customers.

  • Complex entity and ownership structures

    Handling layered or cross-border business structures during business screening.

05

Alert Adjudication & Investigator Workflow

The workflow where generated alerts are triaged, investigated and dispositioned by compliance staff, including what evidence the investigator is given.

Mapped capabilities

4 capabilities

  • Alert triage and prioritization

    Ordering and presenting alerts so investigators work the highest-risk items first.

  • Match evidence and rationale

    Showing why a record matched, including the matched fields and source list entry.

  • Disposition and escalation paths

    Recording a decision on an alert and routing items that require escalation.

  • Alert volume and fatigue management

    Behavior when queues are large, including duplicate and repeat-match handling.

06

Data Quality, Assurance & Regulatory Alignment

Cross-cutting properties the platform is sold on: data quality at ingest, scale, privacy and security of sensitive customer and transaction data, and auditable alignment with AML regulations.

Mapped capabilities

4 capabilities

  • Data enrichment and validation at ingest

    Improving match accuracy by validating and enriching screening data before matching.

  • Audit trail and decision explainability

    Reconstructing a past screening decision and the data state it was made against.

  • Data privacy and security handling

    Treatment of sensitive customer identity and transaction data under global regulatory obligations.

  • Model validation and AI governance

    Evidence supporting the behavior of AI-driven matching and tuning components.

Coverage is mapped from Facctum's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Facctum test?+

The coverage map is generated from Facctum's own public product surface (AML sanctions screening and financial crime compliance platform): 6 scoring areas — Watchlist Management & List Currency, Customer Screening & Match Precision, and Payment Screening in Real Time, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Facctum evals scored?+

Every case generated for Facctum — across Watchlist Management & List Currency and Customer Screening & Match Precision and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Facctum library include?+

The full Facctum library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, List ingestion and normalization and Update timeliness and propagation under Watchlist Management & List Currency); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Facctum or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Facctum areas and set them up in a Corsac workspace, where you can run every test case against Facctum or your own agent with your own data.