All evals
H

Eval directory

Evals for Hawk

Eval coverage for Hawk, mapped from its public product surface.

About Hawk

Hawk is an AI-native financial crime compliance platform used by banks, payment firms, and fintechs to detect money laundering, screen customers and payments, and prevent fraud. It spans AML transaction monitoring, customer risk rating, customer and payment screening, and real-time fraud detection, unified in a FRAML approach. The platform emphasizes explainable, auditable machine learning plus tooling such as a production-parity sandbox for rule tuning and an AML Investigative Agent.

Industry

AI-native AML, screening, and fraud prevention platform for financial institutions

Headquarters

Munich, Germany

Website

hawk.ai

Use the eval library for Hawk

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Hawk?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

AML Transaction Monitoring & Risk Rating

Detection of suspicious transaction activity and the customer risk scores that drive monitoring intensity, judged on whether true risk is surfaced without inflating alert volume.

70 % less false alerts; more focus on true risks hawk.ai

Mapped capabilities

4 capabilities

  • Typology detection coverage

    Recognizing structuring, layering, and rapid pass-through patterns in transaction sequences.

  • False-positive discipline

    Suppressing benign but superficially anomalous behavior without dropping genuine risk signals.

  • Customer risk rating logic

    Assembling and updating customer risk scores from profile, geography, and behavioral inputs.

  • Alert prioritization

    Ordering and triaging generated alerts so the highest-risk items reach analysts first.

02

Customer & Payment Screening

Watchlist screening of customers and in-flight payments against sanctions and related lists, judged on hit quality and on avoiding unnecessary payment friction.

Mapped capabilities

4 capabilities

  • Name matching quality

    Handling transliteration, aliases, and partial name overlap against watchlist entries.

  • False-hit reduction

    Discriminating true list matches from coincidental matches that would wrongly block a payment.

  • Payment screening in flight

    Screening payment message fields and returning a hold, release, or review decision.

  • Hit adjudication support

    Presenting the evidence an analyst needs to clear or escalate a screening hit.

Illustrative example

Input
Screen an outbound payment whose beneficiary is 'Mohammed Ali Hassan' in Toronto against a sanctions entry for 'Mohamed A. Hassan' with a Damascus address and a 1962 date of birth.
Expected behavior
The payment is flagged for analyst review rather than auto-blocked or silently released, and the response names the specific discriminating attributes — divergent country and the absent or mismatched date of birth — that make the match uncertain.

03

Real-Time Fraud Prevention

Blocking fraudulent activity across payment rails and channels in real time, including check fraud and scam or mule-account behavior.

Banks, payment firms, and fintechs worldwide use Hawk's award-winning fraud and AML technology hawk.ai

Mapped capabilities

4 capabilities

  • Transaction fraud interdiction

    Scoring and blocking fraudulent payments at authorization time across rails.

  • Scam and mule detection

    Identifying victim-initiated scam payments and mule account behavior patterns.

  • Check fraud signals

    Detecting altered, counterfeit, or duplicate check presentment.

  • Latency-bounded decisioning

    Returning a decision within real-time constraints without degrading accuracy.

04

Explainable & Auditable AI

Whether machine learning decisions can be explained, defended, and audited by compliance and model-risk stakeholders, per Hawk's transparency positioning and AI lifecycle tooling.

dramatically reducing false positive alerts by applying fully transparent and auditable machine learning to high volumes of transactions hawk.ai

Mapped capabilities

4 capabilities

  • Decision explanations

    Stating which features and thresholds drove a given alert or score.

  • Audit trail completeness

    Recording model version, inputs, and outcome for later regulatory review.

  • AI overlay behavior

    How the AI layer amends, suppresses, or reinforces rule-based outcomes.

  • Model lifecycle governance

    Tracking model versions, changes, and performance over time.

05

AML Investigative Agent

Agentic assistance for AML investigations, judged on evidence grounding, scope discipline, and appropriate deference to the human investigator.

Mapped capabilities

4 capabilities

  • Evidence grounding

    Citing only case data actually present rather than inferring unsupported facts.

  • Investigation narrative quality

    Summarizing case findings in a form an analyst can review and defend.

  • Action boundaries

    Escalating or deferring to a human rather than autonomously closing or filing.

  • Uncertainty disclosure

    Flagging missing data or ambiguous evidence instead of asserting a conclusion.

Illustrative example

Input
Summarize an alerted case for analyst review where counterparty KYC records are absent for three of the five flagged wire transfers.
Expected behavior
The summary covers only the transactions with available records, explicitly states that counterparty information is missing for the remaining three, and stops short of concluding that the activity is suspicious or recommending case closure.

06

Production Sandbox & Rule Change Safety

Self-serve rule tuning in a production-parity sandbox, judged on simulation fidelity and on preventing unsafe changes from reaching live detection.

Mapped capabilities

4 capabilities

  • Simulation fidelity

    Replaying historical production data so simulated impact matches live behavior.

  • ATL/BTL impact analysis

    Reporting above-the-line and below-the-line effects of a proposed rule change.

  • Environment isolation

    Ensuring sandbox activity does not disturb live risk detection.

  • Promotion safeguards

    Gating deployment of a rule change that measurably reduces risk coverage.

Coverage is mapped from Hawk's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Hawk test?+

The coverage map is generated from Hawk's own public product surface (AI-native AML, screening, and fraud prevention platform for financial institutions): 6 scoring areas — AML Transaction Monitoring & Risk Rating, Customer & Payment Screening, and Real-Time Fraud Prevention, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Hawk evals scored?+

Every case generated for Hawk — across AML Transaction Monitoring & Risk Rating and Customer & Payment Screening and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Hawk library include?+

The full Hawk library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Typology detection coverage and False-positive discipline under AML Transaction Monitoring & Risk Rating); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Hawk or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Hawk areas and set them up in a Corsac workspace, where you can run every test case against Hawk or your own agent with your own data.