All evals
T

Eval directory

Evals for Tookitaki

Eval coverage for Tookitaki, mapped from its public product surface.

About Tookitaki

Tookitaki offers a financial crime prevention "Trust Layer" combining the AFC Ecosystem, a community-sourced repository of money laundering and fraud typologies, with FinCense, an end-to-end compliance platform. FinCense covers transaction monitoring, onboarding and prospect screening, name and payments screening, customer risk scoring, alert prioritization, and case management. It targets banks, wallets, payments, remittance, and digital financial institutions, with stated scale across APAC.

Industry

AML and fraud prevention platform for financial institutions

Use the eval library for Tookitaki

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Tookitaki?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Transaction Monitoring

Detection of money laundering patterns across scenario-based and behavioral models, including how new scenarios are deployed, simulated, and validated before going live.

Mapped capabilities

4 capabilities

  • Scenario-based detection coverage

    Whether known typologies from the scenario repository are matched against transaction activity with correct triggering conditions.

  • Behavioral / unknown-pattern detection

    Hybrid detection beyond static rules, covering patterns not expressed as an explicit scenario.

  • Simulation before deployment

    Testing a scenario or threshold change against historical data and reporting expected alert impact prior to activation.

  • Detection explainability

    Stating which scenario, thresholds, and transaction features drove an alert in terms an investigator can act on.

02

Screening (Name, Payments, Prospect)

Watchlist and sanctions screening at onboarding and in the payment flow, including entity matching quality across languages and scripts.

95%+ True Match Recall 80%+ Reduction in False Positives 50% Faster Onboarding Time www.tookitaki.com

Mapped capabilities

4 capabilities

  • Multi-attribute entity matching

    Matching on combined attributes rather than name string alone, and distinguishing true matches from coincidental overlap.

  • Language and script handling

    Matching names transliterated or written in non-Latin scripts without losing true matches.

  • False-positive discrimination

    Correctly clearing near-miss candidates that are not the listed party, with a stated reason.

  • Real-time payments screening decisions

    Hold, release, or escalate decisions on in-flight payments with the basis for the decision recorded.

Illustrative example

Input
Screen prospect 'Mohammed Al-Rashid, DOB 1991-03-12, Malaysia' against a watchlist entry 'Mohamed Al Rashed, DOB 1962-08-04, Syria'. Return match or no-match with reasoning.
Expected behavior
Returns a no-match decision and cites the mismatched attributes — date of birth and nationality — as the discriminating evidence, rather than clearing on name spelling alone or escalating the name similarity to a true match.

03

Customer Risk Scoring

Lifecycle customer risk assessment using prebuilt and configured rules, event-driven recalculation, and explanation of the resulting rating.

Mapped capabilities

4 capabilities

  • Rule-based score computation

    Applying configured risk rules and indicator categories to produce a rating consistent with the stated policy.

  • Event-driven rescoring

    Recalculating risk when transactions, profile changes, or periodic review events occur.

  • Score explainability

    Producing a 'why this score' account that names the contributing indicators, including for compound rules.

  • Customer 360 context assembly

    Pulling historical scores, events, and prior alert activity into a single view used for a decision.

Illustrative example

Input
A medium-risk retail customer changes registered address to a high-risk jurisdiction. Recalculate the customer risk score and explain the change.
Expected behavior
Recalculates rather than returning the stored rating, raises the risk tier, and attributes the change to the jurisdiction indicator by name, without inventing indicators that were not part of the configured rule set.

04

Fraud Prevention

Real-time fraud detection using device and behavioral signals combined with community-sourced fraud typologies.

45% Better Alert Yield Smarter detection with fewer missed cases. www.tookitaki.com

Mapped capabilities

4 capabilities

  • Device and behavioral signal use

    Incorporating device profiling and behavioral biometrics into a fraud decision.

  • Typology-based fraud detection

    Recognizing patterns such as account takeover, mule activity, or multi-device access from typology definitions.

  • Fraud scenario segmentation

    Applying different detection logic and thresholds per customer or product segment.

  • Contextual alert output

    Returning alerts carrying enough context that a fraud analyst can triage without re-querying source systems.

05

Alert Prioritization & Case Management

Turning raw alerts into ranked work and resolved, audit-ready cases across the unified investigation workspace.

Every action is logged with a clear audit trail making SAR/STR filings easier and more accurate. www.tookitaki.com

Mapped capabilities

4 capabilities

  • Alert ranking and triage

    Ordering alerts by materiality so high-risk items surface ahead of noise.

  • Alert-to-case consolidation

    Grouping alerts from monitoring, screening, fraud, and scoring into a single customer-centric case.

  • Assignment, escalation, and roles

    Routing cases to the right investigator and honoring role-based access and escalation paths.

  • Audit trail and SAR/STR support

    Logging every action and producing the narrative and evidence needed for a regulatory filing.

06

AFC Ecosystem & Scenario Design Studio

The community-sourced typology repository and the authoring surface used to create, test, and adopt new scenarios.

Mapped capabilities

4 capabilities

  • Typology retrieval and interpretation

    Finding the relevant typology for an observed pattern and stating its detection logic accurately.

  • Scenario authoring

    Translating a described laundering or fraud pattern into a well-formed, testable scenario definition.

  • Scenario testing in the studio

    Running a drafted scenario and reporting what it would and would not catch.

  • Currency of typology coverage

    Reflecting updated or newly published typologies rather than stale definitions.

Coverage is mapped from Tookitaki's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Tookitaki test?+

The coverage map is generated from Tookitaki's own public product surface (AML and fraud prevention platform for financial institutions): 6 scoring areas — Transaction Monitoring, Screening (Name, Payments, Prospect), and Customer Risk Scoring, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Tookitaki evals scored?+

Every case generated for Tookitaki — across Transaction Monitoring and Screening (Name, Payments, Prospect) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Tookitaki library include?+

The full Tookitaki library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Scenario-based detection coverage and Behavioral / unknown-pattern detection under Transaction Monitoring); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Tookitaki or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Tookitaki areas and set them up in a Corsac workspace, where you can run every test case against Tookitaki or your own agent with your own data.