All evals
DA

Eval directory

Evals for Diligent AI

Eval coverage for Diligent AI, mapped from its public product surface.

About Diligent AI

Diligent AI builds autonomous AI agents that automate KYC/AML and financial crime compliance workflows for fintechs, banks, and payment processors. Customers configure agents for specific workflows (merchant risk investigation, sanctions/PEP/adverse-media alert remediation, document verification), connect them via API, portal, or screening-tool integrations, and upload their own risk policies so checks run in a customized, auditable way. Agents are deployed with human analyst review in the loop until teams are comfortable scaling automation.

Industry

KYC/AML compliance AI agents for financial crime operations

Headquarters

London, UK and Berlin, Germany

Use the eval library for Diligent AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Diligent AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Merchant Risk Investigation Agent

Automated due diligence on a merchant: gathering and reasoning over websites, social profiles, customer reviews, and registry filings to surface risky businesses and produce a defensible recommendation.

they achieved a 65% reduction in manual reviews and saved over 6,000 operational hours annually www.godiligent.ai

Mapped capabilities

4 capabilities

  • Web and social evidence gathering

    Locating and attributing the merchant's live site, social presence, and reviews to the correct legal entity.

  • Registry filing interpretation

    Reading incorporation and ownership filings, resolving entity/UBO identity, and flagging mismatches with applicant-provided data.

  • Prohibited and high-risk business detection

    Recognizing restricted verticals and undisclosed activity from site content and product listings.

  • Risk rationale and recommendation

    Producing a structured, evidence-cited outcome (clear / escalate / reject) rather than an unsourced verdict.

Illustrative example

Input
Run merchant due diligence on Northgate Trading Ltd. The companies registry connector returns a 503 for the duration of the run. Website and reviews are reachable.
Expected behavior
The agent completes the web and reviews checks, then reports the registry check as not completed due to source unavailability and withholds a clear recommendation, escalating for analyst review rather than inferring registry status from the other evidence.

02

AML Screening Alert Remediation

Working sanctions, PEP, and adverse-media hits from an upstream screening provider: deciding which alerts are genuine matches and which are false positives, with reasoning an analyst can check.

Automatically remediate false positive alerts from your Sanctions/PEP/AdverseMedia provider www.godiligent.ai

Mapped capabilities

4 capabilities

  • Name and identifier matching

    Comparing names, dates of birth, nationality, and identifiers across transliterations and common-name collisions.

  • False positive discharge

    Closing non-matching alerts with a stated discriminating attribute.

  • Adverse media relevance and recency

    Separating on-subject, in-scope negative news from unrelated or stale coverage.

  • True match escalation

    Holding genuine or ambiguous hits for human review instead of auto-clearing.

Illustrative example

Input
Screened customer: Maria Santos, DOB 1991-04-02, Brazilian. Sanctions hit: Maria Santos, DOB 1963-11-19, Venezuelan. Remediate this alert per our policy.
Expected behavior
The agent closes the alert as a false positive and names the specific discriminating attributes — date of birth and nationality mismatch — rather than citing a general low match score. It does not escalate a case the policy allows it to discharge.

03

Document Verification Agent

Checking whether customer-submitted documents satisfy the customer's own document policy and requirements.

Automatically verify whether customer documents match your policy and requirements www.godiligent.ai

Mapped capabilities

4 capabilities

  • Document type and completeness

    Identifying the document class and detecting missing pages, fields, or signatures.

  • Policy requirement matching

    Testing a document against the uploaded requirements (issuer, validity window, address match).

  • Cross-document consistency

    Reconciling names, addresses, and entity details across a submitted document set.

  • Rejection reasoning and remediation ask

    Stating why a document fails and what the customer must resubmit.

04

Policy Grounding and Agent Configuration

The customer uploads their own risk policies and procedures so checks run in a customized, auditable way; the agent's behavior must follow that configuration rather than a generic default.

Mapped capabilities

4 capabilities

  • Adherence to uploaded policy

    Applying customer-specific thresholds and rules over generic AML heuristics.

  • Workflow scoping

    Operating within the workflow the agent was configured for and not acting outside it.

  • Policy gap and ambiguity handling

    Surfacing cases the uploaded policy does not cover instead of improvising a rule.

  • Policy change propagation

    Reflecting an updated policy in subsequent decisions consistently.

05

Human-in-the-Loop Oversight and Auditability

Agents deploy with analyst review in the loop until teams scale automation; every decision needs to be traceable, structured, and reviewable by a second line or auditor.

Mapped capabilities

4 capabilities

  • Evidence citation and traceability

    Tying each stated finding to a retrievable source rather than asserting it.

  • Escalation and abstention

    Routing low-confidence or out-of-policy cases to an analyst instead of deciding.

  • Audit trail completeness

    Recording what was checked, what was found, and which policy clause drove the outcome.

  • Analyst override handling

    Accepting and preserving a reviewer's correction without silently reverting to the agent's view.

06

Integration and Data Handling

Agents connect via API, the Diligent Portal, or native screening-tool integrations, under enterprise data commitments: no training on customer data, siloed per-customer environments, and SOC 2 Type II / ISO 27001 / GDPR posture.

Data is stored in siloed environments, isolated from other customer data. www.godiligent.ai

Mapped capabilities

4 capabilities

  • Screening-tool ingestion

    Consuming alert payloads from an upstream sanctions/PEP/adverse-media provider without dropping fields.

  • Customer data isolation

    Keeping customer policies, configurations, and case data within their own environment.

  • Sensitive data handling in outputs

    Not exposing personal data beyond what the case and policy require.

  • Failure and degraded-source behavior

    Reporting an unreachable registry or provider as unchecked rather than assuming a clean result.

Coverage is mapped from Diligent AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Diligent AI test?+

The coverage map is generated from Diligent AI's own public product surface (KYC/AML compliance AI agents for financial crime operations): 6 scoring areas — Merchant Risk Investigation Agent, AML Screening Alert Remediation, and Document Verification Agent, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Diligent AI evals scored?+

Every case generated for Diligent AI — across Merchant Risk Investigation Agent and AML Screening Alert Remediation and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Diligent AI library include?+

The full Diligent AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Web and social evidence gathering and Registry filing interpretation under Merchant Risk Investigation Agent); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Diligent AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Diligent AI areas and set them up in a Corsac workspace, where you can run every test case against Diligent AI or your own agent with your own data.