All evals
SE

Eval directory

Evals for Silent Eight

Eval coverage for Silent Eight, mapped from its public product surface.

About Silent Eight

Iris 7 is Silent Eight's policy-bound agentic AI platform that makes financial crime compliance decisions at scale by replicating the investigative judgement of experienced analysts. Its AI Agents adjudicate alerts across sanctions, AML, fraud, and compliance risk for institutions facing high volume and regulatory scrutiny. The site emphasizes explainability, auditability, continuous learning from investigator feedback, and deployment in live production at large financial institutions.

Industry

agentic AI for financial crime compliance (AML, sanctions, fraud)

Headquarters

Singapore (HQ), with offices in New York, London, and Warsaw

Use the eval library for Silent Eight

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Silent Eight?

6 scoring areas · 23 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Alert adjudication and investigative judgement

The core claim: AI Agents adjudicate alerts across sanctions, AML, fraud, and compliance risk, reproducing the reasoning an experienced analyst would apply rather than emitting a score.

Iris 7 makes financial crime decisions at scale by replicating the investigative judgement of experienced analysts www.silenteight.com

Mapped capabilities

4 capabilities

  • Sanctions and watchlist name alert adjudication

    Resolving true matches versus false positives on watchlist hits, including name and entity ambiguity.

  • AML transaction and behaviour alert adjudication

    Reaching a disposition on transaction monitoring alerts with a documented investigative narrative.

  • Fraud and compliance-risk alert handling

    Adjudicating alert types outside sanctions and AML that the platform covers.

  • Decision versus score discipline

    Producing a defensible disposition rather than a probability, consistent with the product's decisions-not-scoring stance.

Illustrative example

Input
A payment for "Maria Santos", born 1991 in Lisbon, hits a sanctions list entry for "Maria Santos", born 1962 in Caracas, with no other shared identifiers.
Expected behavior
The agent discharges the alert as a false positive and states the discriminating attributes — mismatched date of birth and place of birth — as the basis for the decision.

02

Policy binding and risk appetite

Iris 7 is described as policy-bound: decisions must follow the institution's own written policy and configured risk appetite, not a generic model default.

Our decision-making AI Agents are deployed by institutions responsible for high-stakes decisions across sanctions, AML, fraud, and compliance risk www.silenteight.com

Mapped capabilities

4 capabilities

  • Adherence to institution-specific policy rules

    Decisions trace to the controlling policy clause for the alert type.

  • Risk appetite configuration and effect on outcomes

    Changing appetite settings changes dispositions in the expected direction.

  • Customisation per institution and jurisdiction

    Honouring tenant-specific customisation without leaking one institution's policy into another's.

  • Escalation and hand-back to human investigators

    Recognising when policy requires the decision be escalated rather than auto-adjudicated.

03

Explainability and auditability

The site foregrounds explainability and auditability as differentiators, so the rationale, evidence, and audit record are themselves testable surfaces.

Mapped capabilities

4 capabilities

  • Human-readable decision rationale

    Narratives an investigator or examiner can follow end to end.

  • Evidence and data-source attribution

    Every asserted fact in the rationale is tied to the source record it came from.

  • Audit trail completeness and reconstruction

    A past decision can be reconstructed with the inputs and policy version in force at the time.

  • Regulatory examination readiness

    Output holds up when an examiner asks why a specific alert was closed.

Illustrative example

Input
Ask the agent to explain a closed AML transaction-monitoring alert, then request the source record behind each factual claim in its explanation.
Expected behavior
Every factual claim in the explanation maps to a named source record — transaction, customer profile, or watchlist entry — and the agent asserts no customer fact it cannot attribute.

04

Continuous learning and model integrity

Silent Eight states models are maintained through continuous learning loops combining investigator feedback, regulatory updates, and risk signals, which raises drift and bias questions the FAQ itself names.

maintained through continuous learning loops that combine investigator feedback, regulatory updates, and risk signals to adapt in real time www.silenteight.com

Mapped capabilities

4 capabilities

  • Investigator feedback incorporation

    Corrections from investigators change subsequent behaviour on similar alerts.

  • Regulatory and watchlist currency

    Adapting to updated sanctions lists and regulatory changes.

  • Model drift detection and control

    Detecting and surfacing degradation rather than silently drifting.

  • Bias avoidance across names and geographies

    Comparable alerts are treated consistently regardless of name origin or region.

05

Deployment, integration, and data governance

Deployed in live production at large financial institutions, so integration with existing screening and case systems and the handling of sensitive customer data are first-class concerns.

with AI Agents delivering consistent, transparent outcomes across high-volume compliance operations www.silenteight.com

Mapped capabilities

4 capabilities

  • Integration with screening and case management systems

    Ingesting alerts and writing dispositions back into upstream and downstream systems.

  • Data security, privacy, and residency handling

    Sensitive customer and transaction data is handled per stated governance controls.

  • Scalability under high alert volume

    Behaviour holds up at the volumes the product is positioned for.

  • Availability and failure handling

    Degrading safely when an upstream data source or watchlist feed is unavailable.

06

Investigator workflow and review experience

The product is used by investigation teams, so the surfaces where a human inspects, accepts, or overrides an agent decision are part of what a product team would evaluate.

Mapped capabilities

3 capabilities

  • Review and override of an agent decision

    An investigator can disagree and record the override with reasons.

  • Queue handling and prioritisation

    Presenting adjudicated and escalated alerts in a workable order.

  • Consistency across repeat and related alerts

    The same fact pattern yields the same disposition across time and reviewers.

Coverage is mapped from Silent Eight's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Silent Eight test?+

The coverage map is generated from Silent Eight's own public product surface (agentic AI for financial crime compliance (AML, sanctions, fraud)): 6 scoring areas — Alert adjudication and investigative judgement, Policy binding and risk appetite, and Explainability and auditability, and more — spanning 23 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Silent Eight evals scored?+

Every case generated for Silent Eight — across Alert adjudication and investigative judgement and Policy binding and risk appetite and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Silent Eight library include?+

The full Silent Eight library is built on request. The coverage map spans 6 areas and 23 capabilities (for example, Sanctions and watchlist name alert adjudication and AML transaction and behaviour alert adjudication under Alert adjudication and investigative judgement); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Silent Eight or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Silent Eight areas and set them up in a Corsac workspace, where you can run every test case against Silent Eight or your own agent with your own data.