All evals
A

Eval directory

Evals for Alloy

Eval coverage for Alloy, mapped from its public product surface.

About Alloy

Alloy is an AI-powered identity, fraud, and compliance platform used across the full customer lifecycle by financial institutions and fintechs. It covers onboarding decisioning, ongoing fraud detection (including a machine learning model called Fraud Signal), perpetual KYC/KYB and AML checks, and risk-based authentication on login events. Recent product work centers on an AI Assistant that automates KYC/KYB review while exposing its chain of reasoning for auditability.

Industry

identity verification and fraud prevention platform for banks and fintechs

Website

alloy.com

Use the eval library for Alloy

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Alloy?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Onboarding decisioning

Applicant verification and approve/review/deny decisioning at account opening across acquisition channels, where the tension is approving good customers quickly without letting unnecessary friction drive drop-off.

AI-powered identity and fraud prevention platform that accelerates onboarding, stops fraud, and simplifies compliance www.alloy.com

Mapped capabilities

4 capabilities

  • Identity verification outcomes

    Resolving applicant identity from submitted attributes and returning a decision with the signals that drove it.

  • Decision policy execution

    Applying configured rules and thresholds to produce approve, manual review, or deny consistently.

  • Cross-channel consistency

    Same applicant, same policy outcome whether the application arrives digitally or in branch.

  • Friction calibration

    Adding step-up or document checks only where risk warrants it rather than uniformly.

02

Lifecycle fraud prevention

Fraud detection that continues past onboarding into transactions and account activity, including Fraud Signal, Alloy's machine learning model positioned as an alternative to static rules.

Ditch static rules for Fraud Signal, our machine learning model that adapts as threats evolve www.alloy.com

Mapped capabilities

4 capabilities

  • Fraud Signal scoring behavior

    Producing a risk score on account and transaction events and explaining its dominant contributors.

  • Adaptive vs. static rule handling

    Behavior when model output and legacy static rules disagree on the same event.

  • Coordinated attack detection

    Linking related applications or events that indicate ring activity rather than isolated actors.

  • Unified risk view

    Assembling onboarding, transaction, and account-activity signals into one view of a customer.

03

Perpetual KYC/KYB and AML compliance

Ongoing portfolio monitoring rather than one-time checks: real-time risk signals, business verification, watchlist screening, and the auditability a compliance program has to produce on demand.

Streamline perpetual KYC/KYB, AML, and fraud checks with fully auditable agentic automation www.alloy.com

Mapped capabilities

4 capabilities

  • Ongoing portfolio monitoring

    Re-screening existing customers as their risk profile or the underlying data changes.

  • KYB and business entity review

    Verifying businesses and associated parties, including beneficial-ownership-style structures.

  • Watchlist and sanctions match handling

    Surfacing potential matches and distinguishing true hits from name collisions.

  • Audit trail completeness

    Retaining a reviewable record of what was checked, when, and why a disposition was reached.

04

Risk-based authentication

Contextual, adaptive assessment of login and authentication events rather than a single pass/fail moment at sign-in, including self-service aggregation rules on login events.

Mapped capabilities

4 capabilities

  • Login event risk assessment

    Scoring a login using device, velocity, and behavioral context available at the time.

  • Step-up challenge decisioning

    Choosing when to escalate to additional verification versus allowing the session through.

  • Aggregation rules on login

    Self-service rules that aggregate multiple login events into a single risk determination.

  • Account-change event handling

    Treating password resets and profile changes as risk-bearing events, not routine ones.

Illustrative example

Input
A returning customer logs in successfully with correct credentials from a new device and a country not seen in their prior 90 days of activity.
Expected behavior
The login is scored as elevated risk and a step-up challenge is issued rather than an outright block, and the response identifies the new-device and new-geography signals as the drivers of the escalation.

05

AI Assistant review automation

The agentic layer that automates KYC/KYB review, where recent product work emphasizes exposing the agent's chain of reasoning in real time so analysts and auditors can trust and reconstruct each recommendation.

You can now follow the agent's chain of reasoning in real-time as it reviews an application. www.alloy.com

Mapped capabilities

4 capabilities

  • Chain-of-reasoning transparency

    Showing step-by-step which sources were checked, what was found, and how the conclusion followed.

  • Review recommendation quality

    Producing a defensible recommendation on an application a human reviewer would act on.

  • Source attribution

    Tying each asserted fact in the output back to the specific source that supports it.

  • Escalation to human review

    Handing off rather than deciding when evidence is thin, conflicting, or out of scope.

Illustrative example

Input
A KYB application where the submitted business address matches a registered agent's address but the secretary of state filing is active and in good standing.
Expected behavior
The assistant flags the registered-agent address as a known benign pattern rather than an adverse finding, and its reasoning trace names each source consulted and what that source returned before stating the recommendation.

06

Data orchestration and integrations

The pre-integrated data layer beneath decisioning: routing to third-party data sources, handling their responses, and keeping journeys performant as vendors change or degrade.

Alloy’s experience with 700+ clients and 250+ pre-integrated data solutions www.alloy.com

Mapped capabilities

4 capabilities

  • Data source routing

    Selecting and sequencing vendor calls to answer a given verification question.

  • Vendor response handling

    Behavior when a data source returns partial, conflicting, or empty results.

  • Journey configuration

    Composing and modifying decisioning journeys without breaking existing outcomes.

  • Degradation and recovery

    Failing safe when an upstream data provider is slow or unavailable.

Coverage is mapped from Alloy's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Alloy test?+

The coverage map is generated from Alloy's own public product surface (identity verification and fraud prevention platform for banks and fintechs): 6 scoring areas — Onboarding decisioning, Lifecycle fraud prevention, and Perpetual KYC/KYB and AML compliance, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Alloy evals scored?+

Every case generated for Alloy — across Onboarding decisioning and Lifecycle fraud prevention and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Alloy library include?+

The full Alloy library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Identity verification outcomes and Decision policy execution under Onboarding decisioning); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Alloy or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Alloy areas and set them up in a Corsac workspace, where you can run every test case against Alloy or your own agent with your own data.