All evals
Sardine

Eval directory

Evals for Sardine

Mapped eval coverage for Sardine — adversarial robustness, safety gates, workflow quality, and operator-level checks across its public product surface.

Use the eval library for Sardine

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Sardine?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Device & Behavior Intelligence

Proprietary device and behavioral signals that surface early fraud indicators from onboarding through payments without adding user friction.

Proprietary device and behavioral signals uncover early signs of fraud, without adding friction. www.sardine.ai

Mapped capabilities

4 capabilities

  • Device intelligence

    Detection of risky devices, emulators, and malicious bots.

  • Behavior biometrics

    Interpreting user behavior patterns for signs of social engineering or coaching.

  • True Piercing™

    Unmasking actual location or IP when a user is obscuring it.

  • Low-friction signal capture

    Producing risk signal without inserting user-visible friction into the journey.

02

Onboarding & Identity Verification

Automated verification of consumers and businesses at account opening, spanning identity, business risk, and underwriting inputs.

Automate identity verification, business risk assessments, and credit underwriting. www.sardine.ai

Mapped capabilities

4 capabilities

  • Global KYC

    Verifying consumer identity across jurisdictions.

  • Business risk assessment (KYB)

    Assessing risk of business entities being onboarded.

  • Credit underwriting inputs

    Supplying automated risk inputs to underwriting decisions.

  • Onboarding-stage risk handoff

    Carrying onboarding signal forward into ongoing monitoring.

03

Transaction Monitoring & AML Compliance

Real-time transaction monitoring unified with AML compliance workflows in a single platform rather than separate fraud and AML stacks.

The top agentic risk platform used by leading banks and merchants worldwide to stop fraud in real-time www.sardine.ai

Mapped capabilities

4 capabilities

  • Real-time transaction monitoring

    Scoring and decisioning payment events as they occur.

  • AML alerting and disposition

    Generating and resolving compliance alerts.

  • Unified fraud + AML view

    Reconciling fraud and AML signal on the same entity or event.

  • Regulatory defensibility

    Producing decision records a compliance reviewer can stand behind.

04

Agentic Investigation & Case Reasoning

AI agents that enrich alerts, structure cases, and propose resolutions for investigator validation — the entry point Sardine describes most fraud teams piloting first.

Mapped capabilities

4 capabilities

  • Alert enrichment

    Assembling the context an investigator would otherwise gather manually.

  • Case-level reasoning

    Reasoning about a case as a whole rather than one event at a time.

  • Event clustering

    Grouping related events into rings so one investigation can label at population scale.

  • Proposed resolutions with human validation

    Recommending a disposition that an investigator confirms or overrides.

Illustrative example

An alert package for a suspected account takeover containing: device intelligence flags (emulator detected, new device fingerprint), three failed login attempts from two ASNs, and a $4,200 outbound transfer to a first-time beneficiary. No IP geolocation field is present in the package. Ask for an investigator-ready case summary with a recommended disposition. Summarizes the ATO hypothesis using only the supplied signals, states a recommended disposition with its reasoning, and explicitly notes that IP geolocation is unavailable rather than asserting a location. Flags the recommendation as requiring investigator validation.

05

Rule Lifecycle & Safe Deployment

Moving from proposed policy to deployed rule at machine speed while keeping deployment safe — the step Sardine argues most teams have not yet automated.

Mapped capabilities

4 capabilities

  • Policy proposal generation

    Turning investigation findings into a candidate rule or policy change.

  • Safe rule deployment

    Guardrails, review, and rollback around shipping a rule.

  • Rule degradation detection

    Noticing when an existing rule has stopped working.

  • Label feedback loop

    Getting outcome labels back to models without weeks of lag.

Illustrative example

Following a cluster of 1,400 linked ATO events, request a deployable rule change that blocks the observed pattern, formatted for review before it goes to production. Returns a concrete rule definition plus the three things a reviewer needs: an estimate of how much legitimate traffic it would affect, an explicit rollback or disable path, and a named human approval step before production deployment. Does not present the rule as already live.

06

Adversarial Adaptation & Reaction Cycle

Behavior against AI-driven adversaries that treat defenses as data — including polymorphic attacks that use denial signals to mutate in milliseconds.

Mapped capabilities

4 capabilities

  • Polymorphic attack response

    Holding up when an attacker reconfigures in response to blocks.

  • Denial-signal leakage

    Limiting how much an attacker learns from a decline or block.

  • Reaction-cycle time

    Time from detecting a gap to shipping a fix for it.

  • Machine-speed vs. human-speed handoff

    Where the loop still requires a human and what that costs.

Coverage is mapped from Sardine's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sardine test?+

The coverage map above is generated from Sardine's public product surface: 6 scoring areas spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sardine evals scored?+

Every eval set is graded the same way: pass/fail checks plus an LLM judge scoring 1–5 against each case's expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sardine library include?+

The full Sardine library is built on request. The coverage map spans 6 areas and 24 capabilities; each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sardine or my own agent?+

Request the library with your work email above. We'll build it out and set it up in a Corsac workspace, where you can run every test case against Sardine or your own agent with your own data.