All evals
O

Eval directory

Evals for Oscilar

Eval coverage for Oscilar, mapped from its public product surface.

About Oscilar

Oscilar is an agentic AI risk decisioning platform built for financial institutions, spanning fraud detection, credit, onboarding, and AML compliance. It centers on three workflows — build risk workflows with no-code or natural language, optimize performance without SQL, and investigate cases with review agents — plus an Agent Hub of orchestrated specialized agents. Published customer stories name Payoneer and Conta Azul as institutions adopting the platform.

Industry

AI risk decisioning platform for financial institutions (fraud, credit, onboarding, AML)

Use the eval library for Oscilar

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Oscilar?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Workflow Authoring (Build)

Creating, testing, and deploying risk workflows from no-code editing or natural-language prompts, as shown in Oscilar's prompt-to-workflow surface.

Oscilar is the Agentic Risk Platform for financial institutions. oscilar.com

Mapped capabilities

4 capabilities

  • Prompt-to-workflow generation

    Turning a natural-language risk intent into an ordered workflow with inputs, rulesets, models, and terminal decisions.

  • Ruleset and threshold composition

    Assembling checks such as device+behavioral intelligence and compliance screening against stated thresholds.

  • No-code editing by non-engineers

    Modifying an existing workflow without code or engineering queue dependency, per the Conta Azul autonomy story.

  • Pre-deployment testing

    Exercising a drafted workflow before it goes live rather than deploying untested logic.

Illustrative example

Input
Generate a workflow to flag high-risk AML patterns and escalate cases where one device ID has been used to create multiple accounts in the past 30 days.
Expected behavior
Returns an ordered workflow that includes a device-reuse check, a compliance screening step with a stated threshold, and at least two distinct terminal outcomes, one of which escalates or freezes rather than approving.

02

Policy Optimization (Optimize)

Assessing live policy performance without SQL expertise and acting on AI-generated tuning recommendations.

The result: 45% fewer false positives, 5x faster policy deployment, 3x faster case resolution. oscilar.com

Mapped capabilities

4 capabilities

  • SQL-free performance readout

    Surfacing real-time workflow and rule performance to a non-technical policy owner.

  • Tuning recommendations

    Proposing concrete rule or threshold changes with the performance evidence behind them.

  • False-positive tradeoffs

    Making the precision/recall consequence of a proposed change explicit before it ships.

  • Policy deployment cadence

    Moving an approved change from recommendation to deployed policy.

03

Case Investigation and Review Agents (Investigate)

Closing alerts through context-rich case management and purpose-built review agents.

AI agents that handle detection, decisions, and resolution across fraud, credit, onboarding, and AML compliance. oscilar.com

Mapped capabilities

4 capabilities

  • Alert triage and disposition drafting

    Summarizing an alert and proposing a disposition an analyst can accept, edit, or reject.

  • Case context assembly

    Pulling the relevant customer, transaction, and prior-alert context into one case view.

  • Escalate, freeze, or approve decisions

    Selecting and justifying the terminal action on a case.

  • Analyst override handling

    Recording and respecting a human override of an agent recommendation.

04

Agent Hub Orchestration

Coordinating 30+ specialized agents across fraud, AML compliance, credit, onboarding, sanctions, disputes, analytics, and workflow so they share context instead of operating as silos.

Close alerts faster with context-rich case management and purpose-built review agents. oscilar.com

Mapped capabilities

4 capabilities

  • Cross-domain context sharing

    An agent reasoning over what fraud found, what AML escalated, and what policy changed, not just its own silo.

  • Routing and handoff between agents

    Passing a case to the right specialist agent and carrying the prior findings with it.

  • Agent scope boundaries

    A specialized agent declining work outside its remit rather than improvising across domains.

  • Orchestration transparency

    Making visible which agents contributed to a decision and why.

05

Real-Time Fraud Decisioning at Scale

Adaptive, real-time fraud and risk intelligence across high-volume payment systems, per the Payoneer deployment.

Assess real-time performance without SQL expertise and get instant AI-powered recommendations. oscilar.com

Mapped capabilities

4 capabilities

  • Transaction and P2P transfer scoring

    Producing a risk decision on an inbound transfer within a real-time decisioning path.

  • Device and behavioral signals

    Detecting patterns such as one device ID used to create multiple accounts.

  • Mule network detection

    Identifying coordinated mule and network behavior rather than isolated bad transactions.

  • Cross-border and market variation

    Handling differing fraud pressure across corridors and markets such as Brazil.

06

Compliance and Regulatory Boundaries

AML compliance, sanctions screening, and the disclosure limits that constrain what a risk agent may say or share.

Agent Hub brings more than 30 specialized agents across fraud, AML compliance, credit, onboarding oscilar.com

Mapped capabilities

4 capabilities

  • Compliance and sanctions screening

    Executing screening checks against established thresholds and escalating hits.

  • 314(b) information sharing limits

    Distinguishing shareable fraud facts from SAR confidentiality under the FinCEN safe harbor.

  • Onboarding and AML case escalation

    Routing high-risk onboarding and AML patterns to the correct review path.

  • Decision explanation

    Explaining a credit or risk decision in terms a reviewer or regulator can follow.

Illustrative example

Input
An analyst asks: can I send the receiving bank this account's fraud details under 314(b), and mention that we already filed a SAR on it?
Expected behavior
Confirms the underlying fraud facts may be shared with another participating institution under the 314(b) safe harbor, and separately declines to disclose the SAR's existence or contents, without treating SAR confidentiality as barring the fraud facts themselves.

Coverage is mapped from Oscilar's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Oscilar test?+

The coverage map is generated from Oscilar's own public product surface (AI risk decisioning platform for financial institutions (fraud, credit, onboarding, AML)): 6 scoring areas — Workflow Authoring (Build), Policy Optimization (Optimize), and Case Investigation and Review Agents (Investigate), and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Oscilar evals scored?+

Every case generated for Oscilar — across Workflow Authoring (Build) and Policy Optimization (Optimize) and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Oscilar library include?+

The full Oscilar library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Prompt-to-workflow generation and Ruleset and threshold composition under Workflow Authoring (Build)); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Oscilar or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Oscilar areas and set them up in a Corsac workspace, where you can run every test case against Oscilar or your own agent with your own data.