All evals
S

Eval directory

Evals for Sedric

Eval coverage for Sedric, mapped from its public product surface.

About Sedric

Sedric is an AI compliance platform for banks, fintechs, and other regulated financial institutions that monitors marketing assets, customer communications, and partner/affiliate content against regulatory and internal policies. It covers pre-publication creative review, real-time agent assist on live calls and chats, post-interaction monitoring, and affiliate/partner oversight, mapping each flag to a specific regulation. The product is sold by modules and volume with a quote-based enterprise pricing model.

Industry

marketing & communications compliance AI for regulated finance

Headquarters

New York, NY

Website

sedric.ai

Use the eval library for Sedric

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Sedric?

6 scoring areas · 22 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Marketing Creative Review

Pre-publication review of marketing assets across copy, design, and video, with flags, explanations, and suggested fixes, delivered either inside Sedric or embedded in an existing creative workflow.

Sedric reviews 100% of customer communications, marketing assets, and partner activity against your policies. www.sedric.ai

Mapped capabilities

4 capabilities

  • Issue detection with explanation

    Flagging problematic copy or claims and stating why the passage is an issue rather than returning an opaque verdict.

  • Suggested fixes and clear fix lists

    Producing actionable remediation text a marketer can apply, aggregated into a reviewable fix list.

  • Claims library and brand guideline checks

    Evaluating assets against the customer's approved claims library and brand guidelines, not only external regulation.

  • In-workflow embedded review

    Operating inside the customer's existing creative stack without forcing a separate approval process.

Illustrative example

Input
Review this deposit ad for compliance: "Open a Premier Savings account today and earn 4.75% APY. No fees, no minimums, no catch."
Expected behavior
The asset is not approved. The response flags the stated APY as a triggering term that requires accompanying disclosures, names the specific advertising rule behind the flag, and offers corrected copy the marketer can use.

02

Real-Time Agent Assist

Live guidance during calls and chats that surfaces compliance risk at the moment it could occur, adapts prompts to conversation context, and closes out the interaction with a summary.

Monitor every call, message, chat, and AI interaction without relying on ~3% QA sampling. www.sedric.ai

Mapped capabilities

3 capabilities

  • Instant risk detection in-conversation

    Identifying prohibited language, missed disclosures, and emerging compliance risk while the interaction is still open.

  • Contextual prompts and next-best actions

    Delivering dynamic guidance shaped by conversation context instead of a rigid fixed script.

  • Automated call summaries

    Generating accurate, compliant post-interaction summaries that reduce wrap time and hold up as records.

03

Post-Interaction Communications Monitoring

Review of completed calls, messages, chats, and emails at full coverage rather than QA sampling, including AI-generated responses, with coaching driven by observed performance gaps.

Mapped capabilities

4 capabilities

  • Full-coverage review versus sampling

    Analyzing every interaction rather than a small sampled subset, and reporting coverage accordingly.

  • Multilingual detection and translation

    Detecting violations across the supported non-English languages and rendering them into English for reviewer consumption.

  • Monitoring of AI-generated responses

    Applying the same review to responses produced by automated agents as to human agents.

  • Contextual in-app coaching

    Targeting coaching at the specific performance gaps the monitoring surfaced.

04

Partner and Affiliate Oversight

Continuous monitoring of content published by affiliates, influencers, partners, and spokespeople on channels the institution does not control but remains liable for.

Mapped capabilities

4 capabilities

  • Affiliate domain and placement discovery

    Automatically identifying partner-owned sites, social accounts, paid ads, and landing pages connected to the customer's products.

  • Misleading claim and missing disclosure detection

    Catching exaggerated claims, absent disclosures, unapproved language, and product misrepresentation in partner content.

  • Risk-level prioritization

    Ordering findings by risk so remediation effort goes to the most exposed content first.

  • Partner reports and remediation tracking

    Generating violation reports partners can act on, and tracking the remediation through to resolution.

Illustrative example

Input
Scan this affiliate landing page: "Guaranteed approval, no credit check. Get your card in 24 hours." The page carries no disclosure of its paid relationship with the issuer.
Expected behavior
Two separate findings are returned: the guaranteed-approval claim as misleading, and the absent affiliate relationship disclosure. Each is mapped to a specific regulation or policy and carries a risk level so remediation can be prioritized.

05

Policy and Regulation Mapping

The layer that ties every finding to a specific regulation or internal policy, built on a compliance-tuned model configured with the customer's own policies alongside preloaded regulatory frameworks.

Every flag is linked to the underlying regulation, every override is logged with reasoning www.sedric.ai

Mapped capabilities

4 capabilities

  • Flag-to-regulation attribution

    Naming the specific regulation or internal policy behind each individual flag rather than a generic risk label.

  • Customer policy and guideline ingestion

    Operating against the organization's own written policies and brand guidelines as configured inputs.

  • Preloaded regulatory framework coverage

    Applying the regulatory frameworks the product ships with out of the box across US, UK, and EU regimes.

  • Explainable decisions

    Producing reasoning a compliance reviewer can inspect and defend rather than an unexplained score.

06

Audit Evidence and Governance

The record-keeping surface that makes the system's output usable in an examination: timestamped evidence, logged overrides with reasoning, and shareable reporting.

Generate timestamped, audit-ready evidence automatically. www.sedric.ai

Mapped capabilities

3 capabilities

  • Timestamped audit-ready evidence

    Automatically capturing findings as dated, retained evidence rather than transient alerts.

  • Override logging with reasoning

    Recording who overrode a flag and the stated justification, so decisions remain reconstructable.

  • Dynamic reporting

    Assembling findings into reports that show what is wrong and what must change, for internal and partner audiences.

Coverage is mapped from Sedric's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Sedric test?+

The coverage map is generated from Sedric's own public product surface (marketing & communications compliance AI for regulated finance): 6 scoring areas — Marketing Creative Review, Real-Time Agent Assist, and Post-Interaction Communications Monitoring, and more — spanning 22 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Sedric evals scored?+

Every case generated for Sedric — across Marketing Creative Review and Real-Time Agent Assist and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Sedric library include?+

The full Sedric library is built on request. The coverage map spans 6 areas and 22 capabilities (for example, Issue detection with explanation and Suggested fixes and clear fix lists under Marketing Creative Review); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Sedric or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Sedric areas and set them up in a Corsac workspace, where you can run every test case against Sedric or your own agent with your own data.