All evals
C

Eval directory

Evals for Cardamon

Eval coverage for Cardamon, mapped from its public product surface.

About Cardamon

Cardamon AI provides AI agents that automate regulatory compliance work for regulated firms. The platform spans four products — Horizon Scanning, Obligation Mapping, Gap Analysis, and AI Regulation Search — covering monitoring of regulatory change, extraction of obligations from legal text, comparison of internal policies and controls against those obligations, and natural-language search over mapped regulations. Outputs are tailored to a firm's products, activities, markets, and risk methodology, with editing, assignment, and audit trails.

Industry

regulatory compliance AI agents for financial services

Headquarters

New York, US and London, UK

Use the eval library for Cardamon

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Cardamon?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Horizon Scanning — Change Detection and Triage

Ingesting regulatory updates, enforcement actions, and industry news from global feeds and configurable sources, then filtering to what is actually relevant to a specific firm's products, activities, and markets.

Continuously monitor and detect new regulatory updates, enforcement actions and industry news that is relevant specifically to you. cardamon.ai

Mapped capabilities

4 capabilities

  • Relevance filtering against firm profile

    Correctly keeps in-scope updates and suppresses ones that do not touch the firm's products, activities, or markets, with a stated reason.

  • Summarisation and categorisation

    Faithful summaries of the source update and stable category assignment, without introducing claims absent from the source.

  • Urgency and business-impact assessment

    Assigns urgency and impact tied to the firm's risk methodology, and links the change to affected business areas and products.

  • Recommended actions and findings reports

    Produces actionable next steps and findings output that trace back to the specific update rather than generic compliance advice.

Illustrative example

Input
Firm profile: UK-authorised e-money institution, payments only, no crypto or investment services. New feed item: an FCA consultation on financial promotions for cryptoasset firms.
Expected behavior
The update is marked not relevant to this firm, with a reason naming the mismatch between the crypto scope of the consultation and the firm's payments-only activities. It is not assigned a high urgency or a recommended action.

03

Gap Analysis — Controls Versus Obligations

Comparing ingested internal policies and control statements against mapped obligations to identify coverage, missing coverage, and remediation, with every decision traceable.

Our agents identify relevant changes, extract your obligations and conduct gap analyses in minutes not months cardamon.ai

Mapped capabilities

4 capabilities

  • Policy and control ingestion at usable granularity

    Internal documentation is chunked finely enough that a control statement can be matched to a specific obligation.

  • Coverage and missing-coverage identification

    Correctly reports an obligation as covered, partially covered, or unaddressed, and does not claim coverage from loosely related text.

  • Remediation suggestions tailored to frameworks

    Proposed remediations reference the firm's own frameworks and the specific gap, not boilerplate.

  • Traceability of match, miss, and suggestion

    Each result cites the obligation and the internal document passage that drove it.

04

AI Regulation Search — Grounded Regulatory Q&A

Natural-language questions answered strictly from the customer's mapped regulations, company context, and supplied policies and controls, including cross-jurisdiction comparison.

Our AI agents map your internal documentation against your regulatory obligations - identifying missing coverage in minutes cardamon.ai

Mapped capabilities

4 capabilities

  • Answer grounding and source citation

    Answers cite the mapped regulatory provisions they rest on and stay within the customer's corpus.

  • Abstention outside the mapped corpus

    Declines or flags scope limits when a question concerns a jurisdiction or instrument not mapped, instead of answering from general knowledge.

  • Cross-jurisdiction and L1–L4 comparison

    Compares requirements across jurisdictions and regulatory levels without conflating them.

  • Harmonisation of common versus unique obligations

    Identifies overlapping requirements across frameworks and isolates the jurisdiction-specific remainder.

Illustrative example

Input
Only EU and UK regulations are mapped for this tenant. Question: what are the customer due diligence thresholds for a money services business in Singapore?
Expected behavior
The agent states that Singapore requirements are not in the tenant's mapped regulations and does not supply thresholds from general knowledge. It may offer the EU or UK equivalent, clearly labelled as a different jurisdiction.

05

Firm Context and Tailoring

The configuration layer that makes outputs firm-specific: products, activities, markets, risk methodology and risk tags, and configurable sources — and whether that context is applied consistently across all four products.

We pull in data from regulatory feeds, newsletters, and government sources across the globe cardamon.ai

Mapped capabilities

4 capabilities

  • Firm profile applied to outputs

    Products, activities, and markets in the profile visibly drive relevance, applicability, and impact decisions.

  • Custom risk methodology and risk tags

    Uses the customer's own risk model and tag taxonomy rather than a default severity scheme.

  • Source configuration honoured

    Respects which regulatory feeds and sources the customer has enabled or excluded.

  • Consistency of context across products

    The same firm context yields non-contradictory conclusions between obligation mapping, gap analysis, and search.

06

Review Workflow, Editing and Audit Trail

The human-in-the-loop surface: filtering, editing, assignment, progress tracking, sharing, and the audit trail that records what changed, who approved it, and why.

Mapped capabilities

4 capabilities

  • Edits preserved with audit trail

    Human edits to an assessment persist and are recorded alongside the original AI output.

  • Assignment and review progress

    Items can be assigned to teammates and their review state is reflected accurately in dashboards.

  • Filtering on assessment points

    Results can be filtered on any assessment attribute and the filtered set matches the underlying data.

  • Sharing with intact source links

    Shared findings retain working links back to the originating regulatory source.

Coverage is mapped from Cardamon's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Cardamon test?+

The coverage map is generated from Cardamon's own public product surface (regulatory compliance AI agents for financial services): 6 scoring areas — Horizon Scanning — Change Detection and Triage, Obligation Mapping — Legal Text to Structured Obligations, and Gap Analysis — Controls Versus Obligations, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Cardamon evals scored?+

Every case generated for Cardamon — across Horizon Scanning — Change Detection and Triage and Obligation Mapping — Legal Text to Structured Obligations and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Cardamon library include?+

The full Cardamon library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Relevance filtering against firm profile and Summarisation and categorisation under Horizon Scanning — Change Detection and Triage); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Cardamon or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Cardamon areas and set them up in a Corsac workspace, where you can run every test case against Cardamon or your own agent with your own data.