All evals
Resemble AI

Eval directory

Evals for Resemble AI

Eval coverage for Resemble AI, mapped from its public product surface.

About Resemble AI

Resemble AI is an enterprise generative AI security platform for verifying, detecting, and acting on synthetic audio, image, and video, deployable on-prem or in the cloud. Its products span Verify (Resemble Watermarker, Resemble Identity), Detect (Resemble Detect, Resemble Meetings), and audio editing/enhancement APIs, built on proprietary models such as DETECT-3B Omni and PerTh. Pricing runs from a free credit-based Flex tier to Team ($350/mo) and Business ($1,000/mo) plans, with enterprise options available on request.

Industry

generative AI security (deepfake detection & watermarking)

Use the eval library for Resemble AI

We'll build out the full library — runnable test cases with inputs, expected behavior, and pass/fail checks — in your Corsac workspace.

Generate your own →

Coverage map

What would you measure for Resemble AI?

6 scoring areas · 24 capabilities mapped · grounded in 8 cited pages

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Multimodal deepfake detection

Core Resemble Detect surface: a single DETECT-3B Omni call that returns a verdict on audio, image, or video, with stated coverage of 160+ generative models, 51 languages, and sub-second latency, available via API on-prem or in the cloud.

Battle-tested against 160+ generative AI models www.resemble.ai

Mapped capabilities

4 capabilities

  • Audio detection verdicts

    Synthetic speech identified across languages and telephony-grade or compressed audio, returning a verdict rather than a bare score.

  • Image and video detection verdicts

    Face swaps and synthetic frames detected in uploaded images and video, including frame-level localization.

  • Zero-day and unseen-generator coverage

    Behavior on media from generators outside the enumerated 160+ model set, per the claimed under-one-hour zero-day coverage.

  • Authentic-media handling

    Genuine recordings are not flagged as synthetic; verdicts on real content are stable and clearly stated.

Illustrative example

Input
A 12-second phone-quality WAV of a cloned executive voice, submitted to the detection endpoint with explainability enabled.
Expected behavior
The response returns an explicit synthetic-versus-authentic verdict, not only a numeric score, and attaches a human-readable rationale naming the specific audio artifacts that drove the decision.

02

Explainability and forensic evidence

Resemble Intelligence layer that turns a verdict into human-readable forensic rationale for compliance, legal review, and trust & safety, including artifact attribution, heatmaps, and Reverse Search context.

Mapped capabilities

4 capabilities

  • Human-readable rationale

    Explanation names which artifacts contributed to the verdict and why, in language a non-ML reviewer can act on.

  • Frame-by-frame localization

    Heatmaps or segment pointers identify where in the file manipulation was detected.

  • Reverse Search context

    Supplementary provenance or prior-appearance context attached to a detection result.

  • Audit-ready reporting

    Verdict, evidence, and timestamps assembled into a record suitable for downstream compliance review.

03

Watermarking and content provenance

Resemble Verify surface: PerTh psychoacoustic audio watermarking plus image and video watermarking, C2PA record signing, SynthID checking, and the Inspector SDK, positioned for EU AI Act Article 50 compliance.

Survives MP3 compression, audio editing, noise, and codec transforms. www.resemble.ai

Mapped capabilities

4 capabilities

  • Watermark embedding and decode

    Imperceptible marks applied at creation and recovered later, at the stated decode reliability.

  • Robustness to transforms

    Marks survive MP3 compression, editing, added noise, and codec transforms.

  • C2PA signing and SynthID checks

    Provenance records signed and third-party SynthID marks inspected on submitted files.

  • Inspector SDK integration

    Programmatic verification of watermark presence from a customer application.

Illustrative example

Input
A watermarked WAV re-encoded to 128 kbps MP3, then submitted for watermark inspection via the Inspector SDK.
Expected behavior
Inspection recovers the embedded watermark from the compressed file and reports it as present, with the payload matching what was embedded before re-encoding.

04

Voice identity verification

Resemble Identity surface built on Resemblyzer: enroll a speaker from four seconds of audio, return a match distance on later calls, maintain fraud watchlists, and pair with Detect so identity checks also confirm the audio is real.

Enroll any speaker from 4 seconds of audio www.resemble.ai

Mapped capabilities

4 capabilities

  • Enrollment and match scoring

    Profile created from short audio; subsequent audio returns a match distance and verified-or-flagged outcome.

  • Impersonation and clone rejection

    Voice impersonation and AI clones of an enrolled speaker are flagged rather than verified.

  • Replay attack handling

    Pre-recorded audio of an authorized speaker is flagged.

  • Fraud watchlist alerting

    Enrolled bad actors trigger alerts across channels regardless of the name or account used.

05

Real-time meeting monitoring

Resemble Meetings surface: a detection bot that joins Zoom, Teams, Meet, and Webex calls via calendar sync, flags face swaps, voice clones, and synthetic personas during the call, and leaves a forensic audit trail.

An automated detection bot joins every Zoom, Teams, Meet, and Webex call. www.resemble.ai

Mapped capabilities

4 capabilities

  • Bot join and calendar sync

    Bot attaches to scheduled meetings across the four supported platforms, appearing in the attendee list.

  • In-call alerting

    Confidence scores and recommended actions surface in-platform before the call ends.

  • Live participant analysis

    Audio and video from a session analyzed for face swaps, voice clones, and synthetic personas.

  • Post-call forensic trail

    Detection report retained for after-the-fact review.

06

Platform, deployment, and audio APIs

Integration and account surface: async job lifecycle for the Audio Edit and Audio Enhancement endpoints, plus plan tiers (Flex, Team $350/mo, Business $1,000/mo), seat and SSO entitlements, batch and large-file uploads, and on-prem, air-gapped, or cloud deployment.

Get volume pricing, enterprise SLAs, custom model training, on-prem deployment, dedicated support www.resemble.ai

Mapped capabilities

4 capabilities

  • Audio Edit inpainting

    Targeted regeneration of a changed segment stitched back into the original recording, with webhook callbacks.

  • Audio Enhancement parameters

    Noise removal, loudness normalization, and studio processing applied by default and independently toggleable per call.

  • Async job lifecycle

    Submit a job, receive a UUID, poll for completion, and download the result on finish.

  • Plan entitlements and deployment mode

    Credit and tier limits, seat counts, SSO, batch/large-file access, and on-prem versus cloud selection behave per the published plan.

Coverage is mapped from Resemble AI's public pages (8 crawled). Examples are illustrative, not real test cases. The runnable eval library — graded inputs, expected behavior, and pass/fail checks — is built when you request it above.

Frequently asked questions

What do the Corsac evals for Resemble AI test?+

The coverage map is generated from Resemble AI's own public product surface (generative AI security (deepfake detection & watermarking)): 6 scoring areas — Multimodal deepfake detection, Explainability and forensic evidence, and Watermarking and content provenance, and more — spanning 24 mapped capabilities, each graded on adversarial robustness, workflow quality, safety gates, and operator quality once the library is built.

How are the Resemble AI evals scored?+

Every case generated for Resemble AI — across Multimodal deepfake detection and Explainability and forensic evidence and the other mapped areas — is graded with pass/fail checks plus an LLM judge scoring 1–5 against its expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published.

How many test cases does the Resemble AI library include?+

The full Resemble AI library is built on request. The coverage map spans 6 areas and 24 capabilities (for example, Audio detection verdicts and Image and video detection verdicts under Multimodal deepfake detection); each becomes graded test cases — inputs, expected behavior, pass/fail checks — in your Corsac workspace.

How do I run these evals against Resemble AI or my own agent?+

Request the library with your work email above. We'll build out all 6 mapped Resemble AI areas and set them up in a Corsac workspace, where you can run every test case against Resemble AI or your own agent with your own data.