01
Multimodal deepfake detection
Core Resemble Detect surface: a single DETECT-3B Omni call that returns a verdict on audio, image, or video, with stated coverage of 160+ generative models, 51 languages, and sub-second latency, available via API on-prem or in the cloud.
“Battle-tested against 160+ generative AI models” www.resemble.ai
Mapped capabilities
4 capabilities
Audio detection verdicts
Synthetic speech identified across languages and telephony-grade or compressed audio, returning a verdict rather than a bare score.
Image and video detection verdicts
Face swaps and synthetic frames detected in uploaded images and video, including frame-level localization.
Zero-day and unseen-generator coverage
Behavior on media from generators outside the enumerated 160+ model set, per the claimed under-one-hour zero-day coverage.
Authentic-media handling
Genuine recordings are not flagged as synthetic; verdicts on real content are stable and clearly stated.
Illustrative example
- Input
- A 12-second phone-quality WAV of a cloned executive voice, submitted to the detection endpoint with explainability enabled.
- Expected behavior
- The response returns an explicit synthetic-versus-authentic verdict, not only a numeric score, and attaches a human-readable rationale naming the specific audio artifacts that drove the decision.




