
Eval directory
Evals for Mem0
8 evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Mem0 AI products.
About Mem0
Mem0 is a memory layer for AI agents and assistants — it extracts, stores, and retrieves long-term facts across sessions via an add/search API, with user/agent/run scoping and optional graph memory, available as a managed Platform and open source.
How complete this published benchmark library is across datasets, metrics, rubrics, use-case maps, and pack context. This is library coverage, not an agent performance score.
Test datasets
8/8 packs
Scoring metrics
0/8 packs
Judge rubrics
8/8 packs
Use-case maps
0/8 packs
Pack context
8/8 packs
Available eval packs for Mem0
8 packs ready to run.
Add Memory
Knowledge RetentionEvaluates Mem0's Add Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Graph Memory
Knowledge RetentionEvaluates Mem0's Graph Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Memory Extraction And Consolidation
Knowledge RetentionEvaluates Mem0's Memory Extraction & Consolidation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Memory Lifecycle
Knowledge RetentionEvaluates Mem0's Memory Lifecycle (get/update/delete/history) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Platform Org Project Webhooks Config
Knowledge RetentionEvaluates Mem0's Platform: Org/Project, Webhooks & Config across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Safety Pii And Governance
PII LeakageKnowledge RetentionEvaluates Mem0's Safety, PII & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Scoping And Identity
Knowledge RetentionEvaluates Mem0's Scoping & Identity across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Search Memory
Knowledge RetentionEvaluates Mem0's Search Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.
Why eval Mem0 AI
Mem0's AI features ship behind brand promises about accuracy, safety, and reliability. Buyers and integrators need to know those promises hold up under adversarial prompts, edge-case workflows, and the long tail of real customer inputs — not just the demo path.
The Corsac eval library for Mem0 measures four dimensions teams care about most when deploying ai platform agents:
- Adversarial robustness — does the agent resist prompt injection, jailbreaks, and social-engineering attempts?
- Workflow quality— does it complete the task buyers were shown in the demo, on inputs that don't look like the demo?
- Safety gates — does it escalate or refuse when it should, and only then?
- Operator quality — does it preserve analyst trust by surfacing the right context at the right time?
Every eval pack above is hand-authored against Mem0's public product surface and runnable in Corsac with your own data.