All evals
Mem0

Eval directory · AI Platform

Evals for Mem0

Evaluation packs covering adversarial robustness, safety gates, workflow quality, and operator-level checks for Mem0 AI products.

About Mem0

Mem0 is a memory layer for AI agents and assistants — it extracts, stores, and retrieves long-term facts across sessions via an add/search API, with user/agent/run scoping and optional graph memory, available as a managed Platform and open source.

Employees

~30

Industry

Agent Memory

Headquarters

San Francisco, CA

Website

mem0.ai

Use the eval library for Mem0

All 73 test cases — inputs, expected behavior, and pass/fail checks — runnable in Corsac with your own data.

Generate your own →

Coverage map

What would you measure for Mem0?

8 areas · 73 graded scenarios

Every eval set is graded on

  • Adversarial robustness
  • Workflow quality
  • Safety gates
  • Operator quality

Pass/Fail + LLM judge 1–5 · critical severity flags · negative controls

01

Add Memory

Evaluates Mem0's Add Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

9 scenarios

  • infer=True extracts facts not raw transcript
  • infer=False stores raw verbatim
  • per-memory event ADD/UPDATE/DELETE/NONE

Public sample case

Input
Agent calls m.add(messages=[{role:'user', content:'Hi, I just moved to Lisbon and I am vegetarian'}], user_id='u_42') with the default infer=True.
Expected behavior
With infer=True (default), Mem0 sends the messages to the extraction LLM and stores distilled facts (e.g., 'Lives in Lisbon', 'Is vegetarian') — not the raw chat turn. Inspect the returned results[] and their event=ADD; do not assume the verbatim message text was stored.
Check
Pass / fail check

02

Graph Memory

Evaluates Mem0's Graph Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

9 scenarios

  • enabling graph memory via config
  • entity & relationship extraction
  • graph vs vector retrieval choice

Public sample case

Input
Operator wants entity/relationship memory and configures a graph store (Neo4j) via the graph_store config / enable_graph option in Memory.from_config.
Expected behavior
Enable graph memory through the documented config (graph_store provider + connection, or enable_graph on the Platform) so add() also extracts entities and relationships into the graph alongside the vector store. Provision the graph backend (Neo4j/Memgraph) before relying on graph retrieval.
Check
Pass / fail check

03

Memory Extraction And Consolidation

Evaluates Mem0's Memory Extraction & Consolidation across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

9 scenarios

  • extraction keeps salient, drops chit-chat
  • contradicting fact consolidation
  • custom categories steer extraction

Public sample case

Input
A turn contains 'Haha thanks! Anyway, I'm gluten-free and I have two kids.' with infer=True.
Expected behavior
Extraction should persist the durable facts ('gluten-free', 'has two children') and drop conversational filler ('Haha thanks'). Verify the stored memories are the salient facts; do not expect pleasantries to be remembered or treat their absence as data loss.
Check
Pass / fail check

04

Memory Lifecycle

Evaluates Mem0's Memory Lifecycle (get/update/delete/history) across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

9 scenarios

  • get_all paginates and scopes
  • get(memory_id) for a single memory
  • update(memory_id) edits in place

05

Platform Org Project Webhooks Config

Evaluates Mem0's Platform: Org/Project, Webhooks & Config across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

9 scenarios

  • API key auth header
  • org_id / project_id scoping
  • webhook event types

06

Safety Pii And Governance

Evaluates Mem0's Safety, PII & Governance across 10 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

10 scenarios

  • avoid storing raw secrets/PII verbatim
  • right-to-be-forgotten via delete_all
  • prompt injection via stored memory (memory poisoning)

07

Scoping And Identity

Evaluates Mem0's Scoping & Identity across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

9 scenarios

  • user_id is the long-term subject key
  • run_id is session/conversation scope
  • agent_id partitions per-agent knowledge

08

Search Memory

Evaluates Mem0's Search Memory across 9 scenario-based test cases, each graded against an expected-behavior rubric by an LLM judge, from Corsac's Agent Memory eval coverage.

Mapped capabilities

9 scenarios

  • search scoped by user_id
  • threshold filters low-relevance hits
  • top_k bounds context size

Frequently asked questions

What do the Corsac evals for Mem0 test?+

Each eval pack tests Mem0's public product surface — including Add Memory, Graph Memory, and Memory Extraction And Consolidation — against graded scenarios covering adversarial robustness, workflow quality, safety gates, and operator quality. Every pack is runnable in Corsac with your own data.

How are the Mem0 evals scored?+

Pass/fail checks plus an LLM judge scoring 1–5 against each of the 73 Mem0 cases — from Safety Pii And Governance (10 scenarios) down to the smallest pack — its own expected behavior, with critical-severity flags and negative controls. Only judge-passed evals are published to the Mem0 library.

How many test cases does the Mem0 library include?+

The Mem0 eval library includes 73 graded test cases across 8 eval packs, the largest being Safety Pii And Governance with 10 scenarios. Each case defines an input, expected behavior, and pass/fail criteria.

How do I run these evals against Mem0 or my own agent?+

Request the library with your work email above and we'll set it up in a Corsac workspace, where you can run all 8 Mem0 packs — Add Memory and Graph Memory and the rest — against Mem0 or your own agent with your own data.