Mem0 vs Hindsight: Agent Memory Layers Compared

7 min read

Mem0 vs Hindsight compared (Oct 2026): how each stores and retrieves memory, open-source vs hosted features, benchmarks and token cost, pricing, and which fits.

Mem0 vs Hindsight is a choice between two memory layers that sit beside your agent, not inside it. Mem0 keeps memory lean: one LLM pass extracts facts, and search returns a small set of them cheaply. Hindsight does more work per memory: it links entities and time, merges facts into observations in the background, and can reason over them with reflect.

Both are open source and both have a hosted version. This page compares them on memory model, retrieval, what’s free versus paid, benchmarks and cost, using each project’s README, docs and pricing pages as checked on 8 October 2026.

What are Mem0 and Hindsight?

Mem0 is an Apache 2.0 memory layer that uses an LLM to pull durable facts from conversations, stores them per user, agent or session, and returns the relevant ones on search. Hindsight is an MIT-licensed memory server from Vectorize that extracts facts, entities and time data into memory banks, searches them four ways in parallel, and reasons over them with a reflect operation.

Mem0 runs as a Python or TypeScript library, a self-hosted Docker server, or the hosted Mem0 Platform. It had about 66,800 GitHub stars in October 2026, the most of any agent memory project. Our page on what Mem0 is covers it in depth.

Hindsight runs as a Docker container, a pip-installed server, a Helm chart, an embedded Python server, or the managed Hindsight Cloud. It had about 47,100 stars. The Hindsight overview covers its design and limits.

Neither runs your agent. Your code (or a framework integration) calls them before and after model turns.

Mem0 vs Hindsight at a glance

Mem0Hindsight
LicenseApache 2.0MIT
Core callsadd, search, update, deleteretain, recall, reflect
What gets storedShort extracted facts with embeddings and entitiesWorld facts, experiences, observations, mental models
Write pathOne LLM pass, ADD-onlyLLM extraction of facts, entities, time and causal links, then background consolidation
RetrievalSemantic + BM25 + entity matching, fusedSemantic + BM25 + graph + temporal, fused and reranked
Reasoning over memoryNo; your prompt does itreflect returns a cited answer
Graph in open sourceRemoved in April 2026 (Platform only)Included
SDKsPython, TypeScript, REST, CLIPython, TypeScript, Go, REST, CLI
MCPHosted MCP server (Platform key)MCP endpoint per bank, self-hosted or cloud
StorageVector store of your choice; server uses Postgres + pgvectorPostgreSQL + pgvector, Oracle 23ai, or embedded pg0
Hosted pricingFree; Starter $19/mo; Pro $249/moUsage-based, free credits to start

Sources: the two READMEs, Mem0’s Platform vs Open Source page and pricing page.

How Mem0 stores and retrieves memory

Mem0 is built to be cheap per call. After an exchange, your code sends the messages to add. Since the April 2026 algorithm, an LLM extracts facts in a single pass and Mem0 only adds: “Memories accumulate; nothing is overwritten,” in the README’s words. A changed fact sits next to the old one, and you call update or delete for corrections.

On search, Mem0 scores semantic similarity, BM25 keyword match and entity overlap in parallel and fuses them, in a single retrieval call with no agentic loop. You put the top results in your prompt.

The open-source SDK and the hosted Platform share these calls, but several features are now Platform-only: graph memory, temporal reasoning, memory decay and the Dream background consolidation process. The README’s benchmark scores also “reflect Mem0’s managed platform, which includes proprietary optimizations not available in the open-source SDK.”

How Hindsight stores and retrieves memory

Hindsight does more work at write time. retain sends content to an LLM that extracts facts, dates, entities and relationships into an isolated memory bank (often one per user or agent). Entities attach to the memories they appear in, and memories link through shared entities, semantic neighbors and causal links.

A background step merges related facts into observations: deduplicated beliefs with supporting evidence that get refined as new facts arrive. Mental models are standing answers you define, refreshed in the background and read without an LLM call.

recall runs four strategies in parallel: semantic, BM25 keyword, graph and temporal. Results are merged with reciprocal rank fusion, reranked with a cross-encoder and cut to a token budget. reflect is an agent loop that reads mental models, observations and facts and returns an answer with citations. The FAQ puts reflect at 1-10 seconds versus 50-500 ms for recall.

All of this ships in the open-source build; Hindsight Cloud is the managed option.

Code: the same task in each

Store a fact about a user, then fetch it before the next turn. Mem0, with the open-source library (defaults to OpenAI, needs OPENAI_API_KEY):

1from mem0 import Memory
2
3memory = Memory()
4memory.add("Alice moved to Berlin in March and works at a robotics startup.", user_id="alice")
5
6hits = memory.search("Where does Alice live?", filters={"user_id": "alice"}, top_k=5)
7context = "\n".join(h["memory"] for h in hits["results"])

Hindsight, with hindsight-client against a local server on port 8888:

 1from hindsight_client import Hindsight
 2
 3client = Hindsight(base_url="http://localhost:8888")
 4client.retain(
 5    bank_id="alice",
 6    content="Alice moved to Berlin in March and works at a robotics startup.",
 7    document_id="chat-2026-10-08",
 8)
 9
10results = client.recall(bank_id="alice", query="Where does Alice live?", max_tokens=2048)
11context = "\n".join(r.text for r in results.results)
12
13# Optional: let Hindsight reason over the bank and answer directly
14answer = client.reflect(bank_id="alice", query="What should I know before calling Alice?")

The shape is the same: write after a turn, read before the next. The difference is what happens on the server between the two calls.

Benchmarks and token cost

Both vendors publish scores, and both are self-run. Mem0’s own Mem0 vs Hindsight page lists them side by side:

BenchmarkMem0HindsightMem0 tokens / retrievalHindsight tokens / retrieval
LongMemEval94.494.66.7K23.9K
LoCoMo92.5926.9K36.2K
BEAM 1M64.173.96.7K43.6K
BEAM 10M48.664.1under 7Kabout 27K

Hindsight’s benchmark site reports the same Hindsight scores (94.6 LongMemEval, 92 LoCoMo, 73.9 BEAM 1M, 64.1 BEAM 10M) but no token counts. The token columns are Mem0’s measurements.

Read it this way: accuracy is close on the short benchmarks, Hindsight leads on the very long BEAM sets, and Mem0 returns far fewer tokens per retrieval. Mem0’s page frames the gap as Hindsight spending about 4x the tokens for its edge. Token count per retrieval also depends on settings: Hindsight’s max_tokens budget is yours to set. Mem0’s scores come from its managed Platform, not the open-source SDK.

Older comparisons quote much lower Mem0 numbers, from before its April 2026 algorithm. The LLM memory evaluation guide explains what these benchmarks test and why a small test on your own data matters more than vendor tables.

Self-hosting, hosting and cost

Mem0 open source costs only your LLM, embedder and vector store. Every add makes one extraction call. The self-hosted server adds Postgres with pgvector, a dashboard and API keys. The Platform’s free Hobby tier allows 10,000 adds and 1,000 retrievals a month; Starter is $19/month and Pro, which includes graph memory and Dream, is $249/month. Mem0 lists SOC 2 and HIPAA for its Platform.

Hindsight open source costs your LLM calls plus a PostgreSQL database. Retain and consolidation both use an LLM, so writes cost more than Mem0’s single pass; async retain and provider batch APIs reduce that. The FAQ asks for 4 GB of RAM minimum (8 GB recommended) to self-host. Hindsight Cloud is usage-based with no seat fee and free credits to start.

Mem0 or Hindsight: how to choose

If you…Pick
Want the smallest token cost per retrievalMem0
Need per-user preference memory in a chat or support appMem0
Want SOC 2 / HIPAA on a hosted service todayMem0 Platform
Want graph links, time queries and consolidation in open sourceHindsight
Want the memory layer to return a reasoned, cited answerHindsight (reflect)
Have very long histories (millions of tokens per user)Hindsight, on BEAM results
Run simple n8n-style flowsMem0, or plain files; Hindsight’s README calls itself possibly “overkill” there

Other tools take other positions: Zep and Graphiti track fact validity over time, Letta lets the agent edit its own memory, and Cognee builds graphs from documents. See Mem0 alternatives, Mem0 vs Letta and the LLM memory comparison. For the concepts, start with AI agent memory explained.