Agent memory without replaying the whole conversation
In plain terms
Most agents remember by pasting the entire history into every request, which gets slow, expensive and eventually stops fitting. This design splits memory into tiers, built from open-source parts, so each turn carries only what it needs.
The easy way to give an memory is to send it everything, every turn. It works well until conversations get long. Then each turn gets slower and costs more, a decision from forty turns ago gets buried under tool output, and eventually the history doesn't fit in the at all.
Adding a vector database helps less than people expect, because memory isn't one thing. What the user prefers, what happened last Tuesday and how to format a report are different kinds of memory, and they need different handling.
Before building anything
Full history is often the most accurate option while conversations are short. In the Mem0 paper's own results, the full-context baseline beat every memory system on accuracy; the memory systems won on latency and cost. Another benchmark found full context well ahead for the first few dozen conversations.
So I keep full context until it becomes a problem, and when I replace it I keep a full-context baseline in the evals, so I know what the savings cost.
Tiers
Pinned
Who the user is, standing instructions, the current goal. Short and edited in place, because it costs tokens every turn.
Writing memory
Memory gets written after the agent replies, by a background job, so the user never waits for it. The job reads the new turns alongside the rolling summary and extracts what's worth keeping.
The decision that matters is what happens when something changes. If the user says the deadline moved, I add the new fact and mark the old one as no longer valid, instead of overwriting it. Graphiti stores two timestamps per fact (when it was true, and when the system learned it), and Mem0 switched from overwriting to add-only in its 2026 redesign. Overwrites decided by a model lose information you'll want later.
Reading memory
Before each turn, search the stores with the current request using vector similarity, keyword match and, if you have one, the entity graph. Merge, rerank, and take results up to a fixed budget. The published systems retrieve roughly 1.6 to 7 thousand tokens per turn, against full histories of 26 thousand to over 100 thousand.
Options
| Approach | Accuracy, long histories | Cheap per turn | Easy to build | Best for |
|---|---|---|---|---|
| Replay full history | Short conversations; eval baseline | |||
| Last N turns only | When old turns don't matter | |||
| Rolling summary | Continuity on a budget | |||
| Files the agent reads and writes | A simple, strong start | |||
| Tiered memory (this design) | Long-lived agents, many sessions | |||
| Temporal knowledge graph | Facts that change; relationships |
The files option is worth trying first. Letta reported that an agent on a small model, keeping its history in files it could search, scored well on the LoCoMo memory benchmark. A folder of Markdown notes and tools to read and edit them goes a long way.
An open-source stack
- Models: any open-weights model on vLLM, llama.cpp or Ollama. Give memory extraction a stronger model than you'd expect; Graphiti's docs warn that small models without structured output break ingestion.
- Storage: Postgres with pgvector and full-text search handles episodes, facts and hybrid search in one place. Qdrant if you'd rather keep vectors separate.
- Graph, optionally: Graphiti on Neo4j or FalkorDB, when facts change over time and relationships matter.
- Frameworks: Letta if the agent should manage its own memory, Mem0 as a layer beside an existing agent, LangGraph's store and LangMem if you're already in that ecosystem.
One turn, sketched
BUDGET = {"pinned": 800, "summary": 1200, "window": 6000, "retrieved": 4000}
def run_turn(user_id, message):
mem = memory_for(user_id) # isolated per user
context = [
fit(mem.pinned(), BUDGET["pinned"]),
fit(mem.summary(), BUDGET["summary"]),
fit(mem.recent_turns(n=8), BUDGET["window"]),
fit(mem.search(message), BUDGET["retrieved"]),
]
reply = agent.respond(context, message)
mem.log_episode(message, reply)
background(consolidate, user_id) # off the hot path
return reply
def consolidate(user_id):
mem = memory_for(user_id)
for fact in extract_facts(mem.summary(), mem.recent_turns(n=10)):
for old in mem.conflicting(fact):
old.valid_until = fact.valid_from # invalidate, don't delete
mem.add_fact(fact)
mem.refresh_summary()Failure modes
Summary poisoning. One wrong fact in the summary shapes every turn after it. Keep the raw episodes and check against them when it matters.
Leaking between users. Scope every read and write by user and tenant, and make deletion real across every tier.
Never forgetting. Stale facts crowd out useful ones. Decay by recency, expire what hasn't been touched, merge duplicates during consolidation.
Breaking the prompt cache. Every time you clear old turns, cached prompts are invalidated. Clear less often, in bigger batches.
Trusting published scores. Memory vendors have publicly disputed each other's numbers on the same benchmark. Run your own evaluation, full-context baseline included. LongMemEval is a good one to start with: it tests extraction from long histories, reasoning across sessions and over time, updated facts, and knowing when to say "I don't know".
Sources
- Packer et al., MemGPT: Towards LLMs as Operating Systems (2023); Letta docs on memory and sleep-time agents
- Chhikara et al., Mem0 (2025), and Mem0's 2026 algorithm update
- Rasmussen et al., Zep: A Temporal Knowledge Graph Architecture for Agent Memory (2025)
- Park et al., Generative Agents (2023)
- Xu et al., A-MEM: Agentic Memory for LLM Agents (2025)
- Wu et al., LongMemEval (ICLR 2025); Maharana et al., LoCoMo (2024)
- Anthropic, Effective context engineering for AI agents (2025); OpenAI Agents SDK session docs
- Hsieh et al., RULER (NVIDIA, 2024); Modarressi et al., NoLiMa (2025); Chroma, Context Rot (2025)
- Rehberger, SpAIware — persistent memory injection in ChatGPT (2024)
- Anthropic memory tool documentation; Timescale, pgvector vs Pinecone (vendor benchmark)
Repos to explore
Open-source projects worth reading alongside this. Each belongs to its authors; check the licence before using it.
- letta-ai/lettaAgents that manage their own memory (MemGPT lineage), with background consolidation.
- mem0ai/mem0A memory layer beside an existing agent; self-hostable.
- getzep/graphitiTemporal knowledge graph with bi-temporal facts.
- langchain-ai/langgraphShort-term checkpoints and a long-term memory store.
- langchain-ai/langmemExtraction and consolidation for semantic, episodic and procedural memory.
- agiresearch/A-memZettelkasten-style agentic memory (research code).
- pgvector/pgvectorVector search inside Postgres, next to full-text search.
- qdrant/qdrantA dedicated vector database if you'd rather separate it.
- xiaowu0162/LongMemEvalBenchmark for long-term memory, including updated facts and abstention.
Read next
Chunking strategies, compared
How you cut documents into pieces decides what your AI can find. Seven ways to do it, what each costs, and the order I try them in.
RAG for long documents: give it a map
Ordinary RAG answers 'find me the paragraph' well and 'what does this 400-page file say overall' badly. Keeping the document's structure, with summaries at each level, lets one system answer both.