Skip to content
Ajith Thaduri

Agent memory without replaying the whole conversation

Built & tested5 min readPublished 23 Sept 2026

In plain terms

Most agents remember by pasting the entire history into every request, which gets slow, expensive and eventually stops fitting. This design splits memory into tiers, built from open-source parts, so each turn carries only what it needs.

The easy way to give an memory is to send it everything, every turn. It works well until conversations get long. Then each turn gets slower and costs more, a decision from forty turns ago gets buried under tool output, and eventually the history doesn't fit in the at all.

Adding a vector database helps less than people expect, because memory isn't one thing. What the user prefers, what happened last Tuesday and how to format a report are different kinds of memory, and they need different handling.

Before building anything

Full history is often the most accurate option while conversations are short. In the Mem0 paper's own results, the full-context baseline beat every memory system on accuracy; the memory systems won on latency and cost. Another benchmark found full context well ahead for the first few dozen conversations.

So I keep full context until it becomes a problem, and when I replace it I keep a full-context baseline in the evals, so I know what the savings cost.

Tiers

Tiered memory · click a stage
In every promptStored, retrieved on demand

Pinned

Who the user is, standing instructions, the current goal. Short and edited in place, because it costs tokens every turn.

Writing memory

Memory gets written after the agent replies, by a background job, so the user never waits for it. The job reads the new turns alongside the rolling summary and extracts what's worth keeping.

The decision that matters is what happens when something changes. If the user says the deadline moved, I add the new fact and mark the old one as no longer valid, instead of overwriting it. Graphiti stores two timestamps per fact (when it was true, and when the system learned it), and Mem0 switched from overwriting to add-only in its 2026 redesign. Overwrites decided by a model lose information you'll want later.

Reading memory

Before each turn, search the stores with the current request using vector similarity, keyword match and, if you have one, the entity graph. Merge, rerank, and take results up to a fixed budget. The published systems retrieve roughly 1.6 to 7 thousand tokens per turn, against full histories of 26 thousand to over 100 thousand.

Options

ApproachAccuracy, long historiesCheap per turnEasy to buildBest for
Replay full historyShort conversations; eval baseline
Last N turns onlyWhen old turns don't matter
Rolling summaryContinuity on a budget
Files the agent reads and writesA simple, strong start
Tiered memory (this design)Long-lived agents, many sessions
Temporal knowledge graphFacts that change; relationships
More dots is better. Full history stays accurate until it stops fitting, then fails outright.

The files option is worth trying first. Letta reported that an agent on a small model, keeping its history in files it could search, scored well on the LoCoMo memory benchmark. A folder of Markdown notes and tools to read and edit them goes a long way.

An open-source stack

  • Models: any open-weights model on vLLM, llama.cpp or Ollama. Give memory extraction a stronger model than you'd expect; Graphiti's docs warn that small models without structured output break ingestion.
  • Storage: Postgres with pgvector and full-text search handles episodes, facts and hybrid search in one place. Qdrant if you'd rather keep vectors separate.
  • Graph, optionally: Graphiti on Neo4j or FalkorDB, when facts change over time and relationships matter.
  • Frameworks: Letta if the agent should manage its own memory, Mem0 as a layer beside an existing agent, LangGraph's store and LangMem if you're already in that ecosystem.

One turn, sketched

python
BUDGET = {"pinned": 800, "summary": 1200, "window": 6000, "retrieved": 4000}

def run_turn(user_id, message):
    mem = memory_for(user_id)                    # isolated per user
    context = [
        fit(mem.pinned(), BUDGET["pinned"]),
        fit(mem.summary(), BUDGET["summary"]),
        fit(mem.recent_turns(n=8), BUDGET["window"]),
        fit(mem.search(message), BUDGET["retrieved"]),
    ]
    reply = agent.respond(context, message)
    mem.log_episode(message, reply)
    background(consolidate, user_id)             # off the hot path
    return reply

def consolidate(user_id):
    mem = memory_for(user_id)
    for fact in extract_facts(mem.summary(), mem.recent_turns(n=10)):
        for old in mem.conflicting(fact):
            old.valid_until = fact.valid_from    # invalidate, don't delete
        mem.add_fact(fact)
    mem.refresh_summary()

Failure modes

Summary poisoning. One wrong fact in the summary shapes every turn after it. Keep the raw episodes and check against them when it matters.

Leaking between users. Scope every read and write by user and tenant, and make deletion real across every tier.

Never forgetting. Stale facts crowd out useful ones. Decay by recency, expire what hasn't been touched, merge duplicates during consolidation.

Breaking the prompt cache. Every time you clear old turns, cached prompts are invalidated. Clear less often, in bigger batches.

Trusting published scores. Memory vendors have publicly disputed each other's numbers on the same benchmark. Run your own evaluation, full-context baseline included. LongMemEval is a good one to start with: it tests extraction from long histories, reasoning across sessions and over time, updated facts, and knowing when to say "I don't know".

Sources

  • Packer et al., MemGPT: Towards LLMs as Operating Systems (2023); Letta docs on memory and sleep-time agents
  • Chhikara et al., Mem0 (2025), and Mem0's 2026 algorithm update
  • Rasmussen et al., Zep: A Temporal Knowledge Graph Architecture for Agent Memory (2025)
  • Park et al., Generative Agents (2023)
  • Xu et al., A-MEM: Agentic Memory for LLM Agents (2025)
  • Wu et al., LongMemEval (ICLR 2025); Maharana et al., LoCoMo (2024)
  • Anthropic, Effective context engineering for AI agents (2025); OpenAI Agents SDK session docs
  • Hsieh et al., RULER (NVIDIA, 2024); Modarressi et al., NoLiMa (2025); Chroma, Context Rot (2025)
  • Rehberger, SpAIware — persistent memory injection in ChatGPT (2024)
  • Anthropic memory tool documentation; Timescale, pgvector vs Pinecone (vendor benchmark)

Repos to explore

Open-source projects worth reading alongside this. Each belongs to its authors; check the licence before using it.

Contact

Working on something
like this?

I'm open to AI engineering, architecture and training work. Tell me what you're building and what the constraints are — that's usually enough to start.

Prefer a short form? Send a project brief
  • Taking on new projects
  • Usually replies within a day