Agent Memory — Short-Term and Long-Term Memory for AI Agents
An AI agent that forgets what it just did after every step isn’t an agent — it’s a chatbot with an extra hop. Agent memory is the umbrella term for the mechanisms that let an agent keep information available beyond the current step: what currently sits in the context window (short-term memory) and what lives in an external store that gets pulled back in on demand (long-term memory). This article draws a clean line between the two layers, walks through the architecture patterns common in 2026, and points out where memory systems break in practice.
Short-term memory: the context window as working memory
A language model’s context window is inherently volatile — it exists only for the duration of a single request and is gone afterward. For an agent running several steps in a row (call a tool, read the result, plan the next step), that means the history so far has to be resent with every new call, or the model “forgets” what it’s doing mid-task.
That’s the simplest form of memory — the agent loop just appends observations, tool results and intermediate steps to a growing prompt. It works fine for short tasks, but hits two limits fast:
- Size. Even a large window is finite. Long agent runs with many tool calls (file dumps, search results, logs) fill it up surprisingly quickly.
- Quality. A full window suffers from the lost-in-the-middle effect — information in the middle of a long history is used worse than information near the edges.
The practical answer to this is context editing: older tool results are automatically stripped from the active context once a threshold is crossed — provided the important parts were written to an external store first. That’s exactly where long-term memory begins.
Long-term memory: external memory stores and vector DBs
Once information needs to survive beyond a single request — across many steps, past the end of a session, or across multiple sessions — the context window stops being enough. The agent needs an external store it can pull the relevant bits from on demand, instead of dragging everything along at all times.
Three architectures are common in 2026:
- Vector-store-based. Past interactions, extracted facts or user preferences get stored as embeddings in a vector database and retrieved via semantic search — essentially RAG, except the knowledge base is the agent’s own interaction history rather than a document collection.
- File-based. The agent writes and reads structured notes as files — this is how Anthropic’s memory tool for the Claude Developer Platform works: Claude writes findings, intermediate state and summaries to local files before the corresponding tool output is cleared from context, then reads them back on demand later.
- Graph-based. Instead of flat vector hits, a time-annotated knowledge graph is built (entities, relationships, validity periods), which answers questions like “what was true at time X” more robustly than plain similarity search — the approach behind tools like Zep’s Graphiti engine.
A useful distinction that has become standard practice by 2026 is four memory types: working memory (the current context), episodic memory (specific past events — “on Monday the user said X”), semantic memory (extracted, generalized facts and preferences — “the user prefers short answers”), and procedural memory (the agent’s own learned procedures). A production memory system typically combines several of these layers rather than relying on just one.
How agents keep context across steps and sessions
Technically, holding state happens on two levels:
- Within a single run (step to step). Agent frameworks like LangGraph thread an explicit state object through the graph — each node reads and writes a shared state, and a checkpointer persists it to a database after every step. If the process crashes, the run can resume from the last checkpoint instead of starting over.
- Across sessions (day to day, user to user). Here a session or thread ID ties the agent to its external memory store. When the same user or the same job comes back, the agent loads exactly the memories tied to that ID — not the full history, but a curated, usually summarized selection.
The consolidation step in between matters: a raw chat transcript is rarely stored 1:1 as long-term memory. Instead, a separate step (often a second, cheaper LLM call) extracts the relevant facts and only those get persisted — raw data otherwise grows without bound and becomes imprecise to retrieve.
In practice: frameworks and tools
A few concrete implementations that make the concepts tangible:
- Anthropic’s memory tool + context editing (Claude Developer Platform, in beta since 2025): file-based memory combined with automatic cleanup of old tool results from the active context.
- Letta (formerly MemGPT): applies the operating-system idea of virtual memory to LLMs — the agent itself decides, via function calls, what moves from “RAM” (active context) to “disk” (archival storage) and when to page it back in.
- Mem0: a dedicated memory layer that extracts memories from interactions, stores them, and serves them back through an API — framework- and vector-DB-agnostic, built for personalization across many sessions.
- Zep: builds a temporal knowledge graph instead of flat vector hits, positioned as an answer to the problem of stale or contradictory memories.
Which tool fits depends on the use case: a support bot tracking user preferences often does fine with a simple vector store. An agent that works on a project for weeks and needs to avoid contradictory intermediate states benefits from a graph-based approach or an OS-style memory model.
Pitfalls in practice
- Stale memories. A fact that was true last month (“customer is on the Basic plan”) can be wrong today. Without timestamps or graph structure, contradictory memories sit side by side in storage without comment — the agent then acts on outdated information.
- Hallucinated memories. When extracting facts from conversations, the extracting model can invent things that were never actually said. These false memories look just as convincing on retrieval as real ones — a reason to spot-check extracted facts.
- Context pollution. Pulling old memories back in too aggressively bloats the context again and brings back the lost-in-the-middle effect that memory was supposed to avoid in the first place. Not everything stored belongs in every new prompt.
- Cost and latency. Every store and retrieval operation costs additional LLM and embedding calls. For very frequent, short interactions, the memory overhead can end up costing more compute than the actual task.
- Privacy. Long-term memory means personal data (preferences, conversation content, behavior patterns) sits in permanent storage, not just fleetingly in the context of one request. Deletion policies, retention limits and an opt-out aren’t optional extras — under GDPR they’re a requirement as soon as user data is involved.
FAQ
Is agent memory the same thing as RAG? Not quite. RAG typically retrieves knowledge from an external, mostly static document collection. Agent memory stores and retrieves the agent’s own interaction history and facts learned from it — the retrieval technique underneath often overlaps (embeddings, vector search), but the content is different.
Doesn’t a bigger context window solve memory problems on its own? No. A bigger window just moves the point at which external storage becomes necessary — it’s still finite, still costs more per request, and still suffers from lost-in-the-middle. For knowledge that needs to survive across sessions, an external store is needed regardless of window size.
What’s the difference between episodic and semantic memory? Episodic memory stores specific events with context (“on September 3rd the customer asked X”). Semantic memory stores the generalized, de-timestamped takeaway from that (“the customer prefers short answers”). Production systems typically use both — episodic for traceability, semantic for fast, compact retrieval.
How do I stop an agent from accumulating contradictory memories? The most robust approach is a time-annotated store (graph or per-entry timestamps) that explicitly checks new facts against existing ones and marks outdated entries as superseded instead of leaving them standing unremarked. Without that check, contradictions accumulate unnoticed.
Do I need to build agent memory myself, or are there ready-made solutions? For most use cases, it’s worth reaching for an existing memory layer (Mem0 or Zep, for example) or the persistence mechanism built into your agent framework (LangGraph’s checkpointer, for example) rather than building extraction, storage and retrieval from scratch. Building it yourself only pays off for very specific privacy or storage-structure requirements.
Discover more
AI Workflows by Keyword: How We Make Recurring Routines Enforceable
A typed keyword triggers a fixed AI routine — and every single step must be committed before the next one appears. Why that's the actual trick.
GlossaryEnsemble / Multi-Model Orchestration
Ensemble means combining several deliberately varied LLM runs or models whose findings complement each other. Multi-model orchestration drives these runs via orchestrators with sub-agents, so the union of results is larger than any single run.
EncyclopediaFunction Calling / Tool Use
How an LLM uses tools: define a tool as a schema, the model picks the function and arguments, the result returns to the chat — the basis of every agent.