AI Agent Long-Term Memory Architecture & MemGPT Systems
Without a persistent memory subsystem, an AI agent treats every conversation session in isolation—forgetting user preferences, architectural decisions, and past debugging attempts. A useful memory design separates working context, retrievable history, and durable structured facts. MemGPT, now continued through Letta, is an operating-system-inspired approach to managing limited model context and external memory; it is not itself a knowledge graph.
1. The 3-Tier Memory Hierarchy
| Memory Tier | Storage Medium | Retention Duration | Primary Purpose |
|---|---|---|---|
| Working Memory | Prompt Window Buffer | Current Active Session | Immediate reasoning, scratchpad thoughts, tool outputs |
| Episodic Memory | Vector DB (Chroma / Pinecone) | Cross-Session (Days/Months) | Recalling past user conversations, error logs, and code fixes |
| Epistemic Memory | Versioned records or a knowledge graph | Permanent System Knowledge | Project architecture rules, user personas, API key policies |
2. MemGPT: OS-Style Memory Management
MemGPT models LLM memory like OS virtual memory—using explicit tool calls (core_memory_append, archival_memory_search) to page information in and out of the LLM context window:
# Python MemGPT Function Calling Schema Example
def core_memory_append(section: str, content: str):
"""Appends critical user facts to the agent's core memory block."""
# Append to system prompt core memory block
pass
def archival_memory_search(query: str, page: int = 0):
"""Searches long-term vector database for past interaction logs."""
pass
3. Start with a Memory Contract, Not a Vector Database
Before selecting storage, define what the agent is allowed to remember. A production memory record should carry more than text: include the tenant and subject, source event, creation time, expiry policy, sensitivity label, confidence, and the evidence that supports the statement. A preference such as “use concise answers” may be safe to retain; an inferred medical condition, credential, or private document excerpt usually is not. The model may propose a memory, but application code should validate and authorize the write.
| Field | Why it matters | Example control |
|---|---|---|
| Scope | Prevents one user or project from seeing another's memory | Tenant and workspace IDs enforced in the query |
| Provenance | Lets a reviewer trace a remembered claim | Message, ticket, file revision, or tool-result ID |
| Validity | Distinguishes current facts from stale observations | valid_from, expires_at, and superseded-by fields |
| Sensitivity | Controls storage, retrieval, and logging | Public, internal, confidential, or prohibited |
| User control | Makes correction and deletion possible | Readable memory view plus edit and forget actions |
4. A Safe Write and Retrieval Lifecycle
- Observe: capture the candidate fact and its source without treating it as true merely because the model stated it.
- Classify: reject secrets and prohibited personal data; select scope, retention, and required approval.
- Normalize: store a concise claim separately from the supporting event, preserving both identifiers.
- Deduplicate: merge an equivalent fact or create a new version instead of accumulating contradictory copies.
- Retrieve: filter by authorization and validity before semantic ranking. Similarity search must never be the permission layer.
- Use visibly: tell the model which records are memories, include provenance, and allow it to abstain when records conflict.
- Correct or delete: propagate user changes to indexes, caches, summaries, and future context construction.
5. Evaluation Plan for Long-Term Memory
Test memory as a subsystem rather than judging a few pleasant conversations. Build fixtures for correct recall, irrelevant recall, conflicting updates, expired facts, cross-tenant access, deletion, and malicious instructions stored inside old content. Report recall precision at the actual context budget, unsupported-memory rate, stale-memory rate, deletion propagation time, added tokens, and latency. A useful memory system should also know when not to remember: measure how often prohibited or low-value candidates are correctly rejected.
Run a temporal split so the test set contains updates that occur after earlier facts. This reveals whether the system prefers a newer authoritative record over an older semantically similar one. For high-impact actions, memory can provide context but must not replace current authorization or a fresh read from the system of record.
6. Primary References
7. Related Architecture & Builder Tools
- 🛠️ Interactive AI Agent Builder Lab & Code Studio
- 🔄 Agentic RAG Architecture & Self-Correction Pipelines
The cost of carrying context forward
A memory tier only pays for itself if retrieving history is cheaper than re-sending it, and that is a question with a numeric answer. Below: one memory-augmented turn at 30K input tokens, 24K of them served from the cache that a stable prefix produces, and 2K output. Of the 30,000 input tokens, 24,000 are billed at the cache-read rate and 6,000 at full input rate.
| Model | Cost per turn | Monthly at 5,000 turns |
|---|---|---|
| DeepSeek V4.1 Flash | $0.0022 | $10.86 |
| GPT-5.6 Luna | $0.0041 | $20.40 |
| Gemini 3.8 Flash | $0.014 | $69.00 |
| Claude Sonnet 5 | $0.037 | $184 |
This is where the memory design earns its keep. Because the prefix is stable, most of the input bills at the cache rate rather than the input rate, and on Sonnet 5 that is a tenfold difference. An architecture that reshuffles the prefix on every turn throws that away.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.