AI Agent Long-Term Memory Architecture & MemGPT Systems

By AI Agent Hub Editorial Desk · Review method · Corrections

Architecture · 5 min read · Reviewed September 18, 2026

Without a persistent memory subsystem, an AI agent treats every conversation session in isolation—forgetting user preferences, architectural decisions, and past debugging attempts. A useful memory design separates working context, retrievable history, and durable structured facts. MemGPT, now continued through Letta, is an operating-system-inspired approach to managing limited model context and external memory; it is not itself a knowledge graph.

1. The 3-Tier Memory Hierarchy

Memory Tier Storage Medium Retention Duration Primary Purpose
Working Memory Prompt Window Buffer Current Active Session Immediate reasoning, scratchpad thoughts, tool outputs
Episodic Memory Vector DB (Chroma / Pinecone) Cross-Session (Days/Months) Recalling past user conversations, error logs, and code fixes
Epistemic Memory Versioned records or a knowledge graph Permanent System Knowledge Project architecture rules, user personas, API key policies

2. MemGPT: OS-Style Memory Management

MemGPT models LLM memory like OS virtual memory—using explicit tool calls (core_memory_append, archival_memory_search) to page information in and out of the LLM context window:

# Python MemGPT Function Calling Schema Example
def core_memory_append(section: str, content: str):
    """Appends critical user facts to the agent's core memory block."""
    # Append to system prompt core memory block
    pass

def archival_memory_search(query: str, page: int = 0):
    """Searches long-term vector database for past interaction logs."""
    pass

3. Start with a Memory Contract, Not a Vector Database

Before selecting storage, define what the agent is allowed to remember. A production memory record should carry more than text: include the tenant and subject, source event, creation time, expiry policy, sensitivity label, confidence, and the evidence that supports the statement. A preference such as “use concise answers” may be safe to retain; an inferred medical condition, credential, or private document excerpt usually is not. The model may propose a memory, but application code should validate and authorize the write.

FieldWhy it mattersExample control
ScopePrevents one user or project from seeing another's memoryTenant and workspace IDs enforced in the query
ProvenanceLets a reviewer trace a remembered claimMessage, ticket, file revision, or tool-result ID
ValidityDistinguishes current facts from stale observationsvalid_from, expires_at, and superseded-by fields
SensitivityControls storage, retrieval, and loggingPublic, internal, confidential, or prohibited
User controlMakes correction and deletion possibleReadable memory view plus edit and forget actions

4. A Safe Write and Retrieval Lifecycle

  1. Observe: capture the candidate fact and its source without treating it as true merely because the model stated it.
  2. Classify: reject secrets and prohibited personal data; select scope, retention, and required approval.
  3. Normalize: store a concise claim separately from the supporting event, preserving both identifiers.
  4. Deduplicate: merge an equivalent fact or create a new version instead of accumulating contradictory copies.
  5. Retrieve: filter by authorization and validity before semantic ranking. Similarity search must never be the permission layer.
  6. Use visibly: tell the model which records are memories, include provenance, and allow it to abstain when records conflict.
  7. Correct or delete: propagate user changes to indexes, caches, summaries, and future context construction.

5. Evaluation Plan for Long-Term Memory

Test memory as a subsystem rather than judging a few pleasant conversations. Build fixtures for correct recall, irrelevant recall, conflicting updates, expired facts, cross-tenant access, deletion, and malicious instructions stored inside old content. Report recall precision at the actual context budget, unsupported-memory rate, stale-memory rate, deletion propagation time, added tokens, and latency. A useful memory system should also know when not to remember: measure how often prohibited or low-value candidates are correctly rejected.

Run a temporal split so the test set contains updates that occur after earlier facts. This reveals whether the system prefers a newer authoritative record over an older semantically similar one. For high-impact actions, memory can provide context but must not replace current authorization or a fresh read from the system of record.

6. Primary References

7. Related Architecture & Builder Tools

The cost of carrying context forward

A memory tier only pays for itself if retrieving history is cheaper than re-sending it, and that is a question with a numeric answer. Below: one memory-augmented turn at 30K input tokens, 24K of them served from the cache that a stable prefix produces, and 2K output. Of the 30,000 input tokens, 24,000 are billed at the cache-read rate and 6,000 at full input rate.

Model Cost per turn Monthly at 5,000 turns
DeepSeek V4.1 Flash$0.0022$10.86
GPT-5.6 Luna$0.0041$20.40
Gemini 3.8 Flash$0.014$69.00
Claude Sonnet 5$0.037$184

This is where the memory design earns its keep. Because the prefix is stable, most of the input bills at the cache rate rather than the input rate, and on Sonnet 5 that is a tenfold difference. An architecture that reshuffles the prefix on every turn throws that away.

Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.