Introduction
An AI agent can maintain a conversation for several turns, but that does not automatically mean it has useful memory.
A user may explain a project requirement today, return next week, and expect the agent to remember the important decisions. If the agent starts from an empty context every time, the user has to repeat information.
The solution is persistent agent memory.
Modern agent frameworks increasingly separate conversation state from long-term memory. Microsoft Agent Framework, for example, provides sessions for maintaining conversation context and context providers for injecting persistent information. Microsoft Foundry Agent Service also provides managed long-term memory designed to retain useful information across sessions, devices, and workflows.
But adding a database to an AI agent does not automatically create good memory.
A production memory system must answer several questions:
What information should be remembered?
How long should it be retained?
Who can access it?
How should old information be updated?
How should relevant memories be retrieved?
What happens when two memories conflict?
How can a user correct or remove information?
The goal is not to remember everything.
The goal is to remember the right things and retrieve them at the right time.
Session State vs Long-Term Memory
The first architectural decision is understanding the difference between session state and persistent memory.
Session State
Session state represents the current conversation.
For example:
User:
I am building an invoice processing application.
Agent:
What database are you using?
User:
PostgreSQL.
The agent needs this information during the current interaction.
Conceptually:
Conversation
↓
Session
↓
Current context
If the session ends, that context may no longer be available unless it is explicitly persisted.
Agent Framework uses AgentSession to preserve conversation context between invocations and supports serializing and restoring session state.
Long-Term Memory
Long-term memory contains information that should remain useful across sessions.
For example:
User prefers C# examples.
User's application uses PostgreSQL.
User wants concise technical explanations.
Project uses Azure.
The architecture becomes:
Session
↓
Current conversation
Long-term memory
↓
Persistent user/project context
These two systems should not be treated as interchangeable.
Why Saving the Entire Chat History Is Not Enough
A simple implementation might store every message:
{
"user": "I prefer PostgreSQL.",
"assistant": "Understood."
}
Then, on the next conversation, the application sends previous messages back to the model.
This works for small conversations, but it becomes expensive and inefficient as history grows.
Suppose a user has 500 previous conversations.
Sending everything to the model creates several problems:
larger prompts
higher token consumption
higher latency
irrelevant context
conflicting information
more difficult retrieval
increased privacy exposure
Instead of treating memory as an archive, treat it as a curated knowledge layer.
A Better Memory Architecture
A practical architecture looks like this:
User
|
v
Current Request
|
v
Memory Retrieval
|
+-----------+-----------+
| |
v v
Session Context Long-Term Memory
| |
+-----------+-----------+
|
v
Agent
|
v
Response
|
v
Memory Extraction
|
v
Memory Store
The important step is memory extraction.
Not every conversation message should become a permanent memory.
What Should an Agent Remember?
A useful memory system usually stores information that has future value.
Common categories include:
User Preferences
Examples:
Preferred language: C#
Preferred response format: concise
Preferred database: PostgreSQL
Stable User Context
Examples:
Works primarily with .NET applications.
Maintains an internal API platform.
Uses Azure for production workloads.
Project Context
Examples:
Project: Invoice Processing
Database: PostgreSQL
Deployment: Kubernetes
Authentication: Microsoft Entra ID
Decisions
Examples:
The team decided to use REST instead of GraphQL.
The project uses PostgreSQL JSONB for flexible metadata.
Recurring Procedures
Examples:
When deploying the service:
1. Run tests.
2. Build the container.
3. Push the image.
4. Deploy to staging.
5. Run smoke tests.
Microsoft Foundry's current memory model explicitly distinguishes user profile memory, chat-summary memory, and procedural memory.
This is a useful architectural distinction because each type should be retrieved differently.
Do Not Store Every Statement as a Fact
Consider this conversation:
User:
I am testing MongoDB today.
It would be dangerous to automatically create:
User's database = MongoDB
The statement may only describe today's experiment.
Compare that with:
User:
Our production application uses PostgreSQL and we plan to keep it for the foreseeable future.
That is much stronger evidence for durable memory.
Memory extraction therefore needs to consider:
permanence
confidence
relevance
source
scope
expiration
Add Memory Metadata
A memory record should contain more than just text.
For example:
{
"id": "mem_1024",
"scope": "user_123",
"type": "preference",
"content": "Prefers C# examples",
"confidence": 0.94,
"source": "conversation",
"createdAt": "2026-09-28T10:30:00Z",
"updatedAt": "2026-09-28T10:30:00Z",
"expiresAt": null
}
Useful metadata includes:
Field | Purpose |
|---|---|
ID | Identify the memory |
Scope | Define who can access it |
Type | Preference, fact, summary, procedure |
Content | Stored knowledge |
Confidence | Indicate extraction confidence |
Source | Explain where it came from |
Created At | Track creation |
Updated At | Track changes |
Expiration | Remove temporary information |
This becomes particularly important when the memory system grows.
Scope Memory Correctly
Memory must never accidentally cross user or tenant boundaries.
Imagine:
User A
↓
Memory Store
↓
User B
If retrieval does not enforce scope, User B could potentially receive information extracted from User A's conversations.
A safer architecture is:
Tenant A
├── User A1 memory
└── User A2 memory
Tenant B
├── User B1 memory
└── User B2 memory
Microsoft Foundry's memory APIs use a scope value to segment memory and support isolation across users or other identifiers.
The same principle applies whether memory is implemented using a managed service, SQL database, document database, or vector store.
Retrieval Is More Important Than Storage
An agent can have thousands of memories and still behave poorly if retrieval is weak.
Suppose the store contains:
Memory 1: User prefers C#.
Memory 2: User works with PostgreSQL.
Memory 3: User is building an invoice system.
Memory 4: User likes hiking.
Memory 5: User attended a conference in 2024.
The current request is:
How should I structure my PostgreSQL repository layer?
Retrieving memories about hiking or an old conference is unnecessary.
The retrieval system should prioritize:
PostgreSQL
+
Invoice application
+
Relevant development preferences
This is why semantic retrieval is commonly used for long-term agent memory.
Microsoft's current Agent Memory Toolkit for Azure Cosmos DB, for example, supports durable turns, summaries, facts, and user summaries and provides semantic retrieval capabilities.
Use Different Retrieval Strategies for Different Memory Types
Not every memory should be retrieved using the same algorithm.
User Profile Memory
Retrieve early in a conversation.
User preferences
Language
Stable constraints
Conversation Summary
Retrieve when the current conversation relates to a previous thread.
Previous project discussion
Previous decisions
Open tasks
Procedural Memory
Retrieve when the user asks for a workflow the agent has previously handled.
"How do we normally deploy this service?"
Microsoft Foundry's memory documentation uses different retrieval guidance for these memory categories.
This separation can reduce unnecessary context injection.
Retrieval Should Be Selective
A simple retrieval pipeline can look like:
Current request
↓
Generate search representation
↓
Search memory
↓
Filter by scope
↓
Filter by type
↓
Rank relevance
↓
Apply freshness/confidence rules
↓
Select top memories
↓
Inject into agent context
For example:
const memories = await memoryStore.search({
scope: userId,
query: userMessage,
limit: 8
});
const relevantMemories = memories
.filter(memory => memory.confidence >= 0.75)
.slice(0, 5);
The exact implementation depends on the memory provider, but the architectural principle remains the same:
Do not inject everything you retrieve.
Memory Consolidation Prevents Duplication
Over time, agents can store multiple memories about the same subject.
For example:
User prefers PostgreSQL.
User's project uses PostgreSQL.
The application database is PostgreSQL.
Production database: PostgreSQL.
These records may contain overlapping information.
A consolidation process can turn them into:
The user's primary application database is PostgreSQL.
Microsoft Foundry's current memory implementation includes extraction and consolidation behavior intended to merge overlapping information and resolve conflicts.
A custom implementation can follow a similar pattern:
New memory
↓
Find similar memories
↓
Compare information
↓
Merge / replace / reject
↓
Store consolidated memory
Handle Conflicting Memories
Conflicts are inevitable.
Consider:
Memory A:
User's preferred language is Python.
Memory B:
User's preferred language is C#.
The system should not blindly keep both and allow retrieval to select one randomly.
Instead, consider:
timestamp
source reliability
explicit user correction
confidence
scope
whether the new statement supersedes the old one
For example:
{
"content": "Preferred language is C#",
"supersedes": "mem_421",
"updatedAt": "2026-09-28T10:30:00Z"
}
The memory lifecycle should support updates rather than treating memories as immutable facts.
Add Expiration
Some information should not live forever.
Examples:
Current project deadline
Temporary deployment environment
Current travel plan
Temporary software version
Current incident status
A memory record can include:
{
"content": "Production deployment is scheduled for Friday",
"expiresAt": "2026-10-02T23:59:59Z"
}
Microsoft Foundry's current memory preview supports store-level default retention controls and TTL settings for memory entries.
Expiration is especially important for information that changes frequently.
Memory Is Not a Knowledge Base
Another important distinction is between memory and retrieval-augmented generation (RAG).
Memory usually represents information learned from interactions.
A knowledge base contains authoritative external information.
For example:
Memory:
User prefers PostgreSQL.
Knowledge base:
Company database standards require PostgreSQL 17.
The first is personalized context.
The second is organizational knowledge.
Microsoft's current guidance similarly distinguishes long-term agent memory from curated knowledge sources and file search.
Do not use user memory as a replacement for authoritative documentation.
Memory and RAG Can Work Together
A production agent may need both.
User Request
|
+-------------------+
| |
v v
User Memory Knowledge Base
| |
+---------+---------+
|
v
Agent Context
|
v
Response
For example:
Memory:
User prefers C#.
Knowledge:
Company API development standards.
Current request:
Create an API endpoint.
The agent can combine personal preferences with authoritative engineering standards.
Keep Memory Out of the Main Prompt When Possible
A common implementation is to concatenate every memory into the system prompt:
User memories:
1. ...
2. ...
3. ...
4. ...
5. ...
...
This becomes difficult to control as the memory store grows.
Instead, retrieve only relevant information and inject it into a clearly separated context section.
For example:
<user_memory>
The user prefers C# examples.
The current project uses PostgreSQL.
</user_memory>
<current_request>
Design the repository layer.
</current_request>
This separation makes the agent's context easier to inspect and debug.
Memory Writes Should Be Asynchronous
Memory extraction does not always need to block the user's response.
A practical architecture is:
User Request
↓
Agent
↓
Response
↓
Queue memory extraction
↓
Extract
↓
Consolidate
↓
Persist
This prevents memory processing from unnecessarily increasing response latency.
Microsoft Foundry's current memory implementation also uses delayed/debounced updates for long-term memory writes in some flows.
For production systems, asynchronous processing can also make memory failures independent from the main user request.
Add a Memory Audit Trail
When an agent remembers something incorrectly, developers need to know why.
Store information such as:
Memory ID
Created by
Source conversation
Extraction timestamp
Previous value
New value
Reason for update
Confidence
Then a debugging process can look like:
Wrong answer
↓
Retrieved memory
↓
Memory ID
↓
Original source
↓
Extraction decision
↓
Correction
Without an audit trail, debugging memory-related behavior becomes difficult.
Privacy and Security
Memory introduces another data-storage boundary into an AI application.
Before storing information, determine:
Is this information necessary?
Is it sensitive?
How long should it be retained?
Who can access it?
Can the user correct it?
Can the user request deletion?
Is the data encrypted?
Is access logged?
Does tenant isolation apply?
Memory should follow the same security principles as other persistent application data.
Do not automatically store sensitive conversation content simply because the model can extract it.
Give Users Control
Users should be able to understand and correct important memories.
Useful capabilities include:
"What do you remember about me?"
"Forget that preference."
"Update my project database to PostgreSQL."
"Don't remember this."
The underlying memory API should therefore support more than insertion.
A useful lifecycle is:
Create
Read
Update
Delete
Expire
Current Microsoft Foundry memory capabilities include item-level create, read, update, list, and delete operations along with retention controls.
Common Mistakes
Storing Everything
More memory does not necessarily mean better responses.
Irrelevant information increases retrieval noise.
Treating Temporary Statements as Permanent Facts
A statement about today's task should not automatically become a permanent user preference.
No Scope Isolation
Never use a shared memory namespace without strict tenant and user boundaries.
No Expiration
Changing information becomes misleading when old memories remain indefinitely.
Retrieving Too Much
Injecting dozens or hundreds of memories can increase context size while reducing relevance.
No Conflict Resolution
Contradictory memories should have an explicit resolution strategy.
Mixing Memory With Authoritative Knowledge
User-specific memory should not override company policies, product documentation, or other authoritative sources.
No User Controls
A user should have a way to correct important persistent information.
Best Practices
For production AI agents:
Separate session state from long-term memory.
Store durable information rather than every message.
Classify memories by type.
Attach scope to every memory.
Use relevance-based retrieval.
Apply confidence and freshness rules.
Consolidate duplicate memories.
Resolve conflicting facts explicitly.
Expire temporary information.
Keep memory separate from authoritative knowledge bases.
Process memory writes asynchronously where appropriate.
Maintain an audit trail.
Provide user correction and deletion controls.
Test memory retrieval independently from response generation.
A Production Memory Model
A practical database model might look like:
Memory
--------------------------------
id
scope_id
memory_type
content
embedding
confidence
source_id
created_at
updated_at
expires_at
supersedes_id
status
For larger systems, you may also maintain:
MemorySource
--------------------------------
id
conversation_id
message_id
created_at
MemoryAudit
--------------------------------
id
memory_id
operation
old_value
new_value
timestamp
actor
This makes the memory subsystem observable instead of turning it into an opaque collection of model-generated text.
Testing Agent Memory
Memory should be tested like any other production subsystem.
Create test cases for:
Recall
Session 1:
My project uses PostgreSQL.
Session 2:
What database does my project use?
Expected:
PostgreSQL
Non-Recall
Session 1:
I am testing MongoDB today.
Session 2:
What database does my production project use?
The temporary statement should not automatically become a permanent fact.
Conflict
Session 1:
I use Python.
Session 2:
I switched the project to C#.
The newer explicit decision should replace or supersede the previous project-specific memory.
Scope Isolation
User A → Project Alpha
User B → Project Beta
User B must never retrieve User A's project context.
Expiration
Memory expires
↓
Retrieval
↓
Memory excluded
These tests are just as important as testing the model's generated response.
Summary
AI agent memory is not simply a database containing previous conversations.
A useful memory architecture separates current session state, durable user or project context, and authoritative knowledge. Modern agent frameworks already provide different approaches for session persistence and long-term memory. Microsoft Agent Framework uses sessions and context providers, while Microsoft Foundry Agent Service provides managed long-term memory with extraction, consolidation, retrieval, scoping, and retention capabilities.
The most important design principle is:
Do not make the agent remember everything. Make it remember what will still be useful later.
A production memory system should therefore have:
Selective storage
+
Correct scope
+
Relevant retrieval
+
Conflict resolution
+
Expiration
+
User control
+
Observability
When these pieces work together, an agent can maintain useful continuity across sessions without turning its context into an ever-growing transcript.

Join the conversation! Your thoughts help the community grow.