AI agents often need to remember information across multiple interactions. A support agent may need to retain a customer's preferences, a coding agent may need to track decisions made during a long debugging session, and a research agent may need to reuse facts discovered in earlier investigations.
Keeping all this information inside a single conversation context is not a reliable long-term memory strategy. Context windows have limits, sessions can end, and repeatedly sending the entire interaction history to a language model increases token usage and latency.
A more scalable approach is to separate agent memory from the model's immediate context and store it in a dedicated data layer.
Google Cloud AlloyDB for PostgreSQL and Memorystore for Valkey can serve complementary roles in this architecture. AlloyDB can store durable, structured memory and support relational queries and vector search. Memorystore for Valkey can provide fast access to frequently used state, temporary context, and cached retrieval results.
The goal is not to store every conversation forever. It is to build a memory system that preserves useful information, retrieves the right facts, respects access controls, and avoids giving the agent stale or irrelevant context.
Why AI Agents Need a Separate Memory Layer
An AI agent typically operates through a cycle of observing information, reasoning about it, invoking tools, and producing a response or action. In a multi-step workflow, the agent may need to retain facts that are not available in the current prompt.
For example, consider a developer assistant that investigates a recurring production issue. During one session, it discovers that a particular service uses a nonstandard timeout and that a previous deployment introduced a regression.
If the next session begins without that information, the agent may repeat the same investigation. Persisting a concise record of the finding allows the agent to retrieve it when the problem appears again.
A useful memory system addresses several problems:
Persistence: Important information survives the end of a session.
Retrieval: The agent can find relevant information without loading the entire history.
Latency: Frequently accessed state can be retrieved quickly.
Consistency: Memory updates can be tracked and validated.
Isolation: One user's information must not leak into another user's context.
Freshness: Outdated or superseded information should not be treated as current truth.
These requirements are different from ordinary conversation history. An effective memory architecture needs explicit rules for what to retain, how to retrieve it, and when to update or delete it.
Divide Memory Between AlloyDB and Valkey
AlloyDB and Memorystore for Valkey are not interchangeable storage systems. Their value comes from using each for the workload it handles well.
<box border radius="lg" padding={3} gap={3}>
<box gap={1}>
<row align="center" gap={2}>
<icon name="database" size="xl" />
**AlloyDB: Durable memory**
</row>
Store persistent facts, conversation summaries, agent decisions, task history, metadata, and embeddings used for semantic retrieval.
<text color="secondary" size="sm">Best suited to durable records, relational queries, and persistent retrieval indexes.</text>
</box>
<divider color="subtle" />
<box gap={1}>
<row align="center" gap={2}>
<icon name="zap" size="xl" />
**Memorystore for Valkey: Fast memory access**
</row>
Cache recent context, frequently accessed facts, session state, retrieval results, and short-lived coordination data.
<text color="secondary" size="sm">Best suited to low-latency access and data that can be reconstructed or refreshed when necessary.</text>
</box>
<divider color="subtle" />
<box gap={1}>
<row align="center" gap={2}>
<icon name="brain-circuit" size="xl" />
**AI agent: Memory policy**
</row>
Decide what information to save, retrieve, summarize, update, or forget before constructing the model's context.
</box>
</box>This design keeps the authoritative memory in durable storage while using Valkey to reduce repeated retrieval work.
Not every agent needs both systems. A small application may work well with AlloyDB alone. A high-throughput agent platform with frequent repeated reads may benefit from a cache, but only if measurements justify the added complexity.
Design a Durable Memory Schema in AlloyDB
The first step is to define what a memory record represents.
Storing raw conversation messages is straightforward, but it does not necessarily create useful long-term memory. A better design can distinguish between episodic records, which describe what happened, and semantic records, which capture reusable facts.
For example:
An episodic memory might record that an agent investigated an error during a particular deployment.
A semantic memory might record that a service requires a specific connection timeout.
A task memory might store the current status of a multi-step workflow.
A preference memory might store a user's explicitly stated response preferences.
These categories can have different retention and retrieval rules.
A simple relational table can store durable memory records:
CREATE TABLE agent_memory (
memory_id UUID PRIMARY KEY,
tenant_id TEXT NOT NULL,
agent_id TEXT NOT NULL,
memory_type TEXT NOT NULL,
content TEXT NOT NULL,
source_reference TEXT,
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
expires_at TIMESTAMPTZ,
metadata JSONB NOT NULL DEFAULT '{}'::jsonb
);
CREATE INDEX idx_agent_memory_lookup
ON agent_memory (
tenant_id,
agent_id,
memory_type,
updated_at DESC
);
----------------This schema is illustrative and provides a foundation for durable memory. The tenant_id and agent_id fields help scope records, while memory_type supports retrieval based on the type of information needed.
The source_reference field can identify the conversation, task, or event that produced the memory. The metadata field can hold structured attributes such as confidence, source category, or a version identifier.
Do not use a source reference as a substitute for access control. The application must enforce tenant and user permissions when reading or modifying records.
For production use, also define policies for retention, deletion, schema evolution, and how conflicting memories are resolved.
Add Semantic Retrieval With Embeddings
Keyword queries work well when the agent knows the exact terms it needs to find. However, users frequently describe the same problem in different words.
For example, a stored memory might say:
The payment service experienced connection pool exhaustion after a deployment.
A later request might ask why payment requests are failing because database connections are unavailable.
The wording differs, but the concepts overlap. Semantic retrieval can help find the relevant memory even when the query does not repeat the original terms.
AlloyDB supports vector search through PostgreSQL-compatible capabilities, including the pgvector extension where supported by the instance configuration.
A memory table can be extended with an embedding column:
CREATE EXTENSION IF NOT EXISTS vector;
ALTER TABLE agent_memory
ADD COLUMN embedding vector(768);The vector dimension is an example. It must match the embedding model used to generate the stored vectors. Confirm the supported extension version and configuration for your AlloyDB environment before using this schema.
Generate embeddings from the memory content using the same compatible embedding model for both stored memories and retrieval queries. Store the resulting vectors alongside the original text and metadata.
A conceptual retrieval query looks like this:
SELECT
memory_id,
memory_type,
content,
updated_at
FROM agent_memory
WHERE tenant_id = $1
AND agent_id = $2
ORDER BY embedding <=> $3::vector
LIMIT 10;Here, $1 and $2 represent the tenant and agent identifiers, while $3 represents the query embedding.
The <=> operator represents cosine distance in the relevant pgvector configuration. Smaller distances indicate closer vectors under that metric.
This example illustrates the retrieval pattern, not a complete production query. The embedding column must be populated, the model and vector dimensions must be compatible, and the query must apply the application's authorization and freshness rules.
For large collections, evaluate the appropriate vector index and query plan. An exact scan may be adequate for a small memory table, while a larger workload may require approximate nearest-neighbor indexing and careful recall-versus-latency tuning.
Use Valkey to Cache Frequently Accessed Memory
Durable storage does not eliminate the need to optimize repeated reads.
Suppose an agent repeatedly retrieves the same set of service configuration facts during a troubleshooting session. Fetching and processing those records on every step creates unnecessary database work.
Valkey can cache the selected memory records or the result of a recent retrieval operation.
A key might follow a convention such as:
agent-memory:{tenant-id}:{agent-id}:{session-id}The value could contain a compact JSON representation of the memory entries selected for the current session.
The cache should have a defined expiration policy. When the underlying memory changes, the application should invalidate the relevant cache entry or ensure that stale data cannot be used beyond its permitted lifetime.
A simple access pattern is:
Generate a cache key from the authorized tenant, agent, and session identifiers.
Check Valkey for a cached memory result.
If the result is valid, use it.
Otherwise, retrieve the relevant records from AlloyDB.
Store the result in Valkey with an appropriate expiration.
Construct the agent's context from the retrieved records.
The cache key must include every identity dimension that affects access. A key based only on a session identifier or user-provided string can create isolation problems if identifiers collide across tenants.
Do not cache data indefinitely merely because it is expensive to retrieve. Memory may be corrected, revoked, or deleted, and the cache must respect those changes.
Implement a Basic Valkey Cache in Python
The following example demonstrates a cache-aside pattern using the Python redis client, which supports the Redis protocol used by Valkey.
It assumes the application has already authenticated the user and obtained trusted tenant, agent, and session identifiers. It also assumes a function is available to retrieve authorized memory from AlloyDB.
import json
import os
import redis
cache = redis.Redis.from_url(
os.environ["VALKEY_URL"],
decode_responses=True,
socket_connect_timeout=2,
socket_timeout=2,
)
def get_agent_memory(
tenant_id,
agent_id,
session_id,
load_from_alloydb,
):
cache_key = (
f"agent-memory:{tenant_id}:"
f"{agent_id}:{session_id}"
)
try:
cached = cache.get(cache_key)
if cached is not None:
return json.loads(cached)
except redis.RedisError:
# Cache failures should not block the durable read.
pass
memories = load_from_alloydb(
tenant_id=tenant_id,
agent_id=agent_id,
session_id=session_id,
)
try:
cache.set(
cache_key,
json.dumps(memories),
ex=300,
)
except redis.RedisError:
# The durable result remains usable.
pass
return memoriesThis example uses a five-minute expiration to illustrate the mechanism. The correct TTL depends on how frequently the underlying memory changes and how much staleness the application can tolerate.
The load_from_alloydb function is intentionally abstract because its implementation depends on the database driver, query structure, and authorization model.
A production implementation should also validate the serialized payload, enforce size limits, handle malformed cache values, and define explicit invalidation behavior after memory updates or deletion.
Build the Agent's Memory Retrieval Pipeline
Storing information is only half the problem. The application also needs to decide which memories belong in the model's current context.
A useful retrieval pipeline can combine several steps.
Step 1: Extract the current task
Identify what the agent is trying to accomplish. A troubleshooting task might need recent incidents, service configuration, and deployment history. A customer-support task might need account preferences and relevant previous interactions.
Avoid retrieving every memory associated with the user or agent. Broad retrieval increases latency and introduces irrelevant context.
Step 2: Apply access and relevance filters
Filter records by tenant, agent, task scope, and any relevant authorization constraints before using them.
Then retrieve candidates through keyword search, relational filters, vector similarity, or a combination of these techniques.
Step 3: Rank and deduplicate the results
The first retrieval pass may return similar or overlapping memories. Rank candidates based on relevance, recency, source reliability, and the application's memory policy.
Do not assume the closest vector is always the most useful memory. A semantically similar record may be outdated or based on an unverified hypothesis.
Step 4: Apply freshness and confidence rules
A memory that was correct during an earlier incident may no longer describe the current environment.
Store useful metadata such as the source, observation time, confidence, and whether the information has been verified. Use this information to determine whether a memory should be included, refreshed, or treated as historical context.
Step 5: Build a compact context
Pass only the selected information to the model. Prefer concise records with clear provenance over long, unfiltered conversation histories.
This reduces unnecessary token consumption and helps the agent distinguish verified facts from prior observations.
Keep Memory Consistent Across AlloyDB and Valkey
Using a durable database and a cache introduces a consistency problem: the same memory can exist in two places.
Suppose the application updates a memory record in AlloyDB but the old version remains in Valkey. The next agent request may retrieve stale information even though the database contains the correction.
A simple approach is to invalidate the relevant cache key after a successful database update.
However, the application should not assume that a database update and a cache invalidation are one atomic transaction. A process can fail between those operations.
For workloads where stale memory could lead to a harmful decision, consider stronger mechanisms:
Store a version number or update timestamp with each memory.
Include that version in the cache payload.
Invalidate cache entries after successful durable updates.
Use a durable event or outbox pattern when cache invalidation must survive process failures.
Define a maximum staleness window for cached results.
Recheck critical facts against AlloyDB before performing sensitive actions.
For many applications, a short cache TTL combined with explicit invalidation is sufficient. More demanding workflows may need a version-aware cache or a durable change-processing mechanism.
Choose the consistency model according to the consequence of using stale information, not simply the desired cache hit rate.
Security and Privacy Considerations
Long-term agent memory can contain customer details, internal system information, access patterns, and other sensitive data. Persisting such information introduces responsibilities that do not disappear when the data is summarized or embedded.
Start with data minimization. Save information only when it has a clear future use, and avoid retaining secrets or unnecessary personal details.
Enforce tenant isolation at the database query layer. Do not retrieve a broad set of memories and rely on the language model to ignore records belonging to another user.
Use appropriate authentication and network controls for both AlloyDB and Valkey. Restrict database permissions, manage credentials through approved secret-management mechanisms, and encrypt data according to the requirements of the deployment.
Remember that embeddings are derived from source content. They should receive access controls and retention treatment appropriate to the information they represent.
Finally, provide a deletion workflow that removes or invalidates a memory across the durable database, vector indexes, caches, and any downstream copies. A deletion that affects only AlloyDB may leave the same information available through a cache.
Monitor the Memory System
A memory architecture should be evaluated using both infrastructure metrics and retrieval quality.
Useful infrastructure metrics include:
Metric | What it reveals |
|---|---|
Cache hit rate | How often Valkey avoids a durable database read |
Retrieval latency | Time required to find relevant memories |
Database query latency | Cost of retrieving authoritative records |
Memory count and growth | Whether storage usage is increasing unexpectedly |
Cache invalidation failures | Risk of serving outdated records |
Retrieval errors | Whether memory is unavailable during agent execution |
Also evaluate application-level outcomes. A high cache hit rate does not prove that the agent retrieves the right information.
Create representative tasks with known relevant memories and measure retrieval recall, irrelevant-result rates, and the frequency with which outdated facts are selected. Track whether retrieved memories actually improve task completion or reduce repeated work.
The best architecture is not necessarily the one with the most cached data. It is the one that supplies the right information with acceptable latency, cost, consistency, and privacy guarantees.
Summary
AlloyDB and Memorystore for Valkey can form a complementary memory layer for AI agents. AlloyDB provides durable storage, structured queries, and vector-based retrieval, while Valkey can accelerate access to frequently used memory and session-specific context.
The implementation starts with a well-defined memory schema, adds semantic retrieval where needed, and uses a cache-aside pattern for repeated reads. The agent then applies relevance, authorization, freshness, and confidence rules before inserting selected memories into its context.
The main engineering challenges are not simply storing data. They are deciding what deserves to be remembered, preventing stale information from being reused, enforcing tenant isolation, and keeping the cache consistent with the authoritative store.
Start with the simplest architecture that meets the workload's requirements. Add Valkey when measured retrieval patterns justify caching, and evaluate memory quality alongside latency and infrastructure costs. A reliable agent memory system should improve continuity without sacrificing correctness, security, or control over retained information.

Join the conversation! Your thoughts help the community grow.