Introduction
In the lifecycle of a Retrieval-Augmented Generation (RAG) system, data is never static. As businesses update their knowledge bases, the underlying embedding models also evolve. A model trained on 2023 data might represent "cloud computing" differently than a 2026 model optimized for AI-native infrastructure. When you swap an embedding model or significantly retrain it, the vector space shifts. If your RAG system doesn't account for this, you end up with "version mismatch"—where queries are embedded using Model V2 but searched against vectors created by Model V1. This article explores Embedding Versioning and provides an enterprise-grade POC using a Multi-Agent LangGraph architecture to manage these transitions seamlessly.
The Challenge: Why Embedding Versioning Matters
Embedding versioning is the practice of tagging vector records with the specific model and version used to generate them. Without it, enterprises face two critical risks:
Semantic Incompatibility: Cosine similarity is only valid within the same vector space. Comparing a V1 query vector against V2 document vectors yields mathematically meaningless results, leading to poor retrieval accuracy.
Operational Blindness: During a migration from one embedding provider (e.g., OpenAI
text-embedding-ada-002) to another (e.g.,text-embedding-3-large), there is often a transition period where both sets of vectors coexist. Without version tags, the retriever cannot distinguish which documents are "current" and which are "legacy."
Real-Time Use Case: Global Product Catalog & Search
Consider a multinational e-commerce platform with a catalog of 10 million products. In early 2026, the engineering team decides to upgrade their search relevance by switching from a generic embedding model to a domain-specific model fine-tuned on retail terminology.
During the migration, the database contains:
Legacy Vectors (v1): Generated by the old model for 8 million products.
New Vectors (v2): Generated by the new model for 2 million recently updated products.
If a user searches for "sustainable running shoes," a standard RAG system might retrieve V1 vectors that don't align with the V2 query vector's nuanced understanding of "sustainability." Our multi-agent system will identify the active embedding version, filter the search to only include compatible vectors, and if necessary, trigger a background agent to re-embed legacy items on the fly.
Architecture: Multi-Agent LangGraph with Version-Aware State
To handle this complexity, we use LangGraph. Its stateful nature allows us to track the "active embedding version" as part of the global state.
Supervisor Agent: Manages the workflow and maintains the
active_embedding_versionin the state.Version Router Agent: Checks the incoming query and determines which embedding model should be used based on current system configuration.
Retriever Agent: Queries the Vector DB (e.g., Qdrant) with a strict metadata filter for the correct
embedding_version.Re-Embedding Agent (Optional): If no results are found in the current version, this agent can trigger a background job to convert legacy data.
Generator Agent: Produces the final response.
Step-by-Step POC: Backend Implementation (FastAPI + LangGraph)
We will build a FastAPI backend that uses LangGraph to route queries based on embedding versions.
Prerequisites
pip install fastapi uvicorn langgraph langchain-openai langchain-core pydantic streamlit requests qdrant-client
Backend Code (main.py)
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from typing import TypedDict, Annotated, List
from langgraph.graph import StateGraph, END
from langgraph.checkpoint.memory import MemorySaver
from langchain_openai import ChatOpenAI, OpenAIEmbeddings
from langchain_core.messages import HumanMessage, AIMessage
import operator
app = FastAPI(title="Embedding Versioning RAG API")
# 1. Define the State
class AgentState(TypedDict):
messages: Annotated[List, operator.add]
query: str
embedding_version: str
retrieved_context: str
final_answer: str
# Initialize LLM and Memory
llm = ChatOpenAI(model="gpt-4o", temperature=0)
memory = MemorySaver()
# Mock Embedding Models
def get_embeddings_for_version(version: str):
if version == "v1":
return OpenAIEmbeddings(model="text-embedding-ada-002")
else:
return OpenAIEmbeddings(model="text-embedding-3-large")
# 2. Define Agent Nodes
def version_router_node(state: AgentState):
# In production, this would check a config service or database
# For this POC, we assume the system is migrating to v2
active_version = "v2"
return {"embedding_version": active_version}
def retriever_node(state: AgentState):
version = state["embedding_version"]
query = state["query"]
# Get the correct embedding model
embedder = get_embeddings_for_version(version)
query_vector = embedder.embed_query(query)
# Mock Qdrant Search with Metadata Filter
# In real code: client.search(collection_name="products", query_vector=query_vector,
# query_filter=Filter(must=[FieldCondition(key="embedding_version", match=MatchValue(value=version))]))
mock_context = f"[Retrieved from {version} vectors]: Latest sustainable running shoes made from recycled ocean plastics..."
return {"retrieved_context": mock_context}
def generator_node(state: AgentState):
prompt = f"Answer the query based on the context.\nQuery: {state['query']}\nContext: {state['retrieved_context']}"
response = llm.invoke(prompt)
return {"final_answer": response.content, "messages": [AIMessage(content=response.content)]}
# 3. Build the LangGraph
workflow = StateGraph(AgentState)
workflow.add_node("route_version", version_router_node)
workflow.add_node("retrieve", retriever_node)
workflow.add_node("generate", generator_node)
workflow.set_entry_point("route_version")
workflow.add_edge("route_version", "retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)
# Compile with Memory
graph = workflow.compile(checkpointer=memory)
# FastAPI Endpoints
class QueryRequest(BaseModel):
query: str
thread_id: str = "default_thread"
@app.post("/chat")
async def chat_endpoint(req: QueryRequest):
config = {"configurable": {"thread_id": req.thread_id}}
initial_state = {
"messages": [HumanMessage(content=req.query)],
"query": req.query,
"embedding_version": "",
"retrieved_context": "",
"final_answer": ""
}
try:
final_state = graph.invoke(initial_state, config)
return {
"answer": final_state["final_answer"],
"used_embedding_version": final_state["embedding_version"]
}
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
Step-by-Step POC: Frontend Implementation (Streamlit)
The frontend will display the answer and inform the user which embedding version was used to retrieve the data, ensuring transparency.
Frontend Code (app.py)
import streamlit as st
import requests
import uuid
st.set_page_config(page_title="Version-Aware Enterprise RAG", layout="centered")
st.title("🏢 Product Search RAG (Version-Aware)")
# Initialize session state for thread memory
if "thread_id" not in st.session_state:
st.session_state.thread_id = str(uuid.uuid4())
if "messages" not in st.session_state:
st.session_state.messages = []
# Display chat history
for msg in st.session_state.messages:
with st.chat_message(msg["role"]):
st.markdown(msg["content"])
# Chat input
if prompt := st.chat_input("Search for products..."):
st.session_state.messages.append({"role": "user", "content": prompt})
with st.chat_message("user"):
st.markdown(prompt)
with st.chat_message("assistant"):
with st.spinner("Routing to correct embedding version and retrieving..."):
# Call FastAPI Backend
response = requests.post(
"http://localhost:8000/chat",
json={"query": prompt, "thread_id": st.session_state.thread_id}
)
if response.status_code == 200:
data = response.json()
answer = data["answer"]
version = data["used_embedding_version"]
# Render UI
st.markdown(answer)
st.info(f"🔍 **Data Source:** Retrieved using Embedding Model Version: `{version}`")
st.session_state.messages.append({"role": "assistant", "content": answer})
else:
st.error("Error connecting to the RAG backend.")
Running the POC
Start the backend:
uvicorn main:app --reloadStart the frontend:
streamlit run app.pyTest Case: Ask "Show me eco-friendly footwear." The system will route to
v2, retrieve from thev2vector space, and display the version used in the UI.
Conclusion
Embedding versioning is not just a technical detail; it is a fundamental requirement for maintaining the integrity of enterprise RAG systems. As models improve and data evolves, the ability to isolate vector spaces by version prevents semantic incompatibility and ensures high-quality retrieval. By leveraging a Multi-Agent LangGraph Architecture, we can automate the routing logic, ensuring that every query is matched with the correct embedding context. This approach provides the scalability, transparency, and reliability needed for mission-critical enterprise applications.

Join the conversation! Your thoughts help the community grow.