Introduction

In the rapidly evolving landscape of Generative AI, Retrieval-Augmented Generation (RAG) systems are rarely static. Enterprises frequently update database schemas, modify API contracts, or restructure knowledge bases to accommodate new business requirements. However, when a RAG pipeline relies on rigid vector embeddings and hardcoded retrieval logic, schema evolution becomes a breaking point. Documents indexed under an old schema may return irrelevant context, hallucinations may increase due to mismatched metadata filters, and multi-agent orchestration can fail when state definitions change.

Handling schema evolution in RAG is not merely a database migration task; it is an architectural challenge involving semantic consistency, state management, and backward compatibility. In enterprise environments using multi-agent frameworks like LangGraph, this requires a system that understands both the old and new data structures during transition periods. This article demonstrates a production-grade Proof of Concept (PoC) for managing schema evolution in a multi-agent RAG system with persistent memory and state, utilizing a real-time HR policy migration use case.

Real-Time Use Case: HR Policy System Migration

Consider an enterprise transitioning its HR knowledge base from a legacy flat-document structure (doc_id, text, category) to a richly typed, compliance-aware schema (policy_id, effective_date, jurisdiction, applies_to_roles, superseded_by). During this three-month migration window, the RAG system must accurately answer employee queries by retrieving relevant policies regardless of whether they reside in the legacy table or the new compliant store. A naive re-indexing would cause downtime and loss of historical context. Our solution uses a LangGraph multi-agent architecture where a "Schema Router Agent" dynamically directs retrieval based on query intent and temporal context, while a "State Manager" maintains conversation continuity across schema boundaries.

Step-by-Step PoC Implementation

Technology Stack

  • Orchestration: LangGraph (Python)

  • Backend: FastAPI + Pydantic V2

  • Vector Store: PostgreSQL with pgvector

  • LLM: Azure OpenAI GPT-4o

  • Frontend: Streamlit

  • Memory: LangGraph Checkpointer (PostgreSQL)

Step 1: Define Dual-Schema Pydantic Models

We use Pydantic V2’s model_config and computed fields to create a unified interface that abstracts schema differences from the agents.

from pydantic import BaseModel, Field, computed_field
from typing import Literal, Optional
from datetime import date

class LegacyPolicy(BaseModel):
    doc_id: str
    text: str
    category: str

class ModernPolicy(BaseModel):
    policy_id: str
    content: str
    effective_date: date
    jurisdiction: Literal["US", "EU", "APAC"]
    applies_to_roles: list[str]
    superseded_by: Optional[str] = None

    @computed_field
    @property
    def is_active(self) -> bool:
        return self.superseded_by is None and self.effective_date <= date.today()

# Unified schema for agent consumption
class UnifiedPolicyContext(BaseModel):
    id: str
    content: str
    schema_version: Literal["legacy", "modern"]
    metadata: dict

Step 2: Build the LangGraph Multi-Agent State

Define a state that tracks which schema version was used for retrieval, enabling the synthesizer agent to format responses appropriately.

from typing import Annotated, TypedDict
from langgraph.graph.message import add_messages

class RAGState(TypedDict):
    messages: Annotated[list, add_messages]
    query: str
    retrieved_docs: list[UnifiedPolicyContext]
    active_schema_version: str
    migration_phase: Literal["pre", "during", "post"]

Step 3: Implement Schema-Aware Retrieval Node

This node performs hybrid search across both schemas and normalizes results.

async def schema_aware_retriever(state: RAGState):
    query = state["query"]
    # Pseudo-code: Query both tables/views
    legacy_results = await legacy_vector_store.similarity_search(query, k=3)
    modern_results = await modern_vector_store.similarity_search(
        query, 
        filter={"is_active": True}, 
        k=3
    )
    
    unified = []
    for doc in legacy_results:
        unified.append(UnifiedPolicyContext(
            id=doc.doc_id, content=doc.text, 
            schema_version="legacy", metadata={"category": doc.category}
        ))
    for doc in modern_results:
        unified.append(UnifiedPolicyContext(
            id=doc.policy_id, content=doc.content,
            schema_version="modern", 
            metadata={"jurisdiction": doc.jurisdiction, "roles": doc.applies_to_roles}
        ))
    
    # Rank by relevance + recency bias for modern schema
    ranked = sorted(unified, key=lambda x: (x.schema_version == "modern", -relevance_score(x)))
    return {"retrieved_docs": ranked[:5], "active_schema_version": "hybrid"}

Step 4: Synthesizer Agent with Schema Context

The LLM receives explicit instructions about schema versions to prevent conflating outdated and current policies.

SYNTHESIZER_PROMPT = """You are an HR policy assistant. Answer based ONLY on retrieved docs.
If docs include 'legacy' schema versions, explicitly note they may be superseded.
Always cite the schema version and effective date when available.
Retrieved Docs: {docs}
User Query: {query}"""

Step 5: FastAPI Backend with Memory Persistence

Use LangGraph’s PostgreSQL checkpointer to maintain thread state across schema transitions.

from fastapi import FastAPI, Depends
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver
from contextlib import asynccontextmanager

@asynccontextmanager
async def lifespan(app: FastAPI):
    checkpointer = AsyncPostgresSaver.from_conn_string(DATABASE_URL)
    app.state.graph = build_rag_graph().compile(checkpointer=checkpointer)
    yield

app = FastAPI(lifespan=lifespan)

@app.post("/chat")
async def chat(thread_id: str, message: str):
    config = {"configurable": {"thread_id": thread_id}}
    result = await app.state.graph.ainvoke(
        {"messages": [("user", message)], "query": message}, 
        config=config
    )
    return {"response": result["messages"][-1].content}

Step 6: Streamlit Frontend

A simple UI that displays schema version badges alongside citations, helping users understand data provenance during migration.

import streamlit as st
import requests

st.title("HR Policy Assistant (Schema-Evolution Aware)")
thread_id = st.session_state.get("thread_id", "default-thread")

if prompt := st.chat_input("Ask about HR policies..."):
    response = requests.post(f"{API_URL}/chat", json={"thread_id": thread_id, "message": prompt})
    st.chat_message("assistant").write(response.json()["response"])

Best Practices for Schema Evolution in RAG

  1. Versioned Embeddings: Maintain separate embedding spaces or metadata tags for each schema version. Never overwrite old embeddings until deprecation is confirmed.

  2. Semantic Aliasing: Create mapping layers that translate old field names to new ones at query time, not just at index time.

  3. Graceful Degradation: If the new schema returns zero results, automatically fall back to legacy search with a user-facing disclaimer.

  4. State Serialization Safety: When using LangGraph checkpointers, ensure state schemas are forward-compatible. Use optional fields and avoid removing keys abruptly.

  5. Automated Regression Testing: Maintain a golden dataset of Q&A pairs validated against both schemas. Run these nightly during migration windows.

  6. Observability: Log schema version hits per query. A sudden drop in legacy hits indicates successful migration; persistent legacy hits indicate incomplete data transformation.

Conclusion

Schema evolution in enterprise RAG pipelines demands more than database expertise—it requires treating data structure changes as first-class citizens in your AI architecture. By implementing dual-schema awareness in retrieval nodes, maintaining state continuity through LangGraph’s checkpointing, and exposing schema provenance to end users, organizations can navigate migrations without sacrificing RAG reliability. The multi-agent pattern excels here because it decouples schema routing logic from generation logic, allowing each concern to evolve independently. As AI systems become integral to enterprise operations, building resilience against structural change is no longer optional—it is foundational to trustworthy AI deployment.