Introduction

In the era of Generative AI, data privacy is not just a compliance checkbox; it is a fundamental architectural requirement. Enterprises dealing with healthcare, finance, or customer support often possess vast amounts of Personally Identifiable Information (PII) such as names, social security numbers, email addresses, and phone numbers. Feeding this raw data directly into Large Language Models (LLMs)—whether cloud-based or on-premise poses significant risks of data leakage, regulatory fines (under GDPR, HIPAA, or CCPA), and reputational damage. PII Anonymization in LLM systems involves detecting sensitive entities in user inputs and retrieved context, replacing them with placeholders or pseudonyms before processing, and then optionally re-identifying them in the final output if necessary for the user experience. However, simple regex-based masking is often insufficient for complex natural language. A robust solution requires semantic understanding to distinguish between a "credit card number" and a random string of digits, or a "person's name" and a common noun.

This article details how to build an enterprise-grade PII anonymization pipeline using LangGraph for multi-agent orchestration, Microsoft Presidio for high-accuracy entity recognition, and Graph RAG for context-aware retrieval. We will implement a secure Customer Support Agent that ensures no PII ever leaves the trusted environment in plain text during LLM processing.

Real-Time Use Case: Secure Healthcare Patient Portal

Consider a healthcare provider implementing an AI assistant to help patients understand their lab results. A patient might ask: "Hi, I'm John Doe, DOB 1985-07-08. Can you explain my recent blood test results from Dr. Smith?"

If this query is sent to an LLM without protection, the model processes John’s name and date of birth. If the model is hosted externally or if logs are retained, this constitutes a PHI (Protected Health Information) breach. Our system must:

  1. Detect "John Doe" and "1985-07-08" as PII/PHI.

  2. Replace them with tokens like [PATIENT_NAME] and [DATE_OF_BIRTH].

  3. Retrieve relevant medical records using Graph RAG without exposing other patients' data.

  4. Generate a response based on the anonymized context.

  5. Return the answer to the user while maintaining an audit trail of what was masked.

Architecture Overview

  • Anonymization Agent: Uses Microsoft Presidio to scan and mask input text.

  • Graph RAG Engine: Retrieves medical context using Neo4j, ensuring row-level security.

  • LLM Processor: Processes the sanitized query and context.

  • Re-identification Agent (Optional): Maps placeholders back to original values for the final UI display if required by business logic.

  • LangGraph State: Tracks the original input, anonymized input, and the mapping dictionary for audit purposes.

Step-by-Step Implementation

Step 1: Environment Setup

# requirements.txt
langgraph==0.2.0
langchain-community==0.3.0
neo4j==5.14.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.6.0
openai==1.12.0
presidio-analyzer==2.2.0
presidio-anonymizer==2.2.0

Step 2: Defining the Privacy-Aware State

We maintain a mapping of original PII to their anonymized tokens to ensure we can track what was hidden.

from typing import TypedDict, List, Optional, Dict
from pydantic import BaseModel, Field

class PrivacyState(TypedDict):
    original_input: str
    anonymized_input: str
    pii_mapping: Dict[str, str]  # Maps token -> original value
    retrieved_context: Optional[dict]
    llm_response: str
    final_output: str
    audit_log: List[str]

class PatientQuery(BaseModel):
    message: str = Field(..., description="Patient's question")
    patient_id: str = Field(..., description="Verified patient ID")

Step 3: The Anonymization Agent

We use Microsoft Presidio, which uses NLP models to identify PII with high accuracy.

from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
from presidio_anonymizer.entities import OperatorConfig

analyzer = AnalyzerEngine()
anonymizer = AnonymizerEngine()

def anonymize_input(state: PrivacyState) -> PrivacyState:
    """Detects and replaces PII in the user's input"""
    # Analyze text for PII
    analyzer_results = analyzer.analyze(text=state['original_input'], language='en')
    
    # Anonymize detected entities
    anonymized_result = anonymizer.anonymize(
        text=state['original_input'],
        analyzer_results=analyzer_results,
        operators={"DEFAULT": OperatorConfig("replace", {"new_value": "[REDACTED]"})}
    )
    
    state['anonymized_input'] = anonymized_result.text
    state['audit_log'].append(f"Anonymized {len(analyzer_results)} PII entities")
    
    # Store mapping for audit/re-identification if needed
    for res in analyzer_results:
        original_text = state['original_input'][res.start:res.end]
        state['pii_mapping'][f"[{res.entity_type}]"] = original_text
        
    return state

Step 4: Secure Graph RAG Retrieval

We retrieve medical context using Neo4j. Since the input is already anonymized, we rely on the patient_id from the authenticated session for data scoping.

from neo4j import GraphDatabase

class MedicalGraphRetriever:
    def __init__(self):
        self.driver = GraphDatabase.driver("bolt://localhost:7687", auth=("neo4j", "password"))

    def get_patient_records(self, patient_id: str) -> dict:
        with self.driver.session() as session:
            query = """
            MATCH (p:Patient {id: $pid})-[:HAS_RECORD]->(r:MedicalRecord)
            RETURN r.test_name, r.result, r.date
            """
            result = session.run(query, pid=patient_id)
            return [record.data() for record in result]

def retrieve_secure_context(state: PrivacyState) -> PrivacyState:
    retriever = MedicalGraphRetriever()
    # In production, patient_id comes from auth token, not user input
    state['retrieved_context'] = retriever.get_patient_records("PATIENT_123")
    state['audit_log'].append("Retrieved medical records securely")
    return state

Step 5: LLM Processing and Final Output Generation

The LLM only sees the anonymized input and the structured context. It never sees the actual PII.

from langchain_openai import ChatOpenAI

llm = ChatOpenAI(model="gpt-4-turbo", temperature=0)

def generate_safe_response(state: PrivacyState) -> PrivacyState:
    context_str = str(state['retrieved_context'])
    prompt = f"""
    You are a helpful medical assistant. 
    Context: {context_str}
    User Question (Anonymized): {state['anonymized_input']}
    
    Provide a clear explanation. Do not attempt to guess any redacted information.
    """
    
    state['llm_response'] = llm.invoke(prompt).content
    # For this use case, we return the response as-is since the user knows their own name.
    # In a logging scenario, we would keep it anonymized.
    state['final_output'] = state['llm_response']
    state['audit_log'].append("Generated response using anonymized context")
    return state

Step 6: Orchestrating with LangGraph

from langgraph.graph import StateGraph, END

workflow = StateGraph(PrivacyState)
workflow.add_node("anonymize", anonymize_input)
workflow.add_node("retrieve", retrieve_secure_context)
workflow.add_node("respond", generate_safe_response)

workflow.set_entry_point("anonymize")
workflow.add_edge("anonymize", "retrieve")
workflow.add_edge("retrieve", "respond")
workflow.add_edge("respond", END)

app = workflow.compile()

Step 7: FastAPI Backend and React Frontend

# Backend
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware

api_app = FastAPI(title="Secure Medical AI")
api_app.add_middleware(CORSMiddleware, allow_origins=["*"], allow_methods=["*"])

@api_app.post("/ask-medical")
async def ask_medical(query: PatientQuery):
    initial_state = PrivacyState(
        original_input=query.message,
        anonymized_input="",
        pii_mapping={},
        retrieved_context=None,
        llm_response="",
        final_output="",
        audit_log=[]
    )
    result = await app.ainvoke(initial_state)
    return {
        "response": result['final_output'],
        "audit_summary": result['audit_log']
    }
// Frontend
import React, { useState } from 'react';
import axios from 'axios';

const MedicalChat = () => {
  const [msg, setMsg] = useState('');
  const [res, setRes] = useState(null);

  const send = async () => {
    const data = await axios.post('http://localhost:8000/ask-medical', {
      message: msg,
      patient_id: 'PATIENT_123'
    });
    setRes(data.data);
  };

  return (
    <div className="p-4 max-w-md mx-auto">
      <h2 className="text-xl font-bold">Secure Patient Portal</h2>
      <input 
        className="border p-2 w-full mt-2" 
        value={msg} 
        onChange={e => setMsg(e.target.value)} 
        placeholder="Ask about your results..."
      />
      <button onClick={send} className="bg-blue-600 text-white p-2 mt-2 rounded">Send</button>
      {res && (
        <div className="mt-4 p-3 bg-gray-100 rounded">
          <p>{res.response}</p>
          <p className="text-xs text-gray-500 mt-2">Audit: {res.audit_summary.join(', ')}</p>
        </div>
      )}
    </div>
  );
};
export default MedicalChat;

Conclusion

Implementing PII anonymization in LLM systems is a critical step toward trustworthy enterprise AI. By integrating tools like Microsoft Presidio into a LangGraph workflow, we create a pipeline where privacy is enforced by design, not by policy alone. This multi-agent approach ensures that sensitive data is masked before it reaches the LLM, reducing the risk of leakage while still allowing the model to provide valuable, context-aware responses. As regulations tighten, such architectures will become the standard for any organization leveraging generative AI in regulated industries.