Introduction

As enterprises integrate Generative AI into their core operations, the security landscape shifts from traditional perimeter defense to protecting intellectual property (IP) and sensitive data within the AI pipeline. In a Multi-Agent Retrieval-Augmented Generation (RAG) system, three critical assets are at risk: proprietary prompts (the "secret sauce" of agent behavior), model endpoints (expensive infrastructure vulnerable to abuse), and sensitive context (private user data or corporate secrets injected into the LLM). This article provides a comprehensive guide to securing these assets within an enterprise-grade FastAPI application. We will build a Proof-of-Concept (PoC) for a "Secure Legal Research Assistant," a multi-agent system that analyzes confidential legal documents. We will demonstrate how to encrypt prompt templates, enforce strict access controls on model endpoints, and implement dynamic data masking to protect sensitive context before it ever reaches the Large Language Model (LLM).

The Security Triad in GenAI

In a multi-agent system, security cannot be an afterthought.

  • Proprietary Prompts: These define the agent’s persona and logic. If leaked, competitors can replicate your service.

  • Model Endpoints: Direct access to LLM APIs can lead to massive cost inflation via "denial-of-wallet" attacks.

  • Sensitive Context: RAG systems inject retrieved documents into the prompt. If these contain Personally Identifiable Information (PII) or trade secrets, they must be protected from both leakage and unauthorized access.

Securing Proprietary Prompts

Storing prompts in plain text in code repositories is a critical vulnerability. Instead, we use HashiCorp Vault to store encrypted prompt templates. The application retrieves them at runtime using short-lived tokens.

import hvac
from functools import lru_cache

class PromptManager:
    def __init__(self):
        self.client = hvac.Client(url='http://vault:8200', token=os.getenv('VAULT_TOKEN'))
    
    @lru_cache(maxsize=128)
    def get_prompt_template(self, agent_name: str, version: str = "latest") -> str:
        # Retrieve encrypted prompt from Vault
        secret = self.client.secrets.kv.v2.read_secret_version(
            path=f'prompts/{agent_name}',
            mount_point='secret'
        )
        return secret['data']['data']['template']

prompt_mgr = PromptManager()

Protecting Model Endpoints

We never expose direct LLM API keys to the frontend or even the main application logic if possible. Instead, we use a Proxy Pattern with strict authentication and rate limiting.

from fastapi import Depends, HTTPException
from slowapi import Limiter
from slowapi.util import get_remote_address

limiter = Limiter(key_func=get_remote_address)

@app.post("/api/inference")
@limiter.limit("10/minute")
async def secure_inference(request: InferenceRequest, user: dict = Depends(get_current_user)):
    if user['role'] not in ['admin', 'analyst']:
        raise HTTPException(status_code=403, detail="Unauthorized access to model")
    
    # Internal proxy call to LLM provider using stored credentials
    response = await llm_proxy.invoke(request.prompt)
    return response

Safeguarding Sensitive Context: PII Redaction

Before any context is injected into a prompt, it must be scanned for sensitive data. We use Microsoft Presidio to detect and redact PII (Personally Identifiable Information) dynamically.

from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine

analyzer = AnalyzerEngine()
anonymizer = AnonymizerEngine()

def sanitize_context(text: str) -> str:
    results = analyzer.analyze(text=text, language='en')
    anonymized_text = anonymizer.anonymize(text=text, analyzer_results=results)
    return anonymized_text.text

Backend Implementation: Secure Multi-Agent RAG

Our LangGraph workflow integrates these security layers. The Research Agent retrieves documents, sanitizes them, and then uses a secure prompt from Vault to generate insights.

from langgraph.graph import StateGraph, END

class SecureAgentState(TypedDict):
    query: str
    raw_documents: list
    sanitized_documents: list
    final_response: str

def retrieve_and_sanitize_node(state: SecureAgentState):
    docs = vector_db.search(state['query'])
    # Sanitize each document before adding to context
    state['sanitized_documents'] = [sanitize_context(doc) for doc in docs]
    return state

def generate_secure_response_node(state: SecureAgentState):
    # Get secure prompt from Vault
    prompt_template = prompt_mgr.get_prompt_template("legal_analyst")
    context = "\n".join(state['sanitized_documents'])
    
    # Construct final prompt
    final_prompt = prompt_template.format(context=context, query=state['query'])
    
    # Call secured endpoint
    response = llm_proxy.invoke(final_prompt)
    state['final_response'] = response
    return state

workflow = StateGraph(SecureAgentState)
workflow.add_node("retrieve", retrieve_and_sanitize_node)
workflow.add_node("generate", generate_secure_response_node)
workflow.set_entry_point("retrieve")
workflow.add_edge("retrieve", "generate")
workflow.add_edge("generate", END)
app_graph = workflow.compile()

Frontend Implementation: Zero-Trust UI

The React frontend never handles sensitive keys. It sends only the user’s query and receives sanitized responses. We implement Content Security Policy (CSP) headers to prevent XSS attacks that could steal session tokens.

const SecureChat: React.FC = () => {
  const sendMessage = async (text: string) => {
    // No sensitive data in client-side logic
    const res = await axios.post('/api/inference', { query: text });
    setMessages(prev => [...prev, { role: 'ai', content: res.data.final_response }]);
  };
  
  return (
    <div className="secure-chat-interface">
      {/* UI Components */}
    </div>
  );
};

Real-Time Use Case: Confidential Legal Document Analysis

A law firm uses this system to analyze merger agreements.

  1. Input: A lawyer uploads a PDF containing names and financial figures.

  2. Processing: The Retrieve Node fetches relevant clauses. The Sanitize Node replaces names with [PERSON] and amounts with [AMOUNT] using Presidio.

  3. Generation: The Generate Node pulls a proprietary "Legal Risk Assessment" prompt from Vault. It sends the sanitized context to the LLM via a secured, rate-limited endpoint.

  4. Output: The lawyer receives a risk analysis without exposing client identities to the external LLM provider.

Conclusion

Securing an enterprise Multi-Agent RAG system requires a defense-in-depth strategy. By encrypting proprietary prompts in Vault, proxying model endpoints with strict access controls, and dynamically redacting sensitive context before inference, we create a robust security posture. This approach ensures that while the AI remains powerful and intelligent, it does not become a liability for data privacy or intellectual property theft. As AI adoption grows, these security patterns will become the standard for trustworthy enterprise automation.