Introduction
In the enterprise adoption of Large Language Models (LLMs), organizations often face a critical strategic question: How do we adapt a general-purpose model to our specific business needs? The answer lies in understanding three distinct but complementary techniques: Pre-training, Fine-tuning, and Prompt Engineering.
Pre-training is the foundational phase where a model learns general language patterns from massive, unstructured datasets. It’s expensive and rarely done by individual enterprises.
Fine-tuning involves further training a pre-trained model on a smaller, domain-specific dataset to specialize its behavior or knowledge.
Prompt Engineering is the art of crafting inputs to guide the model’s output without changing its weights, leveraging in-context learning.
For modern enterprise applications, relying on just one technique is insufficient. A robust system combines all three: using a pre-trained base, fine-tuned for domain-specific terminology, and guided by sophisticated prompt engineering within a multi-agent workflow. In this article, we will build a complete Proof-of-Concept (PoC) for an Enterprise Legal Contract Review System. This system will use LangGraph to orchestrate agents that utilize these techniques alongside Graph RAG for contextual accuracy and persistent memory for state management.
Real-Time Use Case: Intelligent Legal Contract Review
Imagine a law firm processing hundreds of Non-Disclosure Agreements (NDAs) daily. Manually reviewing them for risky clauses is time-consuming and error-prone. Our system will:
Ingest contract text.
Retrieve relevant legal precedents and company policies using Graph RAG.
Analyze clauses using a fine-tuned legal LLM agent.
Generate a review report using prompt-engineered templates.
Maintain State across multiple revisions and user feedback.
Step-by-Step Implementation
Step 1: Environment Setup
# requirements.txt
langgraph==0.2.0
langchain==0.1.0
langchain-community==0.0.10
torch==2.1.0
transformers==4.35.0
fastapi==0.109.0
uvicorn==0.27.0
pydantic==2.5.0
chromadb==0.4.22
networkx==3.2.1
python-dotenv==1.0.0
Step 2: Simulating the "Fine-Tuned" Model Adapter
Since actual fine-tuning requires significant resources, we will simulate a fine-tuned legal model using a specialized prompt wrapper around a base model. In a real scenario, this would be a LoRA-adapted model.
# models/legal_llm.py
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI # Or any other provider
import os
class LegalLLMAdapter:
def __init__(self):
# In production, this would load a fine-tuned model checkpoint
# For PoC, we use a base model with strong prompt engineering to simulate fine-tuning
self.llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.1)
# This prompt represents the "Fine-Tuned" behavior via few-shot prompting
self.review_prompt = ChatPromptTemplate.from_template("""
You are an expert Legal AI Assistant specialized in Contract Law.
Your task is to review the following clause from an NDA.
Company Policy: {company_policy}
Clause: {clause_text}
Retrieved Precedents: {precedents}
Instructions:
1. Identify any risky or non-standard terms.
2. Compare against the provided precedents.
3. Provide a risk score (1-10).
4. Suggest a revised clause if necessary.
Format your response as JSON:
{{
"risk_score": int,
"issues": [list of strings],
"suggested_revision": string
}}
""")
def review_clause(self, clause: str, policy: str, precedents: str) -> dict:
chain = self.review_prompt | self.llm
response = chain.invoke({
"clause_text": clause,
"company_policy": policy,
"precedents": precedents
})
# In production, parse JSON properly
return {"content": response.content}
Step 3: Graph RAG Service for Legal Precedents
# services/graph_rag.py
import chromadb
import networkx as nx
from typing import List, Dict
class LegalGraphRAG:
def __init__(self):
self.client = chromadb.PersistentClient(path="./legal_chroma")
self.collection = self.client.get_or_create_collection("legal_precedents")
self.graph = nx.DiGraph()
def add_precedent(self, doc_id: str, text: str, metadata: Dict):
self.collection.add(documents=[text], ids=[doc_id], metadatas=[metadata])
self.graph.add_node(doc_id, **metadata)
if metadata.get('cites'):
self.graph.add_edge(doc_id, metadata['cites'])
def search_precedents(self, query: str, n_results: int = 2) -> List[str]:
results = self.collection.query(query_texts=[query], n_results=n_results)
return results['documents'][0] if results['documents'] else []
Step 4: LangGraph Multi-Agent Workflow with State
# agents/workflow.py
from langgraph.graph import StateGraph, END
from typing import TypedDict, List, Optional
from models.legal_llm import LegalLLMAdapter
from services.graph_rag import LegalGraphRAG
class ContractState(TypedDict):
contract_text: str
clauses: List[str]
current_clause_index: int
reviews: List[dict]
final_report: Optional[str]
company_policy: str
memory: List[dict] # Persistent memory for user interactions
class ContractReviewAgent:
def __init__(self):
self.llm_adapter = LegalLLMAdapter()
self.rag = LegalGraphRAG()
# Load some dummy data for PoC
self.rag.add_precedent("p1", "Standard NDA clause: Confidentiality lasts 2 years.", {"type": "confidentiality"})
self.rag.add_precedent("p2", "High Risk: Indefinite confidentiality is non-standard.", {"type": "risk"})
def extract_clauses(self, state: ContractState) -> ContractState:
"""Simple split for PoC. In prod, use an LLM to extract clauses."""
state['clauses'] = state['contract_text'].split('.')
state['current_clause_index'] = 0
state['reviews'] = []
return state
def analyze_clause(self, state: ContractState) -> ContractState:
"""The Core Agent: Uses Fine-Tuned Logic + RAG"""
if state['current_clause_index'] >= len(state['clauses']):
return state
clause = state['clauses'][state['current_clause_index']]
if not clause.strip():
state['current_clause_index'] += 1
return state
# Retrieve context
precedents = self.rag.search_precedents(clause)
precedent_text = "\n".join(precedents)
# Invoke "Fine-Tuned" Model
review = self.llm_adapter.review_clause(clause, state['company_policy'], precedent_text)
state['reviews'].append({
"clause": clause,
"analysis": review
})
state['current_clause_index'] += 1
return state
def should_continue(self, state: ContractState) -> str:
if state['current_clause_index'] < len(state['clauses']):
return "analyze"
else:
return "generate_report"
def generate_report(self, state: ContractState) -> ContractState:
"""Prompt Engineering for Final Output"""
summary = "\n".join([f"- Clause: {r['clause'][:50]}... | Risk: {r['analysis']}" for r in state['reviews']])
state['final_report'] = f"Contract Review Complete.\n\nFindings:\n{summary}"
state['memory'].append({"action": "review_completed", "report": state['final_report']})
return state
def build_workflow():
agent = ContractReviewAgent()
workflow = StateGraph(ContractState)
workflow.add_node("extract", agent.extract_clauses)
workflow.add_node("analyze", agent.analyze_clause)
workflow.add_node("report", agent.generate_report)
workflow.set_entry_point("extract")
workflow.add_edge("extract", "analyze")
workflow.add_conditional_edges(
source="analyze",
path_map={
"analyze": "analyze",
"generate_report": "report"
},
condition=agent.should_continue
)
workflow.add_edge("report", END)
return workflow.compile()
Step 5: FastAPI Backend
# main.py
from fastapi import FastAPI
from pydantic import BaseModel
from agents.workflow import build_workflow
app = FastAPI(title="Legal Contract Review API")
workflow = build_workflow()
class ContractRequest(BaseModel):
text: str
policy: str = "Standard Company Policy: Confidentiality must not exceed 3 years."
class ReviewResponse(BaseModel):
report: str
@app.post("/review", response_model=ReviewResponse)
async def review_contract(request: ContractRequest):
initial_state = {
"contract_text": request.text,
"clauses": [],
"current_clause_index": 0,
"reviews": [],
"final_report": None,
"company_policy": request.policy,
"memory": []
}
result = await workflow.ainvoke(initial_state)
return ReviewResponse(report=result['final_report'])
if __name__ == "__main__":
import uvicorn
uvicorn.run(app, host="0.0.0.0", port=8000)
Step 6: Frontend Interface
<!-- index.html -->
<!DOCTYPE html>
<html>
<head>
<title>Legal Contract Reviewer</title>
<style>
body { font-family: 'Segoe UI', sans-serif; max-width: 900px; margin: 40px auto; padding: 20px; background: #f9f9f9; }
textarea { width: 100%; height: 150px; padding: 10px; border: 1px solid #ccc; border-radius: 5px; }
button { background: #2c3e50; color: white; padding: 12px 25px; border: none; border-radius: 5px; cursor: pointer; font-size: 16px; }
button:hover { background: #34495e; }
.report { background: white; padding: 20px; margin-top: 20px; border-left: 5px solid #2c3e50; box-shadow: 0 2px 5px rgba(0,0,0,0.1); }
</style>
</head>
<body>
<h1>Enterprise Legal Contract Review</h1>
<p>Paste your NDA text below:</p>
<textarea id="contractText">This agreement shall remain in effect indefinitely. The receiving party agrees to keep all information confidential.</textarea>
<br><br>
<button onclick="submitContract()">Review Contract</button>
<div id="output"></div>
<script>
async function submitContract() {
const text = document.getElementById('contractText').value;
const btn = document.querySelector('button');
btn.innerText = "Processing...";
btn.disabled = true;
try {
const response = await fetch('/review', {
method: 'POST',
headers: {'Content-Type': 'application/json'},
body: JSON.stringify({text: text})
});
const data = await response.json();
document.getElementById('output').innerHTML = `
<div class="report">
<h3>Review Report</h3>
<pre>${data.report}</pre>
</div>
`;
} catch (error) {
alert("Error processing contract");
} finally {
btn.innerText = "Review Contract";
btn.disabled = false;
}
}
</script>
</body>
</html>
Conclusion
This PoC demonstrates how Pre-training, Fine-tuning, and Prompt Engineering converge in a modern enterprise architecture. While we used a base model, the LegalLLMAdapter simulates fine-tuned behavior through specialized prompts and few-shot examples. The LangGraph workflow orchestrates the process, ensuring that each clause is processed with context retrieved via Graph RAG, maintaining a stateful memory of the review process. By combining these techniques, enterprises can build systems that are not only intelligent but also compliant, contextual, and capable of handling complex, multi-step reasoning tasks like legal contract review. This approach moves beyond simple chatbots to create true AI agents that augment human expertise.

Join the conversation! Your thoughts help the community grow.