Introduction

In Retrieval-Augmented Generation (RAG), the quality of the final answer depends heavily on the quality of the retrieved context.

Most enterprise RAG systems begin with vector similarity search. Embedding models such as text-embedding-ada-002 or bge-m3 convert queries and documents into vectors and retrieve documents that are semantically similar to the query.

Vector search is fast and effective, but semantic similarity is not always the same as relevance. It can miss exact terminology, subtle differences in intent, negation, or relationships between concepts.

This is where re-ranking becomes important.

Re-ranking is a second-stage retrieval process. Instead of asking a computationally expensive model to search the entire knowledge base, the system first retrieves a larger candidate set, such as the top 50 documents, and then uses a more accurate model to score those candidates against the original query.

In a multi-agent LangGraph architecture, the re-ranker can act as a quality gate between retrieval and answer generation. It ensures that downstream agents such as summarizers, support resolvers, or code generators receive a smaller set of highly relevant documents.

This can improve answer quality while reducing unnecessary context and token usage.

The Retrieval Gap: Why Vector Search Isn't Enough

Vector search is good at finding documents with related meanings, but semantic similarity alone does not guarantee that a document directly answers the user's question.

For example, consider a support query:

How do I reset a password?

A vector search might retrieve documents about:

  • Resetting a password

  • Password reset failures

  • Changing an expired password

  • Recovering a locked account

  • Password security policies

These documents may all be semantically related, but only some are directly relevant.

Common retrieval problems include:

Specificity

Vector search may struggle to distinguish between closely related questions.

For example:

  • "How do I reset a password?"

  • "Why did my password reset fail?"

Both queries contain similar concepts, but they require different documentation.

Negation

Negation can also create retrieval problems.

A document discussing non-compliant data may be semantically close to a query about compliant data formats, even though the distinction is important.

Recency and Priority

Enterprise knowledge bases often contain multiple versions of similar documentation.

A recent engineering document may contain the actual fix for a problem, while older documents contain outdated troubleshooting steps.

Vector similarity does not automatically understand that the newer document should take priority.

How Re-Ranking Works

A typical RAG pipeline can use two retrieval stages:

User Query
    |
    v
Hybrid / Vector Retrieval
    |
    v
Top 50 Candidate Documents
    |
    v
Cross-Encoder Re-Ranker
    |
    v
Top 5 Relevant Documents
    |
    v
LLM / Resolver Agent
    |
    v
Final Answer

The first stage is optimized for speed and recall.

The second stage is optimized for precision and relevance.

Re-ranking models such as BGE Reranker or other Cross-Encoder models evaluate the query and document together. Instead of independently embedding the query and document and comparing vectors, a Cross-Encoder can analyze interactions between the two pieces of text.

This deeper comparison generally makes the second stage ranking more suitable for determining which retrieved documents are actually useful for answering the query.

Real-Time Use Case: Technical Support Ticket Resolution

Consider a SaaS platform that receives thousands of technical support tickets every day.

Engineers need an AI assistant that can search:

  • API documentation

  • Previously resolved support tickets

  • Internal engineering documentation

  • Incident reports

  • Engineering logs

The assistant should then suggest a potential resolution.

The Problem

Suppose an engineer submits:

API 500 error during bulk upload

A vector search might return 50 documents containing information about HTTP 500 errors.

Many of those documents could contain generic troubleshooting steps such as:

  • Restart the service

  • Check server logs

  • Verify authentication

  • Check database connectivity

These documents are related to the query, but they may not identify the actual problem.

Suppose a recent engineering document contains:

Bulk upload requests exceeding the configured request rate can trigger a 500 response because of an internal rate-limiting bug.

That document is much more useful to the resolver, but it might not rank first in the initial vector search.

The Re-Ranking Solution

A LangGraph workflow can separate retrieval from reasoning:

  1. Retriever Agent retrieves the top 50 candidates using vector or hybrid search.

  2. Re-Ranker Agent evaluates those candidates and selects the top 5.

  3. Resolver Agent generates the suggested solution using the refined context.

  4. Memory Store records the resolution path for future retrieval or few-shot learning.

This architecture gives each stage a focused responsibility.

Technology Stack

A proof of concept can use the following technologies:

  • LangGraph

  • LangChain

  • Python

  • FastAPI

  • PostgreSQL

  • pgvector

  • Redis

  • Streamlit

  • Sentence Transformers

  • BGE Reranker

  • OpenAI GPT models

  • Docker

  • Pydantic

  • AsyncIO

  • Hybrid Search

  • Cross-Encoder models

  • Vector databases

  • LangGraph state management

Step-by-Step POC Implementation

Backend: LangGraph State Definition

The LangGraph state keeps track of both the initial retrieval results and the refined results produced by the re-ranker.

from typing import Annotated, List, TypedDict
from langgraph.graph.message import add_messages


class SupportState(TypedDict):
    ticket_query: str
    raw_candidates: List[dict]  # content, metadata, vector_score
    reranked_results: List[dict]  # content, metadata, rerank_score
    messages: Annotated[list, add_messages]
    suggested_fix: str | None

The important distinction is between raw_candidates and reranked_results.

The first contains the broader retrieval set, while the second contains only the documents that passed the re-ranking stage.

The Re-Ranker Node

A Cross-Encoder can be used to score each candidate against the original query.

from sentence_transformers import CrossEncoder


reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")


async def reranker_node(state: SupportState):
    query = state["ticket_query"]
    candidates = state["raw_candidates"]

    if not candidates:
        return {"reranked_results": []}

    docs = [doc["content"] for doc in candidates]

    pairs = [(query, doc) for doc in docs]
    scores = reranker.predict(pairs)

    scored_docs = [
        {
            **doc,
            "rerank_score": float(score)
        }
        for doc, score in zip(candidates, scores)
    ]

    top_results = sorted(
        scored_docs,
        key=lambda x: x["rerank_score"],
        reverse=True
    )[:5]

    return {"reranked_results": top_results}

The workflow first retrieves a relatively large candidate set. The re-ranker then evaluates each query-document pair and sorts the results according to the model's relevance score.

Only the highest-ranked documents are passed to the next stage.

For a production system, the number of candidates and final documents should be tuned against retrieval quality, latency, model cost, and context-window requirements.

Orchestrating the Multi-Agent Workflow

LangGraph can connect the retrieval, re-ranking, and resolution stages.

from langgraph.graph import StateGraph, END


workflow = StateGraph(SupportState)

workflow.add_node("retriever", retriever_node)
workflow.add_node("reranker", reranker_node)
workflow.add_node("resolver", resolver_node)

workflow.set_entry_point("retriever")

workflow.add_edge("retriever", "reranker")

workflow.add_conditional_edges(
    "reranker",
    lambda state: (
        "resolver"
        if state["reranked_results"]
        else END
    )
)

workflow.add_edge("resolver", END)

app = workflow.compile(checkpointer=memory_store)

The workflow creates a simple retrieval pipeline:

Retriever
   |
   v
Re-Ranker
   |
   +---- No relevant results ----> END
   |
   v
Resolver
   |
   v
Final Response

The conditional edge is useful because the resolver should not be called when the retrieval stage produces no usable context.

Frontend: Streamlit Visualization Dashboard

A Streamlit interface can expose the retrieved evidence to engineers and make the re-ranking process easier to inspect.

import streamlit as st
import requests


st.title("Tech Support RAG with Re-Ranking")

query = st.text_input(
    "Enter Ticket Issue:",
    "Bulk upload failing with 500 error"
)

if st.button("Find Solution"):
    with st.spinner("Retrieving and re-ranking..."):
        response = requests.post(
            "http://localhost:8000/support",
            json={"query": query}
        )

        data = response.json()

    st.subheader("Suggested Fix")
    st.info(data["suggested_fix"])

    st.subheader("Re-Ranking Evidence")

    for index, doc in enumerate(data["reranked_results"]):
        with st.expander(
            f"Result {index + 1} "
            f"(Score: {doc['rerank_score']:.4f})"
        ):
            st.write(doc["content"])
            st.caption(
                f"Source: {doc['metadata']['source']}"
            )

The dashboard provides two useful pieces of information:

  • The suggested resolution

  • The documents used as supporting evidence

This is particularly useful during development and debugging because engineers can inspect why a particular document was selected.

Important Production Considerations

Re-ranking improves retrieval quality, but it also introduces additional computation.

A production system should consider:

Candidate Count

Retrieving 50 candidates and re-ranking all 50 may provide better recall than re-ranking only 10, but it also increases latency.

The candidate count should be measured rather than chosen arbitrarily.

Final Context Size

Passing five large documents to an LLM can still create unnecessary context.

The final stage may need additional filtering, chunk selection, or context compression.

Model Hosting

A local Cross-Encoder can help keep sensitive enterprise data within the organization's infrastructure, but model inference still requires CPU or GPU resources.

Hybrid Retrieval

Combining keyword search with vector search can improve initial recall, especially for technical identifiers, API names, error codes, and exact terminology.

Observability

Track metrics such as:

  • Retrieval latency

  • Re-ranking latency

  • Initial candidate count

  • Final context count

  • Re-ranking scores

  • Answer quality

  • Retrieval failures

  • Token usage

These metrics help determine whether re-ranking is actually improving the system.

Conclusion

Re-ranking adds a precision layer to a RAG pipeline.

Vector or hybrid retrieval can quickly identify a broad set of potentially relevant documents. A Cross-Encoder can then examine those candidates more carefully and reduce the context passed to the answer-generation stage.

In the technical support example, the architecture becomes:

Fetch → Rank → Refine → Resolve

LangGraph provides a useful way to represent these stages explicitly. The retriever focuses on recall, the re-ranker focuses on relevance, and the resolver focuses on generating the final answer.

The result is not simply a larger retrieval pipeline. It is a more controlled architecture where each stage has a clear responsibility and where the quality of retrieved context can be measured and improved independently.