A RAG application can retrieve relevant documents and give them to a large language model as context.
But there is another problem.
How does the application show where the answer came from?
For an internal chatbot, saying:
"The company allows 20 days of annual leave."may not be enough.
The user may also need:
Source: Employee Handbook
Page: 18This becomes even more important in enterprise applications where users need to verify answers before acting on them.
Amazon Bedrock Knowledge Bases supports retrieval workflows that can return source information along with retrieved content. Bedrock's Retrieve and RetrieveAndGenerate APIs expose citations and source references that applications can use when presenting generated answers.
This article explains how citations work, how to build a RAG workflow around them, and what developers should consider when turning retrieved source metadata into a trustworthy user experience.
What Is Amazon Bedrock Knowledge Bases?
Amazon Bedrock Knowledge Bases provides managed capabilities for connecting foundation models with external data.
A simplified RAG architecture looks like this:
User Question
|
v
Knowledge Base
|
v
Retrieve Relevant Content
|
v
Foundation Model
|
v
Generated AnswerWithout RAG, the model primarily relies on what it learned during training and the context provided at request time.
With a Knowledge Base, the application can retrieve current information from connected data sources and use that information as context.
The retrieval process can also return metadata about the source documents.
That gives the application two outputs:
Retrieved Content
+
Source MetadataThe second part is what makes citations possible.
Why Do Citations Matter?
Consider two chatbot responses.
Response Without Citation
Our refund policy allows customers to
request a refund within 30 days.The user has no direct way to verify the statement.
Response With Citation
Our refund policy allows customers to
request a refund within 30 days.
Source:
Refund Policy
Section: Customer RefundsThe second response is easier to trust because the user can inspect the supporting material.
This matters especially for:
Internal company policies
Technical documentation
Product documentation
Compliance information
Support knowledge bases
Financial procedures
Legal documentation
Enterprise RAG systems
Citations do not make an AI answer automatically correct.
They make the answer more verifiable.
How Retrieval and Generation Work
A typical Bedrock Knowledge Base flow can be represented as:
Question
|
v
Retrieve
|
+---- Chunk 1
+---- Chunk 2
+---- Chunk 3
|
v
Context
|
v
Foundation Model
|
v
Answer + CitationsThere are two useful API patterns to understand.
Retrieve
Retrieve focuses on retrieval.
The application receives relevant chunks and their metadata.
This is useful when you want to control the generation step yourself.
Application
|
v
Retrieve
|
v
Chunks + Metadata
|
v
Your Prompt
|
v
Your Model CallRetrieveAndGenerate
RetrieveAndGenerate combines retrieval and generation.
Application
|
v
RetrieveAndGenerate
|
+---- Retrieval
|
+---- Generation
|
v
Generated Answer
+
CitationsThis is useful when you want Bedrock to manage the retrieval-to-generation workflow.
What Does a Citation Contain?
A citation is more than a URL.
The response can contain information connecting a generated section to the retrieved source.
Conceptually:
{
"generatedResponsePart": {
"textResponsePart": {
"text": "The refund period is 30 days."
}
},
"retrievedReferences": [
{
"content": {
"text": "Customers may request a refund within 30 days..."
},
"location": {
"type": "S3",
"s3Location": {
"uri": "s3://company-docs/refund-policy.pdf"
}
}
}
]
}The exact response structure depends on the API operation and source type.
The important relationship is:
Generated Text
|
v
Citation
|
v
Retrieved Reference
|
v
Original SourceYour application can use that relationship to build a citation UI.
A Basic RAG Application Architecture
Suppose a company stores documentation in Amazon S3.
The architecture might be:
Amazon S3
|
v
Knowledge Base
|
v
Vector Store
|
v
User Question ---> Retrieval
|
v
Foundation Model
|
v
Answer + Citations
|
v
Web ApplicationThe user might ask:
"What is the process for requesting parental leave?"The Knowledge Base retrieves the relevant policy content.
The model generates:
Employees should submit a parental leave request
through the HR portal and provide the required
documentation.The application can then show:
Sources
1. Employee Leave Policy
Section: Parental LeaveIngesting Documents
Before retrieval can work, documents need to be available to the Knowledge Base.
A typical ingestion pipeline looks like:
Documents
|
v
Amazon S3
|
v
Knowledge Base
|
v
Parsing / Chunking
|
v
Embeddings
|
v
Vector StoreThe exact ingestion behavior depends on the data source and Knowledge Base configuration.
The important thing for developers is that retrieval quality depends heavily on the content being indexed.
A poor document structure can produce poor retrieval.
Chunking Affects Citations
Suppose a PDF contains:
Employee Benefitswith several pages of unrelated information.
If the content is split into poor chunks, a question about health insurance may retrieve a large section containing unrelated material.
That can lead to:
Question
|
v
Poor Chunk
|
v
Weak Context
|
v
Weak AnswerGood chunking produces:
Question
|
v
Relevant Chunk
|
v
Focused Context
|
v
Better AnswerIt also improves citation quality because the source reference is tied to more relevant content.
Retrieval Quality Comes Before Citation Quality
A citation can tell you where the model got information.
It cannot fix bad retrieval.
Consider:
User Question
|
v
Retriever
|
X
Wrong Document
|
v
Model
|
v
Confident AnswerThe answer may contain a valid citation.
But it is still wrong.
That is why RAG evaluation should measure retrieval quality separately from generation quality.
Useful evaluation questions include:
Did the retriever find the right document?
Did it find the correct section?
Did the generated answer stay within the retrieved evidence?
Does the citation actually support the claim?
A .NET Retrieval Example
A .NET application can use the AWS SDK to call Bedrock services.
A simplified retrieval workflow might look like:
using Amazon.BedrockAgentRuntime;
using Amazon.BedrockAgentRuntime.Model;
var client = new AmazonBedrockAgentRuntimeClient();
var request = new RetrieveRequest
{
KnowledgeBaseId = "YOUR_KNOWLEDGE_BASE_ID",
RetrievalQuery = new KnowledgeBaseQuery
{
Text = "What is the refund policy?"
}
};
var response = await client.RetrieveAsync(request);
foreach (var result in response.RetrievalResults)
{
Console.WriteLine(result.Content.Text);
Console.WriteLine(result.Location?.Type);
}In a real application, the Knowledge Base ID and AWS credentials should come from configuration or the application's AWS identity rather than being hard-coded.
The important part of this workflow is that retrieval results can contain both content and location metadata.
Building Your Own Generation Step
Using Retrieve gives you more control over the generation process.
For example:
Question
|
v
Bedrock Retrieve
|
+---- Chunk A
+---- Chunk B
+---- Chunk C
|
v
Build Prompt
|
v
Foundation Model
|
v
AnswerYour prompt can explicitly instruct the model:
Answer the question using only the provided context.
If the context does not contain enough information,
say that the answer cannot be determined.
Do not invent facts.
Return the source identifiers associated with
claims in the answer.The application then controls how citations are rendered.
This approach provides more flexibility but also gives the application more responsibility.
Using RetrieveAndGenerate
If you do not need to control every step, RetrieveAndGenerate can simplify the workflow.
Conceptually:
Question
|
v
RetrieveAndGenerate
|
+---- Retrieve
|
+---- Generate
|
v
Response
|
+---- Answer
|
+---- CitationsA simplified .NET example is:
using Amazon.BedrockAgentRuntime;
using Amazon.BedrockAgentRuntime.Model;
var client = new AmazonBedrockAgentRuntimeClient();
var request = new RetrieveAndGenerateRequest
{
Input = new RetrieveAndGenerateInput
{
Text = "What is the refund policy?"
},
RetrieveAndGenerateConfiguration =
new RetrieveAndGenerateConfiguration
{
Type = RetrieveAndGenerateType.KNOWLEDGE_BASE,
KnowledgeBaseConfiguration =
new KnowledgeBaseRetrieveAndGenerateConfiguration
{
KnowledgeBaseId = "YOUR_KNOWLEDGE_BASE_ID"
}
}
};
var response = await client.RetrieveAndGenerateAsync(request);
Console.WriteLine(response.Output.Text);The SDK model types can vary between versions, so production code should be written against the exact AWS SDK version used by the application.
The important architectural point is that the generated response can be accompanied by citation information.
Rendering Citations in a Web Application
Do not simply dump the complete retrieval metadata onto the screen.
A better UI might look like:
Answer
Customers can request a refund within 30 days.
Sources
[1] Refund Policy
Customer Refunds
[2] Customer Terms
Section 4.2The backend can maintain a mapping:
Claim 1
|
+-- Citation 1
|
+-- Citation 2
Claim 2
|
+-- Citation 3The frontend can then render citations next to the relevant text.
For example:
Customers can request a refund within 30 days. [1]
The request must be submitted through the customer
portal. [2]This is much more useful than displaying a generic "Sources" list at the bottom.
Do Not Trust Citations Blindly
A citation only tells you what source information was associated with the response.
It does not prove that the generated statement is fully supported.
For example:
Source:
Refund Policy
Generated:
"Refunds are always approved within 24 hours."The source might actually say:
"Refund requests are normally reviewed within
24 hours."Those statements are not equivalent.
Your application should therefore evaluate citation quality as well as citation presence.
Citation Validation
For higher-risk applications, add a validation step.
Generated Answer
|
v
Claim Extraction
|
v
Citation Mapping
|
v
Evidence Check
|
v
Final AnswerA simpler approach is to instruct the model to stay strictly within retrieved context.
A stronger architecture can use a second model or deterministic rules to check whether important claims are supported.
The appropriate level depends on the risk of the application.
Metadata Filtering
Enterprise RAG systems often need metadata filters.
Consider a multi-tenant application:
Tenant A
|
+-- HR Documents
+-- Finance Documents
Tenant B
|
+-- HR Documents
+-- Finance DocumentsA user from Tenant A should not retrieve Tenant B's documents.
The retrieval architecture should therefore include authorization metadata:
User
|
+-- TenantId
+-- Role
+-- Department
|
v
Knowledge Base Retrieval
|
v
Authorized DocumentsFiltering should be enforced by the application architecture rather than relying on the model to decide which documents a user is allowed to see.
This is one of the most important security considerations in enterprise RAG.
Source Access and User Permissions
A citation can expose information even if the underlying document is not directly downloadable.
For example:
User asks:
"What is our acquisition strategy?"
Agent:
"Our acquisition strategy is..."
Source: confidential-strategy.pdfEven if the user cannot open the PDF, the generated answer may reveal its contents.
Therefore:
Document Authorization
+
Retrieval Authorization
+
Answer Authorizationshould be considered together.
Do not assume that restricting document URLs is enough.
Handling Missing Information
A good RAG application must know when it does not have enough evidence.
For example:
User:
"What will our revenue be next year?"If the Knowledge Base only contains historical reports, the system should not invent a forecast.
A good response is:
I could not find enough information in the
available documents to answer that question.This behavior should be part of the system prompt and application design.
The goal is:
No Evidence
|
v
No Confident Claimrather than:
No Evidence
|
v
GuessCommon Mistakes
Adding Citations After Generation
Do not generate an answer first and then attach random documents as sources.
The citations should originate from the retrieval context used to produce the answer.
Treating Any Retrieved Document as Evidence
Retrieval relevance does not guarantee that the document supports the generated claim.
Ignoring Authorization
Do not assume that because the Knowledge Base can retrieve a document, every user should be allowed to see its content.
Returning Entire Documents
Retrieve the relevant chunks needed for the question.
Sending excessive context can increase cost and reduce answer quality.
Using Poor Chunking
Large, unrelated chunks make retrieval and citations less useful.
Treating Citations as Proof of Accuracy
A source reference is not the same thing as factual verification.
Troubleshooting
The Answer Has No Citations
Check whether the application is using the API operation and response fields that provide retrieval reference information.
If you are using Retrieve, remember that you are responsible for building the generation and citation experience yourself.
If you use RetrieveAndGenerate, inspect the response for generated response citations.
Citations Point to Irrelevant Documents
Investigate retrieval first.
Check:
Chunk size
Chunk overlap
Embedding model
Metadata
Query formulation
Number of retrieved results
Correct Document but Wrong Section
The document may be relevant while the retrieved chunk is not.
Improve document structure and chunking.
Users See Data They Should Not See
Review authorization and metadata filtering.
Do not depend on the model to enforce tenant boundaries.
The Agent Gives Unsupported Answers
Add stronger grounding instructions:
Use only information present in the retrieved context.
If the context does not contain the answer,
say that the information is unavailable.
Do not infer unsupported facts.Then evaluate the result against a test dataset.
Measuring RAG Quality
A production RAG system should have an evaluation set.
For example:
Question Expected Source
------------------------------------------------
Refund period Refund Policy
Password reset Account Guide
Leave approval HR Handbook
Expense limit Travel PolicyRun these questions regularly.
Measure:
Retrieval Accuracy
Citation Accuracy
Answer Accuracy
Unsupported Claims
Latency
CostA useful test is:
Question
|
+--> Expected Document
|
v
Actual Retrieval
|
v
CompareThis lets you detect changes after modifying chunking, embeddings, retrieval settings, or prompts.
Advantages
Better Trust
Users can inspect the source behind an answer.
Easier Debugging
Developers can investigate whether an incorrect answer came from retrieval or generation.
Better Enterprise Adoption
Teams are more comfortable with AI systems when answers can be traced back to company documentation.
Useful Audit Trail
Source references can help explain why an answer was produced.
Flexible Architecture
Developers can either control retrieval and generation separately or use the combined retrieval-and-generation workflow.
Limitations
Citations Do Not Guarantee Correctness
The source still needs to support the claim.
Retrieval Quality Controls the Result
Bad retrieval leads to bad context.
Authorization Is Complex
Enterprise data requires access control before information reaches the model.
Additional Processing
Citation rendering, validation, and source mapping add application complexity.
Source Documents Can Be Wrong
A perfectly cited answer is still wrong if the underlying company documentation is outdated.
Best Practices
Treat retrieval quality as a first-class engineering problem.
Use citations directly associated with retrieved references.
Display citations next to the claims they support when possible.
Do not expose documents or claims a user is not authorized to access.
Use metadata filtering for tenant and role boundaries.
Instruct the model not to answer beyond retrieved evidence.
Return an explicit "not enough information" response when evidence is missing.
Evaluate retrieval and generation separately.
Test citation accuracy, not just citation presence.
Keep source documents current.
Monitor retrieval latency, model latency, and token usage.
Maintain a regression test set for important business questions.
When Should You Use Knowledge Base Citations?
This pattern is particularly useful for:
Enterprise documentation assistants
Customer support systems
Internal policy assistants
Technical documentation search
Compliance knowledge bases
Product support
RAG applications where users need evidence
It is especially valuable when the answer needs to be explainable.
For a casual chatbot, citations may be optional.
For an enterprise assistant making statements about policies or procedures, they can be essential.
Final Thoughts
RAG is not just about getting an AI model to answer questions from documents.
In production, users also need to understand why they should trust the answer.
Amazon Bedrock Knowledge Bases provides the retrieval layer, while Bedrock retrieval APIs can expose source information that applications can use to build citation-aware experiences.
That makes the architecture more useful:
Question
|
v
Retrieve
|
+---- Evidence
|
v
Generate
|
+---- Answer
|
+---- Citations
|
v
UserBut citations should not become a cosmetic feature.
The real goal is evidence-backed answers.
A production-quality RAG system should retrieve the correct information, respect authorization boundaries, generate an answer grounded in that information, and clearly show users which sources support the answer.
If those four pieces work together, citations become more than links at the bottom of an AI response.
They become part of the trust model for the application.

Join the conversation! Your thoughts help the community grow.