Building a Document Q&A Application with Python, FastAPI, Google Cloud, and RAG
In Part 1, we introduced the PresalesAI architecture and local project setup. In this article, we focus on the ingestion pipeline: how uploaded PDF and Excel files become searchable knowledge.
The complete process is:
Upload
-> Store original file
-> Extract content
-> Create citation-aware chunks
-> Generate embeddings
-> Store vectors
-> Search relevant content
Uploading Documents
The frontend sends PDF and XLSX files to the FastAPI backend:
POST /documents/upload
The backend validates:
- File extension
- Supported format
- Empty files
- Maximum size of 10 MB
The supported formats are:
ALLOWED_UPLOAD_TYPES = {
".pdf": "application/pdf",
".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
}
The original file is stored in Google Cloud Storage before indexing:
gs://bucket-name/documents/<unique-id>/<filename>
Storing the original file separately is important because the extracted content can later be regenerated using a different processor or chunking strategy.
PDF Extraction with Document AI
PDF files are processed by the DocumentAIService.
PresalesAI uses Document AI processors for:
- OCR text extraction
- PDF pages
- Tables
- Layout information
- Form entities where supported
The application prefers OCR for reliable text extraction:
def pdfprocessor(self) -> str:
if self.settings.document_ai_ocr_processor_id:
return "ocr"
if self.settings.document_ai_form_parser_processor_id:
return "form_parser"
if self.settings.document_ai_layout_processor_id:
return "layout"
return "ocr"
This order is intentional. General PDF documents usually need OCR text first. Form Parser and Layout Parser are useful for specialized document structures, but they may not expose text through the same page fields.
The extractor checks several Document AI structures:
for attribute in ("lines", "paragraphs", "blocks", "tokens", "words"):
...
If page-level text is unavailable, the application falls back to the full Document AI text response.
PDF Citations
Each PDF chunk stores its source page:
{
"filename": "proposal.pdf",
"page": 4
}
This allows the final answer to cite the original document:
Citation metadata travels with the text throughout the complete RAG pipeline.
Excel Extraction with openpyxl
Excel files do not require Document AI. PresalesAI uses openpyxl to read workbook content:
workbook = load_workbook(
BytesIO(content),
read_only=True,
data_only=True,
)
The processor loops through every worksheet and reads rows as values.
For example:
The row data is converted to searchable text:
Excel Chunking
Excel content is grouped into chunks of 25 rows by default:
Each chunk includes:
{
"filename": "pricing.xlsx",
"sheet": "Products",
"row_start": 1,
"row_end": 25
}
This produces citations such as:
The row range is more useful than a generic document citation because users can immediately locate the original spreadsheet data.
Why Chunking Is Required
Large documents should not be sent to Gemini in their entirety for every question.
Chunking provides several benefits:
- Reduces prompt size
- Improves retrieval accuracy
- Reduces model cost
- Preserves useful source boundaries
- Makes citations easier
- Allows search across many documents
A chunk should be large enough to preserve meaning but small enough to retrieve precisely.
For PDFs, the current implementation creates page-based chunks. For Excel files, it creates row-range chunks.
Creating Embeddings
Text embeddings convert content into numerical vectors.
The application uses Vertex AI:
EMBEDDING_MODEL=text-embedding-005
Each document chunk is converted into an embedding:
embeddings = self.embedding_model.get_embeddings(
[
TextEmbeddingInput(
chunk.text,
task_type="RETRIEVAL_DOCUMENT",
)
for chunk in chunks
]
)
A question is embedded using a different task type:
TextEmbeddingInput(
question,
task_type="RETRIEVAL_QUERY",
)
The document and query embeddings can then be compared semantically.
This means the search can find relevant content even when the user’s wording does not exactly match the document wording.
Storing Chunks in Firestore
Each chunk is stored in the document_chunks collection:
{
"document_id": document["document_id"],
"filename": document["filename"],
"text": chunk.text,
"citation": chunk.citation,
"embedding": Vector(embedding.values),
}
The collection contains both readable content and its vector representation.
The original document remains in Cloud Storage, while Firestore stores the searchable representation.
Firestore Vector Search
When a user asks a question, the question is converted into an embedding and searched against the complete collection:
self.firestore.collection("document_chunks").find_nearest(
vector_field="embedding",
query_vector=Vector(query_embedding.values),
distance_measure=DistanceMeasure.COSINE,
limit=self.settings.rag_top_k,
)
The default retrieval limit is:
This means the application retrieves up to six relevant chunks from all indexed documents.
There is no conversation-specific document filter. A question can retrieve information from any successfully indexed PDF or Excel file.
Relevance Filtering
Vector search always returns the nearest available chunks, even if they are not genuinely relevant. PresalesAI applies a distance threshold:
Chunks beyond that threshold are discarded.
This helps prevent unrelated content from being passed to Gemini. If all retrieved chunks are beyond the threshold, the application treats the question as unsupported by the indexed documents.
Firestore Vector Index
Firestore requires a vector index for nearest-neighbor queries.
The index is configured for:
- Collection: document_chunks
- Field: embedding
- Dimensions: 768
- Distance measure: Cosine
The index must reach the READY state before vector retrieval can work.
Without the index, Firestore returns a missing vector index error.
End-to-End Indexing Flow
The current upload and indexing flow is:
- User selects PDF or XLSX
- React sends the file to FastAPI
- FastAPI validates the file
- Original file is saved to Cloud Storage
- PDF or Excel content is extracted
- Content is divided into chunks
- Vertex AI creates embeddings
- Chunks are stored in Firestore
- The upload response reports indexing status
A successful response includes:
{
"filename": "pricing.xlsx",
"indexing_status": "complete",
"chunk_count": 12,
"gs_uri": "gs://bucket/documents/example/pricing.xlsx"
}
Handling Indexing Failures
Cloud Storage upload and indexing are separate operations.
This distinction matters. A file can be successfully stored while indexing fails because of:
Missing model configuration
-Document AI extraction failure
-Empty extracted text
-Missing Firestore vector index
-Vertex AI permission errors
-Invalid file contents
The application reports the indexing failure rather than claiming the document is searchable.
The file can then be reprocessed after the configuration problem is fixed.
Conclusion
The ingestion layer converts unstructured PDF and Excel files into searchable, citation-aware knowledge.
The main stages are:
Cloud Storage
-> Document AI or openpyxl
-> Citation-aware chunks
-> Vertex AI embeddings
-> Firestore vector search
This design allows PresalesAI to search across all indexed documents while preserving the source information required for trustworthy answers.
In Part 3, we will build the question-answering API that retrieves these chunks, constructs the Gemini prompt, returns citations, and asks for permission before using general knowledge.

Join the conversation! Your thoughts help the community grow.