Building a Document Q&A Application with Python, FastAPI, Google Cloud, and RAG

In Part 1, we introduced the PresalesAI architecture and local project setup. In this article, we focus on the ingestion pipeline: how uploaded PDF and Excel files become searchable knowledge.

The complete process is:

Upload
-> Store original file
-> Extract content
-> Create citation-aware chunks
-> Generate embeddings
-> Store vectors
-> Search relevant content

Uploading Documents

The frontend sends PDF and XLSX files to the FastAPI backend:

POST /documents/upload

The backend validates:

  • File extension
  • Supported format
  • Empty files
  • Maximum size of 10 MB

The supported formats are:

ALLOWED_UPLOAD_TYPES = {
".pdf": "application/pdf",
".xlsx": "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
}

The original file is stored in Google Cloud Storage before indexing:

gs://bucket-name/documents/<unique-id>/<filename>

Storing the original file separately is important because the extracted content can later be regenerated using a different processor or chunking strategy.

PDF Extraction with Document AI

PDF files are processed by the DocumentAIService.

PresalesAI uses Document AI processors for:

  • OCR text extraction
  • PDF pages
  • Tables
  • Layout information
  • Form entities where supported

The application prefers OCR for reliable text extraction:

def pdfprocessor(self) -> str:
if self.settings.document_ai_ocr_processor_id:
return "ocr"
if self.settings.document_ai_form_parser_processor_id:
return "form_parser"
if self.settings.document_ai_layout_processor_id:
return "layout"
return "ocr"

This order is intentional. General PDF documents usually need OCR text first. Form Parser and Layout Parser are useful for specialized document structures, but they may not expose text through the same page fields.

The extractor checks several Document AI structures:

for attribute in ("lines", "paragraphs", "blocks", "tokens", "words"):
...

If page-level text is unavailable, the application falls back to the full Document AI text response.

PDF Citations

Each PDF chunk stores its source page:

{
"filename": "proposal.pdf",
"page": 4
}

This allows the final answer to cite the original document:

Citation metadata travels with the text throughout the complete RAG pipeline.

Excel Extraction with openpyxl

Excel files do not require Document AI. PresalesAI uses openpyxl to read workbook content:

workbook = load_workbook(
BytesIO(content),
read_only=True,
data_only=True,
)

The processor loops through every worksheet and reads rows as values.

For example:

The row data is converted to searchable text:

Excel Chunking

Excel content is grouped into chunks of 25 rows by default:

Each chunk includes:

{
"filename": "pricing.xlsx",
"sheet": "Products",
"row_start": 1,
"row_end": 25
}

This produces citations such as:

The row range is more useful than a generic document citation because users can immediately locate the original spreadsheet data.

Why Chunking Is Required

Large documents should not be sent to Gemini in their entirety for every question.

Chunking provides several benefits:

  • Reduces prompt size
  • Improves retrieval accuracy
  • Reduces model cost
  • Preserves useful source boundaries
  • Makes citations easier
  • Allows search across many documents

A chunk should be large enough to preserve meaning but small enough to retrieve precisely.

For PDFs, the current implementation creates page-based chunks. For Excel files, it creates row-range chunks.

Creating Embeddings

Text embeddings convert content into numerical vectors.

The application uses Vertex AI:

EMBEDDING_MODEL=text-embedding-005

Each document chunk is converted into an embedding:

embeddings = self.embedding_model.get_embeddings(
[
TextEmbeddingInput(
chunk.text,
task_type="RETRIEVAL_DOCUMENT",
)
for chunk in chunks
]
)

A question is embedded using a different task type:

TextEmbeddingInput(
question,
task_type="RETRIEVAL_QUERY",
)

The document and query embeddings can then be compared semantically.

This means the search can find relevant content even when the user’s wording does not exactly match the document wording.

Storing Chunks in Firestore

Each chunk is stored in the document_chunks collection:

{
"document_id": document["document_id"],
"filename": document["filename"],
"text": chunk.text,
"citation": chunk.citation,
"embedding": Vector(embedding.values),
}

The collection contains both readable content and its vector representation.

The original document remains in Cloud Storage, while Firestore stores the searchable representation.

Firestore Vector Search

When a user asks a question, the question is converted into an embedding and searched against the complete collection:

self.firestore.collection("document_chunks").find_nearest(
vector_field="embedding",
query_vector=Vector(query_embedding.values),
distance_measure=DistanceMeasure.COSINE,
limit=self.settings.rag_top_k,
)

The default retrieval limit is:

This means the application retrieves up to six relevant chunks from all indexed documents.

There is no conversation-specific document filter. A question can retrieve information from any successfully indexed PDF or Excel file.

Relevance Filtering

Vector search always returns the nearest available chunks, even if they are not genuinely relevant. PresalesAI applies a distance threshold:

Chunks beyond that threshold are discarded.

This helps prevent unrelated content from being passed to Gemini. If all retrieved chunks are beyond the threshold, the application treats the question as unsupported by the indexed documents.

Firestore Vector Index

Firestore requires a vector index for nearest-neighbor queries.

The index is configured for:

  • Collection: document_chunks
  • Field: embedding
  • Dimensions: 768
  • Distance measure: Cosine

The index must reach the READY state before vector retrieval can work.

Without the index, Firestore returns a missing vector index error.

End-to-End Indexing Flow

The current upload and indexing flow is:

  1. User selects PDF or XLSX
  2. React sends the file to FastAPI
  3. FastAPI validates the file
  4. Original file is saved to Cloud Storage
  5. PDF or Excel content is extracted
  6. Content is divided into chunks
  7. Vertex AI creates embeddings
  8. Chunks are stored in Firestore
  9. The upload response reports indexing status

A successful response includes:

{
  "filename": "pricing.xlsx",
  "indexing_status": "complete",
  "chunk_count": 12,
  "gs_uri": "gs://bucket/documents/example/pricing.xlsx"
}

Handling Indexing Failures

Cloud Storage upload and indexing are separate operations.

This distinction matters. A file can be successfully stored while indexing fails because of:

Missing model configuration
-Document AI extraction failure
-Empty extracted text
-Missing Firestore vector index
-Vertex AI permission errors
-Invalid file contents

The application reports the indexing failure rather than claiming the document is searchable.

The file can then be reprocessed after the configuration problem is fixed.

Conclusion

The ingestion layer converts unstructured PDF and Excel files into searchable, citation-aware knowledge.

The main stages are:

Cloud Storage
  -> Document AI or openpyxl
  -> Citation-aware chunks
  -> Vertex AI embeddings
  -> Firestore vector search

This design allows PresalesAI to search across all indexed documents while preserving the source information required for trustworthy answers.

In Part 3, we will build the question-answering API that retrieves these chunks, constructs the Gemini prompt, returns citations, and asks for permission before using general knowledge.