Search has become a bigger database problem as applications have started mixing traditional text search with AI-powered semantic search.
A keyword search is good at finding an exact term. Vector search is good at finding something that means roughly the same thing even when the wording is different.
Neither approach is enough for every application.
Consider a user searching an ecommerce catalog for:
comfortable shoes for long-distance runningA vector search may understand the intent and find products described as "cushioned footwear for marathon training."
A traditional full-text search may be better at an exact query such as:
Nike PegasusThe problem comes when an application needs both types of relevance at the same time.
AlloyDB for PostgreSQL now provides a database-native hybrid search function that combines vector and full-text search results using Reciprocal Rank Fusion (RRF). Google Cloud announced the updated hybrid-search capabilities on October 6, 2026. The implementation is exposed through the ai.hybrid_search() function, which combines ranked search components and produces a unified result set.
That sounds like a small SQL feature.
It is actually more useful than that because it removes a common piece of application-side search infrastructure: running two searches, trying to normalize incompatible scores, joining the results, and implementing the final ranking logic yourself.
Why Hybrid Search Is Useful
Traditional full-text search and vector search solve different problems.
Full-text search is based on words and linguistic matching.
For example:
Query:
"PostgreSQL connection pooling"A text search can prioritize documents containing those terms.
Vector search works differently.
The query is converted into an embedding and compared with embeddings stored alongside documents.
User Query
|
v
Embedding Model
|
v
Query Vector
|
v
Vector Similarity Search
|
v
Semantically Similar DocumentsA document that says:
How to reuse database connections efficientlymay be considered relevant even though it does not contain the exact phrase "connection pooling."
The two search methods therefore complement each other.
Google Cloud's AlloyDB documentation describes hybrid search as combining keyword precision with semantic similarity to produce results that are both lexically and semantically relevant.
The Problem With Combining Two Search Systems
Before a database-native approach, an application commonly had to do something like this:
Search Query
|
+----------+----------+
| |
v v
Vector Search Text Search
| |
v v
Ranked List A Ranked List B
| |
+----------+----------+
|
v
Normalize Scores
|
v
Merge Results
|
v
Re-rank
|
v
Final ResultsThere is a subtle problem here.
Vector search and text search usually produce different kinds of scores.
A vector distance might look like:
0.08
0.13
0.19while a text-ranking function might return:
7.4
4.8
2.1Those numbers do not necessarily have a common meaning.
Simply adding them together is not a reliable ranking strategy.
Teams therefore end up writing score-normalization logic or custom ranking code.
Google Cloud specifically identifies this score-normalization step as one of the difficult parts of traditional hybrid-search implementations.
RRF takes a different approach.
What Is Reciprocal Rank Fusion?
RRF does not try to make different search scores directly comparable.
Instead, it uses the position of a result in each ranked list.
Suppose vector search returns:
1. Document A
2. Document B
3. Document Cand text search returns:
1. Document C
2. Document A
3. Document DRRF gives each document a contribution based on its rank.
A simplified formula is:
RRF score = 1 / (k + rank)where k is a constant used to reduce the influence of very high ranks.
If a document appears near the top of both lists, it receives contributions from both.
For example:
Vector rank Text rank
------------ ---------
A: 1 A: 2
B: 2 -
C: 3 C: 1
D: - D: 3The combined ranking can then favor documents that perform well across both search methods.
This is useful because the algorithm works with rank rather than attempting to compare fundamentally different score distributions.
AlloyDB's ai.hybrid_search() currently uses RRF as its supported fusion algorithm.
AlloyDB's ai.hybrid_search() Function
The new function provides a database-level API for defining multiple search components.
A simplified call looks like this:
SELECT *
FROM ai.hybrid_search(
search_inputs => ARRAY[
'{
"data_type": "vector",
"weight": 0.5,
"table_name": "documents",
"key_column": "doc_id",
"vec_column": "text_embedding",
"distance_operator": "<=>",
"limit": 10,
"query_vector": "..."
}'::JSONB,
'{
"data_type": "text",
"weight": 0.5,
"table_name": "documents",
"key_column": "doc_id",
"text_column": "content",
"query_text_input": "database performance",
"limit": 10
}'::JSONB
]
);The exact embedding generation and text-index configuration depend on the application.
The important part is the shape of the API.
Each search method becomes a search component.
AlloyDB executes those components, ranks their results, combines the rankings using RRF, and returns a unified result set.
Vector Search in AlloyDB
For the vector side, AlloyDB supports vector search using its PostgreSQL-based AI capabilities.
A document table might contain:
CREATE TABLE documents (
doc_id TEXT PRIMARY KEY,
content TEXT,
text_embedding vector
);A vector index can then be created for approximate nearest-neighbor search.
For example, AlloyDB supports ScaNN indexing:
CREATE INDEX documents_embedding_scann
ON documents
USING scann (text_embedding cosine);The exact vector dimensions and embedding model depend on the model used by the application.
The important requirement is that the query embedding and stored document embeddings must be compatible.
A simplified vector query looks like:
SELECT
doc_id,
content
FROM documents
ORDER BY text_embedding <=> $1
LIMIT 10;The <=> operator is used for cosine distance in the documented hybrid-search examples.
Full-Text Search Provides the Other Half
The text side can use PostgreSQL full-text search.
A common design is to maintain a generated tsvector column:
CREATE TABLE documents (
doc_id TEXT PRIMARY KEY,
content TEXT,
text_tsv tsvector
GENERATED ALWAYS AS (
to_tsvector('english', content)
) STORED,
text_embedding vector
);An index can then be created:
CREATE INDEX documents_text_gin
ON documents
USING gin (text_tsv);A traditional search might look like:
SELECT
doc_id,
content,
ts_rank(
text_tsv,
websearch_to_tsquery('english', $1)
) AS rank
FROM documents
WHERE text_tsv @@ websearch_to_tsquery('english', $1)
ORDER BY rank DESC
LIMIT 10;This gives the text-search component a ranked result set.
The hybrid layer can then combine that list with the vector results.
AlloyDB also supports RUM indexes and a native BM25 index for different full-text search and ranking requirements.
Why BM25 and RUM Matter
There is another useful part of the recent AlloyDB search improvements.
Traditional PostgreSQL full-text search commonly uses GIN indexes.
For more advanced search behavior, AlloyDB supports RUM indexes, which can improve phrase search and relevance ranking by storing positional information directly in the index.
AlloyDB also provides a native BM25 index for keyword relevance ranking.
BM25 is widely used by search systems because it considers factors such as term frequency and document length.
That gives an application several choices:
Full-Text Search
|
+---- GIN
|
+---- RUM
|
+---- BM25The appropriate choice depends on the workload.
A simple keyword search does not necessarily need BM25.
A search application where relevance ranking and phrase matching are important may benefit from the additional capabilities.
Weighted Hybrid Search
Not every application wants vector and text search to contribute equally.
Consider a developer documentation portal.
A query such as:
HttpClientFactorycontains a very specific technical identifier.
Exact text matching may deserve more weight.
A natural-language query such as:
How should I reuse HTTP connections in a .NET service?may benefit more from semantic search.
AlloyDB's hybrid-search parameters allow each search component to specify a weight. The documented API supports weights between 0.0 and 1.0, with equal distribution when weights are not specified.
Conceptually:
Exact technical query
|
+---- Text: 0.7
+---- Vector: 0.3
Natural-language query
|
+---- Text: 0.4
+---- Vector: 0.6The values should not be chosen arbitrarily.
They should come from search-quality evaluation.
Filtering Can Be Applied to Search Components
Hybrid search does not mean searching every row in the database.
Applications commonly have filters such as:
tenant_id
language
product
category
visibility
publication_status
created_atAlloyDB's hybrid-search parameters support filter conditions for individual search components.
For example:
"filter_condition": "tenant_id = 'tenant-42'"The filter should be treated as part of the security and correctness model, not merely as a performance optimization.
For a multi-tenant application, this distinction is critical.
The vector search, text search, and final result retrieval must all remain within the correct tenant boundary.
Hybrid Search for RAG
This feature fits naturally into Retrieval-Augmented Generation systems.
A typical RAG pipeline looks like:
User Question
|
v
Query Processing
|
+-------------------+
| |
v v
Vector Search Full-Text Search
| |
+---------+---------+
|
v
RRF Re-ranking
|
v
Top Documents
|
v
LLM
|
v
AnswerThe retrieval step is often where RAG quality is won or lost.
If the retriever misses the relevant document, the model cannot reliably use that information later.
Hybrid search helps because it combines two different retrieval signals.
For example, a developer might ask:
How do I configure the retry policy for HttpClientFactory?Semantic search may find documents discussing HTTP resilience and retry strategies.
Full-text search can specifically identify documents containing:
HttpClientFactory
retry
PollyThe combined ranking can surface documents that satisfy both signals.
Keeping Search Inside the Database
One of the strongest architectural benefits is reducing the number of systems involved.
A traditional RAG architecture might look like:
PostgreSQL
|
+---- Vector Search
|
v
Search Service
|
+---- Keyword Search
|
v
Application
|
+---- Score Fusion
|
v
LLMWith AlloyDB hybrid search:
AlloyDB
|
+---------+---------+
| |
Vector Search Full-Text Search
| |
+---------+---------+
|
RRF
|
v
Search Results
|
v
LLMThat can simplify operations because the application does not need to maintain separate ranking infrastructure merely to combine two search result sets.
Google Cloud describes the native function as executing the hybrid workflow in a single database query plan, reducing application-side orchestration and joins.
External Search Engines Can Also Participate
The approach is not limited to data stored entirely inside AlloyDB.
AlloyDB now documents support for combining AlloyDB vector search with token-search results from external systems including:
OpenSearch
Elasticsearch
SolrThis is exposed through the external search Foreign Data Wrapper integration.
That creates another possible architecture:
AlloyDB
|
ai.hybrid_search()
|
+---------+---------+
| |
v v
AlloyDB Vector External Search
Search OpenSearch
|
v
Keyword SearchThis can be useful for organizations that already have a dedicated search engine but want to combine its lexical results with vector retrieval in AlloyDB.
It also means the feature is not necessarily a reason to immediately remove an existing search system.
A Practical RAG Table
A reasonable document table might look like:
CREATE TABLE knowledge_documents (
id UUID PRIMARY KEY,
tenant_id TEXT NOT NULL,
title TEXT NOT NULL,
content TEXT NOT NULL,
content_tsv tsvector
GENERATED ALWAYS AS (
to_tsvector('english', content)
) STORED,
embedding vector,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);Indexes might include:
CREATE INDEX knowledge_documents_tsv_idx
ON knowledge_documents
USING gin (content_tsv);
CREATE INDEX knowledge_documents_embedding_idx
ON knowledge_documents
USING scann (embedding cosine);The application can then use hybrid retrieval rather than maintaining separate ranking code.
The important design decision is not the SQL syntax.
It is deciding which fields should be searchable, which filters should always apply, how embeddings are generated, and how the final search quality will be evaluated.
Common Mistakes
Treating RRF as a Magic Ranking Algorithm
RRF is useful because it avoids direct score normalization.
It does not automatically understand which result is better for your users.
Search quality still needs evaluation.
Giving Vector and Text Search Arbitrary Weights
A 50/50 split is a reasonable starting point, but not a universal answer.
Test real queries.
Measure whether users find the correct documents near the top of the result list.
Forgetting Exact Identifiers
Semantic search is not always the best tool for identifiers such as:
API names
Error codes
Product SKUs
Ticket IDs
Class names
Version numbersFull-text search can be much better for these cases.
Ignoring Tenant Filters
A hybrid search query still needs normal authorization boundaries.
Never assume that combining two search methods automatically preserves tenant isolation.
Using Vector Search for Everything
Not every query needs embeddings.
If someone searches for:
ERR_CONNECTION_RESETthere may be little value in turning that exact identifier into a semantic retrieval problem.
Use the search technique that matches the query.
Advantages and Disadvantages
Advantages
Hybrid ranking becomes much easier to implement. The ai.hybrid_search() function combines multiple ranked search components using RRF rather than forcing developers to build the fusion logic themselves.
It avoids brittle score normalization. RRF operates on result rank, so vector distances and text relevance scores do not need to be converted to a common numerical scale.
Search can remain close to the application data. Teams can combine vector and full-text retrieval inside AlloyDB instead of automatically introducing another search service.
The approach works well for RAG. Semantic retrieval can provide conceptual matches while full-text search preserves exact keyword relevance.
External search systems can participate. Existing OpenSearch, Elasticsearch, or Solr deployments can be incorporated into hybrid retrieval through AlloyDB's external search integration.
Disadvantages
Search quality still requires tuning. RRF makes result fusion easier, but it does not guarantee that the resulting ranking matches user expectations.
More indexes mean more operational considerations. Vector indexes, full-text indexes, RUM indexes, or BM25 indexes all have different performance and maintenance characteristics.
Hybrid search can increase query complexity. Even with a built-in function, developers still need to understand filtering, ranking, embeddings, and search limits.
Not every query benefits from hybrid retrieval. Exact identifiers and simple keyword lookups may be better handled by conventional text search.
When Hybrid Search Makes Sense
AlloyDB hybrid search is particularly useful for applications such as:
RAG applications
Enterprise document search
Developer documentation
Ecommerce search
Knowledge bases
AI assistants
Customer-support search
Product discovery
Agent memoryIt is especially useful when users mix exact terms with natural-language descriptions.
For example:
"How do I configure PostgreSQL connection pooling?"contains both semantic intent and specific technical terminology.
A hybrid retriever can use both signals instead of forcing the application to choose one.
Summary
AlloyDB's native hybrid search capabilities make it easier to combine vector search and full-text search without moving the ranking logic into application code.
The central feature is ai.hybrid_search(), which accepts multiple search components and combines their ranked results using Reciprocal Rank Fusion.
That matters because vector similarity scores and keyword relevance scores are fundamentally different. Trying to normalize them manually often creates fragile ranking logic. RRF avoids that problem by working with result positions instead.
For RAG systems, the approach is particularly useful. Vector search can capture semantic meaning while full-text search protects exact terminology, identifiers, and keyword matches.
The feature also fits well with an existing PostgreSQL architecture. Teams can keep documents, metadata, embeddings, filters, and retrieval logic close to the same database, while organizations that already use OpenSearch, Elasticsearch, or Solr can also incorporate external search results into the hybrid workflow.
The important engineering lesson is not to treat hybrid search as a checkbox.
Good search still requires good indexing, sensible filters, appropriate weights, useful embeddings, and real evaluation against user queries.
RRF simply makes the difficult part of combining those search signals much easier to build and maintain.

Join the conversation! Your thoughts help the community grow.