Search systems often need to identify documents that contain the right words, not just documents that are conceptually similar to a query. A customer searching for a product SKU, an error code, a database identifier, or a specific technical term expects those exact matches to appear near the top of the results.
Vector search is useful for understanding semantic similarity, but it can struggle with exact identifiers and uncommon keywords. Traditional PostgreSQL full-text search can identify matching terms, yet relevance ranking may need additional tuning to produce the desired order.
Google Cloud AlloyDB for PostgreSQL introduces native BM25 ranking through the pg_textsearch extension. BM25, short for Best Matching 25, is a probabilistic ranking algorithm that evaluates term frequency, inverse document frequency, and document length to estimate how relevant a document is to a query.
The feature is available in preview and lets developers perform relevance-ranked keyword searches inside AlloyDB without maintaining a separate search backend solely for BM25 ranking.
For developers building product search, documentation search, support portals, or retrieval-augmented generation (RAG) applications, this provides another option for combining PostgreSQL data with more sophisticated text relevance.
What BM25 Ranking Solves
A basic text search identifies documents containing relevant terms. Ranking determines which of those documents should appear first.
Imagine an application that stores technical documentation. A user searches for database connection timeout. The database might find three documents:
A troubleshooting guide specifically about database connection timeouts.
A general database administration guide that mentions connections several times.
A networking guide that mentions timeouts in a different context.
A useful search system should rank the troubleshooting guide above documents that only partially match the query.
BM25 improves keyword ranking by considering how terms appear in individual documents and across the document collection. It also accounts for document length so that longer documents do not automatically receive an advantage simply because they contain more words.
The main factors are:
Term frequency (TF): How often a query term appears in a document.
Inverse document frequency (IDF): How informative a term is based on how common it is across the collection.
Document length normalization: How the document's length affects its relevance score.
These factors work together to estimate the relevance of a document to a particular query.
BM25 is especially useful when users expect specific words, product identifiers, technical phrases, or error messages to influence ranking. It does not, however, understand the full meaning of a query in the same way a semantic embedding model can.
Why Use Native BM25 in AlloyDB?
Before native BM25 indexing, teams that needed this ranking approach alongside AlloyDB could introduce an external search engine or build additional search infrastructure.
That architecture can work well, especially when an application requires advanced search features, independent scaling, or specialized indexing capabilities. However, it also creates additional operational responsibilities.
Teams may need to synchronize records between the primary database and the search system, monitor indexing delays, troubleshoot inconsistent results, and maintain another service.
With native BM25 ranking, AlloyDB can handle keyword relevance scoring within the PostgreSQL-compatible database environment.
The practical advantages include:
Keeping application records and BM25 search indexes in the same database system.
Avoiding a separate search backend when BM25 is the primary additional requirement.
Using SQL to retrieve and rank documents.
Combining keyword ranking with existing structured filters.
Integrating keyword search with vector similarity search for hybrid retrieval.
Native BM25 does not automatically eliminate every reason to use a dedicated search engine. Applications that require specialized faceting, complex search analytics, extensive distributed indexing, or other search-platform capabilities should evaluate the complete workload before changing architecture.
Enable the pg_textsearch Extension
Before creating a BM25 index, verify that the AlloyDB instance meets the feature requirements.
The native BM25 documentation specifies support for AlloyDB instances running PostgreSQL major versions 17 or 18. The feature is in preview, so check its current availability and applicable service terms before adopting it for a production workload.
You also need the alloydbsuperuser database role to configure the extension.
Connect to the target AlloyDB database using psql or another PostgreSQL-compatible client and run:
CREATE EXTENSION IF NOT EXISTS pg_textsearch;The extension must be enabled in each database where you intend to use BM25 indexing.
The IF NOT EXISTS clause makes the statement safe to repeat when the extension is already installed. It does not guarantee that the extension is available or that the current database role has permission to create it.
If the command fails, verify the PostgreSQL major version, feature availability, database permissions, and extension support before proceeding.
Create a Table and Add Sample Documents
To understand the workflow, consider a table that stores technical documentation.
Each row contains a title and a body of text. The title helps identify the result, while the content column provides the searchable document.
CREATE TABLE documents (
id SERIAL PRIMARY KEY,
title TEXT NOT NULL,
content TEXT NOT NULL
);Insert a few sample records:
INSERT INTO documents (title, content)
VALUES
(
'Database Connection Troubleshooting',
'Troubleshoot database connection timeouts, '
'connection pools, and failed PostgreSQL sessions.'
),
(
'PostgreSQL Administration',
'Learn PostgreSQL administration, indexing, '
'query planning, and database maintenance.'
),
(
'Network Timeout Diagnostics',
'Investigate network timeouts, DNS resolution, '
'firewall rules, and connection failures.'
),
(
'BM25 Search Ranking',
'BM25 ranking evaluates term frequency, inverse '
'document frequency, and document length.'
);This dataset is intentionally small. It is useful for demonstrating query behavior, but meaningful relevance evaluation requires a larger collection with representative documents and realistic search queries.
Create a Native BM25 Index
Create the BM25 index on the text column you want to search.
CREATE INDEX idx_docs_bm25
ON documents
USING bm25 (content)
WITH (
text_config = 'english'
);The USING bm25 clause selects the BM25 index access method, while content specifies the indexed text column.
The text_config option defines the PostgreSQL text-search configuration used to process the text. In this example, english provides English-language tokenization and stemming behavior.
The index configuration also supports two optional parameters:
Parameter | Default | Purpose |
|---|---|---|
| Required | Controls text processing and normalization |
|
| Controls how quickly term-frequency contributions saturate |
|
| Controls document-length normalization |
These parameters influence ranking behavior, not the business meaning of a document. Their best values depend on the dataset and search requirements.
For example, a collection of short product descriptions may behave differently from a collection of long technical manuals. Test parameter changes against a representative query set rather than selecting values based on intuition alone.
Query Documents Using BM25 Ranking
Once the index is created, use the <@> operator to calculate BM25 relevance for a search query.
SELECT
id,
title,
content,
content <@> 'database connection timeout' AS score
FROM documents
ORDER BY content <@> 'database connection timeout' ASC
LIMIT 10;The query calculates a BM25 score for each document and sorts the results in ascending order.
One detail is particularly important: the <@> operator returns a negative BM25 score. A more negative score represents a stronger relevance match, so ascending order puts the most relevant results first.
This differs from many ranking APIs where larger positive scores indicate stronger relevance.
If you sort the results in descending order, you can reverse the intended ranking and place weaker matches first.
The query also returns the score so that you can inspect ranking behavior during development. In a production API, you might return only the fields required by the client while retaining the score for diagnostics.
Search across the appropriate document field
The example indexes only the content column. That is a reasonable starting point for a documentation search application, but it may not be enough for every workload.
For example, product search may need to consider product names, descriptions, and other searchable text. You should decide how those fields contribute to relevance rather than assuming that indexing a single description column will always produce the desired order.
You can also maintain a dedicated searchable text representation when the application requires a particular combination of fields. The representation and index design should match the query behavior you want to support.
Tune BM25 Parameters for Your Dataset
The default parameters provide a starting point, not a universal optimum.
The k1 parameter controls term-frequency saturation. Increasing it allows repeated occurrences of a query term to continue contributing to relevance over a wider range before the contribution saturates.
The b parameter controls document-length normalization. Increasing it generally makes the ranking more sensitive to differences in document length.
For example, create a second index with customized parameters:
CREATE INDEX idx_docs_bm25_tuned
ON documents
USING bm25 (content)
WITH (
text_config = 'english',
k1 = 1.5,
b = 0.8
);These values are illustrative tuning choices, not a recommendation that every application should use them.
Before changing parameters, define what a good result means for your users. A support portal might prioritize documents that directly address a specific error message. A product catalog might prioritize exact product terminology and identifiers. A documentation system might need a balance between specific API references and broader conceptual guides.
Create a small relevance test set with representative queries and expected results. Compare the ordering produced by the default and tuned indexes, and measure whether the changes improve the results that matter to your application.
Avoid creating multiple indexes with different parameter values unless you have a clear reason to maintain them. Additional indexes consume storage and can increase the cost of writes and index maintenance.
Combine BM25 With Vector Search
Keyword relevance and semantic similarity solve different retrieval problems.
BM25 is useful when the query contains important terms that should match the document text. Vector search is useful when the user expresses a concept in different words from those used in the source material.
For example, a user searching for connection timeout might benefit from a troubleshooting article that uses those exact terms. A user asking how to recover a service that cannot reach its database might benefit from a semantically related document even if the wording differs.
A hybrid search combines both result sets and merges their rankings. AlloyDB supports hybrid search workflows that can combine text search and vector similarity using Reciprocal Rank Fusion (RRF).
RRF combines the relative ranks of documents from different retrieval methods instead of directly comparing scores that use different scales.
Conceptually, the combined score can be represented as:
[
\operatorname{RRF}(d)=
\sum_{r \in R}
\frac{1}{k+\operatorname{rank}_{r}(d)}
]
Here, (d) is a document, (R) is the set of ranked result lists, (k) is a smoothing constant, and (\operatorname{rank}_{r}(d)) is the document's position in a result list.
This approach is useful because a BM25 score and a vector-distance score are not directly comparable. RRF combines their ranking positions rather than assuming that their numerical values have the same interpretation.
For applications that already store embeddings in AlloyDB, native BM25 can therefore become another retrieval method within a hybrid search architecture.
However, hybrid search introduces additional choices: how many candidates each method retrieves, how much influence each ranking source should have, and how the final result set is filtered. These choices should be evaluated against application-specific relevance tests.
Improve Search Quality With Structured Filters
BM25 ranking determines relevance among documents, but applications often need to restrict which documents are eligible for a search.
For example, a documentation platform might need to search only a specific product version. A customer-support application might restrict results to a tenant or organization. A product catalog might filter by availability or category.
Structured filters can be applied alongside relevance ranking, provided the query is designed to preserve the intended search behavior.
For example:
SELECT
id,
title,
content,
content <@> 'database connection timeout' AS score
FROM documents
WHERE id > 0
ORDER BY content <@> 'database connection timeout' ASC
LIMIT 10;The id > 0 predicate is only an illustrative filter. A real application should use the appropriate business or authorization constraints.
For multi-tenant applications, access restrictions must be enforced reliably. Search ranking is not an authorization mechanism, and filtering results after retrieval can create unnecessary work or expose data through mistakes in application logic.
Validate the query plan and execution behavior when combining BM25 ranking with selective filters on large tables. Search quality and query performance are separate concerns, and both need testing.
Common Implementation Mistakes
Sorting in the wrong direction: The BM25 <@> operator returns negative scores. Use ascending order to retrieve the strongest matches first.
Assuming BM25 understands semantic meaning: BM25 ranks lexical matches. It does not replace embeddings when a query needs semantic matching across different wording.
Using the wrong text-search configuration: The chosen configuration affects tokenization and stemming. Test language-specific behavior, especially when the corpus contains multiple languages, product identifiers, or code snippets.
Ignoring document length and content quality: Ranking cannot compensate for poorly structured or irrelevant content. Duplicate documents, boilerplate, and inconsistent text fields can affect search quality.
Changing parameters without evaluation: Adjusting k1 and b without a relevance test set can make results worse even if the query executes faster or produces different scores.
Assuming preview features have production guarantees: Review the current preview terms, supported versions, operational limitations, and rollout status before making the feature a critical dependency.
Skipping performance testing: Index creation, index storage, write overhead, and query execution costs depend on the workload. Test with realistic data volumes before committing to the design.
Summary
AlloyDB's native BM25 ranking provides PostgreSQL-compatible applications with a way to perform relevance-ranked full-text search using the pg_textsearch extension.
Developers can enable the extension, create a BM25 index with text_config, and query documents using the <@> operator. The operator returns negative scores, so sorting in ascending order places stronger matches first. The k1 and b parameters provide additional control over term-frequency saturation and document-length normalization.
Native BM25 is especially useful for applications where exact terms, identifiers, and keyword relevance matter. It can also complement vector search in hybrid retrieval workflows, combining lexical precision with semantic similarity.
Before adopting it, verify preview availability and supported PostgreSQL versions, measure relevance against representative queries, and test index maintenance and query performance with realistic data. The result is a more deliberate search architecture that keeps keyword ranking close to the data while preserving the option to combine it with other retrieval techniques.

Join the conversation! Your thoughts help the community grow.