Vector search is becoming a normal part of application architecture.
A traditional SQL query looks for exact values. Vector search looks for data that is semantically similar to a query. That makes it useful for RAG applications, recommendation systems, document search, semantic product discovery, and AI-powered applications.
Azure SQL Database now supports native vector search and DiskANN-based vector indexing. The latest vector index capabilities include full DML support, iterative filtering, optimizer-driven search behavior, and improved storage efficiency through quantization.
That changes the conversation for teams already storing application data in SQL.
Instead of introducing a separate vector database just because an application needs similarity search, developers can evaluate whether Azure SQL Database can keep the relational data and vector data together.
But DiskANN is not simply a faster SELECT.
It changes how vector search is indexed and how developers need to think about recall, latency, filtering, index maintenance, and production workloads.
What Is DiskANN?
DiskANN is an approximate nearest neighbor, or ANN, indexing algorithm designed for efficient vector similarity search.
Suppose an application stores thousands or millions of document embeddings.
A query produces another vector:
User Question
|
v
Embedding Model
|
v
Query Vector
|
v
Vector Search
|
v
Nearest DocumentsThe database needs to determine which stored vectors are closest to the query vector.
One approach is exhaustive nearest-neighbor search.
Query Vector
|
+--> Vector 1
+--> Vector 2
+--> Vector 3
+--> ...
+--> Vector 1,000,000Every vector may need to be considered.
That can become expensive as the dataset grows.
DiskANN uses an approximate nearest neighbor approach to reduce the amount of work required to find relevant vectors.
The result is a trade-off:
Exact Search
Accuracy: Highest
Search Work: Higher
ANN Search
Accuracy: Approximate
Search Work: LowerAzure's architecture guidance lists DiskANN alongside other ANN approaches such as HNSW and IVFFlat, while Azure SQL Database uses DiskANN for vector indexing.
Why Does Vector Indexing Matter?
Without an appropriate vector index, a similarity search may need to compare the query against a large portion of the stored vector dataset.
Imagine a table containing:
10,000 documents
100,000 documents
1,000,000 documents
10,000,000 documentsAs the number of vectors grows, exhaustive comparison becomes increasingly difficult to use for low-latency application requests.
An ANN index changes the search path:
Without Vector Index
Query
|
v
Compare Against Many Vectors
|
v
Calculate Distances
|
v
Sort Results
|
v
Top K
With DiskANN
Query
|
v
DiskANN Index
|
v
Candidate Neighbors
|
v
Top KThe exact performance improvement depends on the dataset, query pattern, hardware, index configuration, filtering, and workload.
There is no universal latency number that should be assumed for every application.
Creating a DiskANN Vector Index
The current Azure SQL syntax uses CREATE VECTOR INDEX.
A basic example is:
CREATE VECTOR INDEX IX_Documents_Embedding
ON dbo.Documents (Embedding)
WITH (
METRIC = 'COSINE',
TYPE = 'DISKANN'
);The vector column must use the SQL vector data type.
Azure SQL supports cosine, dot-product, and Euclidean distance metrics for vector indexes. DiskANN is currently the supported ANN index type in the CREATE VECTOR INDEX syntax.
A typical table might look like:
CREATE TABLE dbo.Documents
(
Id BIGINT PRIMARY KEY,
Title NVARCHAR(500),
Content NVARCHAR(MAX),
Embedding VECTOR(1536)
);The 1536 value must match the dimensionality of the embeddings being stored.
That detail is important.
If the embedding model produces 1,536 dimensions, the database vector column needs to be defined accordingly.
Running Vector Search
Once the vector data and index exist, the application can search for similar vectors.
With the latest vector index syntax, approximate search uses TOP (N) WITH APPROXIMATE.
For example:
DECLARE @QueryVector VECTOR(1536) = '[...]';
SELECT TOP (5) WITH APPROXIMATE
d.Id,
d.Title,
d.Content,
s.distance
FROM VECTOR_SEARCH(
TABLE = dbo.Documents AS d,
COLUMN = Embedding,
SIMILAR_TO = @QueryVector,
METRIC = 'cosine'
) AS s
ORDER BY s.distance;The current syntax is important because the older TOP_N parameter is deprecated for the latest vector index version.
That means new implementations should use the newer syntax rather than copying older examples found in existing projects.
What Does Approximate Mean?
Approximate does not mean random.
It means the search is optimized to find highly relevant nearest neighbors without necessarily evaluating every possible vector.
For many AI applications, this is a reasonable trade.
Consider a RAG system with one million document chunks.
A search that returns the exact mathematical top five documents may require significantly more work than an ANN search that quickly finds highly relevant candidates.
The engineering question becomes:
Is the small potential loss in recall
worth the reduction in search cost and latency?For many production systems, the answer can be yes.
For applications where missing even one particular result has serious consequences, exhaustive search may still be appropriate for some workloads.
DiskANN vs Exact Search
Area | Exact Search | DiskANN |
|---|---|---|
Search method | Exhaustive | Approximate nearest neighbor |
Recall | Exact | Approximate |
Large datasets | More expensive | Better suited |
Query latency | Can increase with data size | Designed for efficient search |
Index required | No ANN index | DiskANN index |
Operational complexity | Lower | Higher |
Best use | Smaller datasets or exact requirements | Larger semantic-search workloads |
The right choice depends on the application.
Do not add DiskANN simply because the application uses embeddings.
First measure whether the existing search method is actually a bottleneck.
DiskANN and RAG Applications
One of the most obvious use cases is Retrieval-Augmented Generation.
A typical RAG pipeline looks like this:
User Question
|
v
Embedding Model
|
v
Query Vector
|
v
Azure SQL DiskANN
|
v
Relevant Documents
|
v
Prompt Context
|
v
Large Language Model
|
v
AnswerThe database can store both application metadata and embeddings.
For example:
CREATE TABLE dbo.KnowledgeChunks
(
Id BIGINT PRIMARY KEY,
DocumentId BIGINT NOT NULL,
TenantId INT NOT NULL,
Title NVARCHAR(500),
Content NVARCHAR(MAX),
Embedding VECTOR(1536)
);This is particularly interesting for applications that already depend heavily on relational SQL data.
Instead of maintaining:
Application Database
+
Vector Database
+
Synchronization Pipelinethe architecture may be able to keep related data together:
Azure SQL Database
|
+-- Business Data
+-- Metadata
+-- Document Chunks
+-- Embeddings
+-- Vector IndexThat can simplify some application architectures.
Filtering Is the Real Production Challenge
A basic vector search is easy.
Production search is usually more complicated.
Imagine a multi-tenant application.
A document might contain:
{
"documentId": 1024,
"tenantId": 42,
"category": "finance",
"language": "en"
}The application should not retrieve semantically similar documents belonging to another tenant.
The desired search is therefore:
Semantic Similarity
+
Tenant Filter
+
Business FiltersThis is where the latest Azure SQL vector indexing improvements become particularly important.
The latest vector indexes support iterative filtering, where predicates can be applied during the vector search process rather than simply retrieving candidates and filtering them afterward.
Conceptually:
Query Vector
|
v
DiskANN Search
|
+---- Tenant = 42
+---- Category = finance
+---- Language = en
|
v
Relevant CandidatesThis can be much more useful for real applications than a vector search that ignores business constraints.
Multi-Tenant Search Example
Consider:
CREATE TABLE dbo.Documents
(
Id BIGINT PRIMARY KEY,
TenantId INT NOT NULL,
Category NVARCHAR(100),
Content NVARCHAR(MAX),
Embedding VECTOR(1536)
);The application may want:
SELECT TOP (10) WITH APPROXIMATE
d.Id,
d.Content,
s.distance
FROM VECTOR_SEARCH(
TABLE = dbo.Documents AS d,
COLUMN = Embedding,
SIMILAR_TO = @QueryVector,
METRIC = 'cosine'
) AS s
WHERE d.TenantId = @TenantId
AND d.Category = @Category
ORDER BY s.distance;The important point is not just the SQL syntax.
The database needs to handle the interaction between vector similarity and relational filtering efficiently.
This is one of the areas where iterative filtering can matter.
Full DML Changes the Architecture
Earlier generations of vector indexing had limitations around modifying tables after vector indexes were created.
The latest vector index capabilities support full DML, including:
INSERT
UPDATE
DELETE
MERGEwhile maintaining the vector index.
That matters for applications where embeddings change continuously.
For example:
New Document
|
v
Generate Embedding
|
v
INSERT
|
v
Vector Index MaintainedAn updated document can follow:
Document Changed
|
v
Generate New Embedding
|
v
UPDATE
|
v
Vector Index MaintainedThis is a much more practical model for applications with continuously changing data.
Optimizer-Driven Search
Another important improvement is that the query optimizer can determine whether to use the DiskANN index or a k-nearest-neighbor search based on query characteristics.
That means developers do not necessarily need to assume:
Vector Query = Always DiskANNInstead, the database can choose an appropriate execution strategy.
This is useful because an ANN index is not automatically the best solution for every query.
For example, if a filter reduces the candidate set to a small number of rows, an exhaustive search over that smaller set may make more sense.
The database engine can make that decision based on the query.
Vector Index Storage and Quantization
Vector data can consume substantial storage.
Suppose an application stores:
1,000,000 vectors
1,536 dimensionsThe raw vector representation alone can become significant before accounting for indexes, metadata, and other columns.
The latest vector indexing capabilities include advanced quantization techniques intended to improve storage efficiency and query performance.
This is important for production because vector search cost is not only about CPU.
You also need to consider:
Storage
Memory
Index build time
Index maintenance
Query throughput
Database compute
Backup and recovery footprint
A vector architecture should therefore be evaluated as a database workload, not just an AI feature.
Choosing the Distance Metric
Azure SQL vector indexes support three distance metrics:
Cosine
Dot product
Euclidean
The metric should match how the embedding model and application interpret similarity.
For example:
WITH (
METRIC = 'COSINE',
TYPE = 'DISKANN'
)Cosine similarity is commonly used for semantic text embeddings.
But do not select a metric simply because it appears frequently in examples.
The application should use the metric appropriate for its embedding model and retrieval strategy.
DiskANN and Traditional SQL Indexes
One of the strongest aspects of Azure SQL is that vector search does not have to exist separately from normal SQL querying.
An application can combine vector retrieval with traditional relational filtering.
For example:
Azure SQL
|
+-----------+-----------+
| |
v v
Traditional Indexes Vector Index
| |
+-----------+-----------+
|
v
Application QueryA query might combine:
Tenant
Status
Date
Category
Vector SimilarityThis is useful for enterprise applications where semantic similarity is only one part of the query.
When Should You Use DiskANN?
DiskANN is a strong candidate when:
The vector dataset is large.
Search latency matters.
Approximate nearest-neighbor results are acceptable.
The application already uses Azure SQL Database.
Vector search needs relational filtering.
Embeddings are continuously inserted or updated.
The team wants to avoid introducing another specialized data store.
It may not be necessary when:
The dataset is very small.
Exact nearest-neighbor results are required.
Vector search is occasional rather than latency-sensitive.
The existing database workload already meets performance requirements.
Measure before migrating.
Common Mistakes
Treating ANN as Exact Search
DiskANN is approximate.
Do not assume that it will always return the exact same top-K results as exhaustive search.
Ignoring Filters
A vector search that works in a benchmark may behave differently when tenant, authorization, category, or date filters are introduced.
Choosing the Wrong Embedding Dimension
The vector column must match the embedding model's output dimensions.
Using Old Syntax
Older examples may use TOP_N.
For current vector indexes, use:
SELECT TOP (N) WITH APPROXIMATEThe older TOP_N parameter is deprecated for the latest vector index version.
Measuring Only Latency
A vector search benchmark should also evaluate recall, throughput, resource consumption, and the quality of retrieved results.
Creating an Index Without Testing the Workload
Indexing decisions should be based on the actual query patterns and dataset.
Troubleshooting Vector Search
Results Are Not Relevant
Check:
Embedding model.
Embedding dimensions.
Distance metric.
Chunking strategy.
Query construction.
ANN search behavior.
Metadata filters.
The vector index may not be the actual problem.
Query Is Fast but Retrieval Quality Is Poor
Increasing database performance does not automatically improve semantic relevance.
Evaluate recall against an exact-search baseline.
A useful test is:
Test Dataset
|
+--> Exact Search
|
+--> DiskANN Search
|
v
Compare Top-K ResultsThis tells you how much retrieval quality changes when moving from exact to approximate search.
Query Is Slow With Filters
Investigate the filter selectivity and the query plan.
Do not assume that adding a vector index automatically solves every filtered search problem.
Index Build Consumes Too Many Resources
Vector index creation is a database operation.
The CREATE VECTOR INDEX syntax supports MAXDOP, which can be used to control parallelism and resulting resource consumption during index creation.
For example:
CREATE VECTOR INDEX IX_Documents_Embedding
ON dbo.Documents (Embedding)
WITH (
METRIC = 'COSINE',
TYPE = 'DISKANN',
MAXDOP = 4
);The appropriate value depends on the database workload and available resources.
Best Practices
Benchmark exact search before introducing ANN indexing.
Measure recall as well as latency.
Choose the distance metric deliberately.
Keep embedding dimensions consistent.
Test vector search with realistic metadata filters.
Use the latest vector index syntax for new implementations.
Evaluate index storage and maintenance costs.
Test INSERT, UPDATE, and DELETE workloads, not just read queries.
Use tenant and authorization filters carefully in multi-tenant systems.
Compare DiskANN results against an exact-search baseline during evaluation.
Tune index-building resources rather than allowing large builds to compete blindly with application traffic.
Monitor the complete database workload after enabling vector search.
DiskANN vs a Separate Vector Database
One of the biggest architectural questions is whether to keep vectors in Azure SQL or introduce a dedicated vector database.
There is no universal answer.
Requirement | Azure SQL + DiskANN | Dedicated Vector Store |
|---|---|---|
Existing SQL application | Strong fit | Additional integration |
Relational filtering | Strong fit | Depends on product |
Existing transactional data | Natural fit | Requires synchronization |
Specialized vector workloads | May fit | Often strong fit |
Operational simplicity | Potential advantage | Additional system |
Large-scale vector workload | Requires benchmarking | Depends on platform |
Existing Azure SQL expertise | Strong advantage | New technology surface |
The correct decision depends on workload characteristics rather than the popularity of a particular database.
If the application already has customer, product, security, and transactional information in Azure SQL, keeping embeddings close to that data can simplify some architectures.
If vector retrieval becomes an extremely specialized workload with requirements that SQL cannot satisfy efficiently, a dedicated system may still be appropriate.
Final Thoughts
DiskANN makes Azure SQL more interesting for production AI workloads because vector search no longer has to be treated as an isolated capability.
The database can store application data, metadata, embeddings, and vector indexes together while allowing vector search to participate in normal SQL workloads.
The latest vector indexing capabilities are especially important because they address several problems that matter in production: full DML support, iterative filtering, optimizer-driven execution, and improved storage efficiency.
But DiskANN should not be treated as a magic performance switch.
Approximate search introduces a recall trade-off. Indexes require storage and maintenance. Filtering changes query behavior. Embedding quality affects retrieval quality. And the best configuration depends heavily on the actual dataset and workload.
For teams already using Azure SQL Database, the practical approach is to benchmark the workload first, compare exact and approximate retrieval, test realistic filters, and then decide whether DiskANN provides enough value to justify the additional indexing and operational considerations.

Join the conversation! Your thoughts help the community grow.