Vector search is becoming a normal part of application architecture.

A traditional SQL query looks for exact values. Vector search looks for data that is semantically similar to a query. That makes it useful for RAG applications, recommendation systems, document search, semantic product discovery, and AI-powered applications.

Azure SQL Database now supports native vector search and DiskANN-based vector indexing. The latest vector index capabilities include full DML support, iterative filtering, optimizer-driven search behavior, and improved storage efficiency through quantization.

That changes the conversation for teams already storing application data in SQL.

Instead of introducing a separate vector database just because an application needs similarity search, developers can evaluate whether Azure SQL Database can keep the relational data and vector data together.

But DiskANN is not simply a faster SELECT.

It changes how vector search is indexed and how developers need to think about recall, latency, filtering, index maintenance, and production workloads.

What Is DiskANN?

DiskANN is an approximate nearest neighbor, or ANN, indexing algorithm designed for efficient vector similarity search.

Suppose an application stores thousands or millions of document embeddings.

A query produces another vector:

User Question
      |
      v
Embedding Model
      |
      v
Query Vector
      |
      v
Vector Search
      |
      v
Nearest Documents

The database needs to determine which stored vectors are closest to the query vector.

One approach is exhaustive nearest-neighbor search.

Query Vector
     |
     +--> Vector 1
     +--> Vector 2
     +--> Vector 3
     +--> ...
     +--> Vector 1,000,000

Every vector may need to be considered.

That can become expensive as the dataset grows.

DiskANN uses an approximate nearest neighbor approach to reduce the amount of work required to find relevant vectors.

The result is a trade-off:

Exact Search
Accuracy:      Highest
Search Work:   Higher

ANN Search
Accuracy:      Approximate
Search Work:   Lower

Azure's architecture guidance lists DiskANN alongside other ANN approaches such as HNSW and IVFFlat, while Azure SQL Database uses DiskANN for vector indexing.

Why Does Vector Indexing Matter?

Without an appropriate vector index, a similarity search may need to compare the query against a large portion of the stored vector dataset.

Imagine a table containing:

10,000 documents
100,000 documents
1,000,000 documents
10,000,000 documents

As the number of vectors grows, exhaustive comparison becomes increasingly difficult to use for low-latency application requests.

An ANN index changes the search path:

Without Vector Index

Query
  |
  v
Compare Against Many Vectors
  |
  v
Calculate Distances
  |
  v
Sort Results
  |
  v
Top K


With DiskANN

Query
  |
  v
DiskANN Index
  |
  v
Candidate Neighbors
  |
  v
Top K

The exact performance improvement depends on the dataset, query pattern, hardware, index configuration, filtering, and workload.

There is no universal latency number that should be assumed for every application.

Creating a DiskANN Vector Index

The current Azure SQL syntax uses CREATE VECTOR INDEX.

A basic example is:

CREATE VECTOR INDEX IX_Documents_Embedding
ON dbo.Documents (Embedding)
WITH (
    METRIC = 'COSINE',
    TYPE = 'DISKANN'
);

The vector column must use the SQL vector data type.

Azure SQL supports cosine, dot-product, and Euclidean distance metrics for vector indexes. DiskANN is currently the supported ANN index type in the CREATE VECTOR INDEX syntax.

A typical table might look like:

CREATE TABLE dbo.Documents
(
    Id BIGINT PRIMARY KEY,
    Title NVARCHAR(500),
    Content NVARCHAR(MAX),
    Embedding VECTOR(1536)
);

The 1536 value must match the dimensionality of the embeddings being stored.

That detail is important.

If the embedding model produces 1,536 dimensions, the database vector column needs to be defined accordingly.

Running Vector Search

Once the vector data and index exist, the application can search for similar vectors.

With the latest vector index syntax, approximate search uses TOP (N) WITH APPROXIMATE.

For example:

DECLARE @QueryVector VECTOR(1536) = '[...]';

SELECT TOP (5) WITH APPROXIMATE
    d.Id,
    d.Title,
    d.Content,
    s.distance
FROM VECTOR_SEARCH(
    TABLE = dbo.Documents AS d,
    COLUMN = Embedding,
    SIMILAR_TO = @QueryVector,
    METRIC = 'cosine'
) AS s
ORDER BY s.distance;

The current syntax is important because the older TOP_N parameter is deprecated for the latest vector index version.

That means new implementations should use the newer syntax rather than copying older examples found in existing projects.

What Does Approximate Mean?

Approximate does not mean random.

It means the search is optimized to find highly relevant nearest neighbors without necessarily evaluating every possible vector.

For many AI applications, this is a reasonable trade.

Consider a RAG system with one million document chunks.

A search that returns the exact mathematical top five documents may require significantly more work than an ANN search that quickly finds highly relevant candidates.

The engineering question becomes:

Is the small potential loss in recall
worth the reduction in search cost and latency?

For many production systems, the answer can be yes.

For applications where missing even one particular result has serious consequences, exhaustive search may still be appropriate for some workloads.

DiskANN vs Exact Search

Area

Exact Search

DiskANN

Search method

Exhaustive

Approximate nearest neighbor

Recall

Exact

Approximate

Large datasets

More expensive

Better suited

Query latency

Can increase with data size

Designed for efficient search

Index required

No ANN index

DiskANN index

Operational complexity

Lower

Higher

Best use

Smaller datasets or exact requirements

Larger semantic-search workloads

The right choice depends on the application.

Do not add DiskANN simply because the application uses embeddings.

First measure whether the existing search method is actually a bottleneck.

DiskANN and RAG Applications

One of the most obvious use cases is Retrieval-Augmented Generation.

A typical RAG pipeline looks like this:

User Question
      |
      v
Embedding Model
      |
      v
Query Vector
      |
      v
Azure SQL DiskANN
      |
      v
Relevant Documents
      |
      v
Prompt Context
      |
      v
Large Language Model
      |
      v
Answer

The database can store both application metadata and embeddings.

For example:

CREATE TABLE dbo.KnowledgeChunks
(
    Id BIGINT PRIMARY KEY,
    DocumentId BIGINT NOT NULL,
    TenantId INT NOT NULL,
    Title NVARCHAR(500),
    Content NVARCHAR(MAX),
    Embedding VECTOR(1536)
);

This is particularly interesting for applications that already depend heavily on relational SQL data.

Instead of maintaining:

Application Database
       +
Vector Database
       +
Synchronization Pipeline

the architecture may be able to keep related data together:

Azure SQL Database
       |
       +-- Business Data
       +-- Metadata
       +-- Document Chunks
       +-- Embeddings
       +-- Vector Index

That can simplify some application architectures.

Filtering Is the Real Production Challenge

A basic vector search is easy.

Production search is usually more complicated.

Imagine a multi-tenant application.

A document might contain:

{
    "documentId": 1024,
    "tenantId": 42,
    "category": "finance",
    "language": "en"
}

The application should not retrieve semantically similar documents belonging to another tenant.

The desired search is therefore:

Semantic Similarity
        +
Tenant Filter
        +
Business Filters

This is where the latest Azure SQL vector indexing improvements become particularly important.

The latest vector indexes support iterative filtering, where predicates can be applied during the vector search process rather than simply retrieving candidates and filtering them afterward.

Conceptually:

Query Vector
     |
     v
DiskANN Search
     |
     +---- Tenant = 42
     +---- Category = finance
     +---- Language = en
     |
     v
Relevant Candidates

This can be much more useful for real applications than a vector search that ignores business constraints.

Multi-Tenant Search Example

Consider:

CREATE TABLE dbo.Documents
(
    Id BIGINT PRIMARY KEY,
    TenantId INT NOT NULL,
    Category NVARCHAR(100),
    Content NVARCHAR(MAX),
    Embedding VECTOR(1536)
);

The application may want:

SELECT TOP (10) WITH APPROXIMATE
    d.Id,
    d.Content,
    s.distance
FROM VECTOR_SEARCH(
    TABLE = dbo.Documents AS d,
    COLUMN = Embedding,
    SIMILAR_TO = @QueryVector,
    METRIC = 'cosine'
) AS s
WHERE d.TenantId = @TenantId
  AND d.Category = @Category
ORDER BY s.distance;

The important point is not just the SQL syntax.

The database needs to handle the interaction between vector similarity and relational filtering efficiently.

This is one of the areas where iterative filtering can matter.

Full DML Changes the Architecture

Earlier generations of vector indexing had limitations around modifying tables after vector indexes were created.

The latest vector index capabilities support full DML, including:

INSERT
UPDATE
DELETE
MERGE

while maintaining the vector index.

That matters for applications where embeddings change continuously.

For example:

New Document
     |
     v
Generate Embedding
     |
     v
INSERT
     |
     v
Vector Index Maintained

An updated document can follow:

Document Changed
     |
     v
Generate New Embedding
     |
     v
UPDATE
     |
     v
Vector Index Maintained

This is a much more practical model for applications with continuously changing data.

Optimizer-Driven Search

Another important improvement is that the query optimizer can determine whether to use the DiskANN index or a k-nearest-neighbor search based on query characteristics.

That means developers do not necessarily need to assume:

Vector Query = Always DiskANN

Instead, the database can choose an appropriate execution strategy.

This is useful because an ANN index is not automatically the best solution for every query.

For example, if a filter reduces the candidate set to a small number of rows, an exhaustive search over that smaller set may make more sense.

The database engine can make that decision based on the query.

Vector Index Storage and Quantization

Vector data can consume substantial storage.

Suppose an application stores:

1,000,000 vectors
1,536 dimensions

The raw vector representation alone can become significant before accounting for indexes, metadata, and other columns.

The latest vector indexing capabilities include advanced quantization techniques intended to improve storage efficiency and query performance.

This is important for production because vector search cost is not only about CPU.

You also need to consider:

A vector architecture should therefore be evaluated as a database workload, not just an AI feature.

Choosing the Distance Metric

Azure SQL vector indexes support three distance metrics:

The metric should match how the embedding model and application interpret similarity.

For example:

WITH (
    METRIC = 'COSINE',
    TYPE = 'DISKANN'
)

Cosine similarity is commonly used for semantic text embeddings.

But do not select a metric simply because it appears frequently in examples.

The application should use the metric appropriate for its embedding model and retrieval strategy.

DiskANN and Traditional SQL Indexes

One of the strongest aspects of Azure SQL is that vector search does not have to exist separately from normal SQL querying.

An application can combine vector retrieval with traditional relational filtering.

For example:

                 Azure SQL
                    |
        +-----------+-----------+
        |                       |
        v                       v
Traditional Indexes        Vector Index
        |                       |
        +-----------+-----------+
                    |
                    v
             Application Query

A query might combine:

Tenant
Status
Date
Category
Vector Similarity

This is useful for enterprise applications where semantic similarity is only one part of the query.

When Should You Use DiskANN?

DiskANN is a strong candidate when:

It may not be necessary when:

Measure before migrating.

Common Mistakes

Treating ANN as Exact Search

DiskANN is approximate.

Do not assume that it will always return the exact same top-K results as exhaustive search.

Ignoring Filters

A vector search that works in a benchmark may behave differently when tenant, authorization, category, or date filters are introduced.

Choosing the Wrong Embedding Dimension

The vector column must match the embedding model's output dimensions.

Using Old Syntax

Older examples may use TOP_N.

For current vector indexes, use:

SELECT TOP (N) WITH APPROXIMATE

The older TOP_N parameter is deprecated for the latest vector index version.

Measuring Only Latency

A vector search benchmark should also evaluate recall, throughput, resource consumption, and the quality of retrieved results.

Creating an Index Without Testing the Workload

Indexing decisions should be based on the actual query patterns and dataset.

Troubleshooting Vector Search

Results Are Not Relevant

Check:

  1. Embedding model.

  2. Embedding dimensions.

  3. Distance metric.

  4. Chunking strategy.

  5. Query construction.

  6. ANN search behavior.

  7. Metadata filters.

The vector index may not be the actual problem.

Query Is Fast but Retrieval Quality Is Poor

Increasing database performance does not automatically improve semantic relevance.

Evaluate recall against an exact-search baseline.

A useful test is:

Test Dataset
     |
     +--> Exact Search
     |
     +--> DiskANN Search
     |
     v
Compare Top-K Results

This tells you how much retrieval quality changes when moving from exact to approximate search.

Query Is Slow With Filters

Investigate the filter selectivity and the query plan.

Do not assume that adding a vector index automatically solves every filtered search problem.

Index Build Consumes Too Many Resources

Vector index creation is a database operation.

The CREATE VECTOR INDEX syntax supports MAXDOP, which can be used to control parallelism and resulting resource consumption during index creation.

For example:

CREATE VECTOR INDEX IX_Documents_Embedding
ON dbo.Documents (Embedding)
WITH (
    METRIC = 'COSINE',
    TYPE = 'DISKANN',
    MAXDOP = 4
);

The appropriate value depends on the database workload and available resources.

Best Practices

  1. Benchmark exact search before introducing ANN indexing.

  2. Measure recall as well as latency.

  3. Choose the distance metric deliberately.

  4. Keep embedding dimensions consistent.

  5. Test vector search with realistic metadata filters.

  6. Use the latest vector index syntax for new implementations.

  7. Evaluate index storage and maintenance costs.

  8. Test INSERT, UPDATE, and DELETE workloads, not just read queries.

  9. Use tenant and authorization filters carefully in multi-tenant systems.

  10. Compare DiskANN results against an exact-search baseline during evaluation.

  11. Tune index-building resources rather than allowing large builds to compete blindly with application traffic.

  12. Monitor the complete database workload after enabling vector search.

DiskANN vs a Separate Vector Database

One of the biggest architectural questions is whether to keep vectors in Azure SQL or introduce a dedicated vector database.

There is no universal answer.

Requirement

Azure SQL + DiskANN

Dedicated Vector Store

Existing SQL application

Strong fit

Additional integration

Relational filtering

Strong fit

Depends on product

Existing transactional data

Natural fit

Requires synchronization

Specialized vector workloads

May fit

Often strong fit

Operational simplicity

Potential advantage

Additional system

Large-scale vector workload

Requires benchmarking

Depends on platform

Existing Azure SQL expertise

Strong advantage

New technology surface

The correct decision depends on workload characteristics rather than the popularity of a particular database.

If the application already has customer, product, security, and transactional information in Azure SQL, keeping embeddings close to that data can simplify some architectures.

If vector retrieval becomes an extremely specialized workload with requirements that SQL cannot satisfy efficiently, a dedicated system may still be appropriate.

Final Thoughts

DiskANN makes Azure SQL more interesting for production AI workloads because vector search no longer has to be treated as an isolated capability.

The database can store application data, metadata, embeddings, and vector indexes together while allowing vector search to participate in normal SQL workloads.

The latest vector indexing capabilities are especially important because they address several problems that matter in production: full DML support, iterative filtering, optimizer-driven execution, and improved storage efficiency.

But DiskANN should not be treated as a magic performance switch.

Approximate search introduces a recall trade-off. Indexes require storage and maintenance. Filtering changes query behavior. Embedding quality affects retrieval quality. And the best configuration depends heavily on the actual dataset and workload.

For teams already using Azure SQL Database, the practical approach is to benchmark the workload first, compare exact and approximate retrieval, test realistic filters, and then decide whether DiskANN provides enough value to justify the additional indexing and operational considerations.