A vector search request fails after an embedding model change. The new vector is shorter than the index expects, so the error is visible. A more difficult failure occurs when the replacement model produces the same number of values: the request may look valid even though the vectors no longer represent compatible embedding spaces.
Validate an embedding against an explicit index contract before storing it or using it as a query. Check its dimension, the embedding model or compatible encoder configuration, and the validity of its numeric values. For cosine similarity, reject an all-zero vector. A matching array length is only one part of compatibility.
This C# example makes those checks explicit using synthetic vectors. It does not call an embedding API or connect to a vector database.
What does the vector dimension describe?
An embedding is a numeric representation produced by an embedding model. Its dimension is the number of coordinates in the vector. It is different from the model's context window, the number of documents in an index, or the length of the original text.
Azure AI Search's index documentation defines a vector field's dimensions property in relation to the output of the embedding model. Its query documentation also recommends using the same embedding models for queries and indexed documents.
The precise schema and API differ across databases. The application-level requirement is stable: the vector presented to a particular index must satisfy that index's configured representation.
Why is a model identifier part of the contract?
Two models can produce arrays with identical lengths without assigning the same meaning to each coordinate. Length equality therefore does not establish that comparing their vectors is meaningful.
Use a stable model key that identifies the actual embedding configuration. Depending on the system, that may include a model revision, requested dimension, preprocessing policy, and an approved pairing of query and document encoders.
Some retrieval models use distinct query and document treatments. In that case, record the compatible pair and its intended roles; do not force identical preprocessing merely because the two paths share a dimension.
The model key in this example is a trusted application label. It is not a cryptographic proof of how a vector was produced. Attach provenance in the embedding service you control rather than accepting an arbitrary label supplied by an external caller.
Create a reproducible C# check
Use a .NET 8 console project with C# 12 and default implicit usings. The example uses only the base class library. It was compiled with the compiler bundled in .NET SDK 8.0.414 and executed with runtime 8.0.20 on Linux x64.
Create the project:
dotnet new console --name EmbeddingContractDemo --framework net8.0
cd EmbeddingContractDemo
Replace Program.cs with the following. The three-dimensional vectors and model names are synthetic fixtures, not the documented outputs of a commercial model.
var index = new EmbeddingContract("catalog-v1", "encoder-a@revision-1", 3);
var good = new Embedding("encoder-a@revision-1", new float[] { 0.2f, 0.4f, 0.8f });
foreach (var test in new[]
{
(Name: "matching", Item: good, Expected: "PASS"),
(Name: "wrong_length", Item: new Embedding("encoder-a@revision-1",
new float[] { 0.2f, 0.4f }), Expected: "dimension_mismatch"),
(Name: "wrong_model", Item: new Embedding("encoder-b@revision-1",
new float[] { 0.2f, 0.4f, 0.8f }), Expected: "model_mismatch"),
(Name: "non_finite", Item: new Embedding("encoder-a@revision-1",
new float[] { 0.2f, float.NaN, 0.8f }), Expected: "non_finite"),
(Name: "zero_vector", Item: new Embedding("encoder-a@revision-1",
new float[] { 0, 0, 0 }), Expected: "zero_norm")
})
{
string result = Validate(index, test.Item);
if (result != test.Expected)
throw new Exception("Unexpectedresultfortest.Name:result");Console.WriteLine("{test.Name}: {result}");
}
static string Validate(EmbeddingContract contract, Embedding candidate)
{
if (contract.Dimensions <= 0 || string.IsNullOrWhiteSpace(contract.ModelKey))
throw new ArgumentException("Invalid index contract.");
if (!string.Equals(contract.ModelKey, candidate.ModelKey,
StringComparison.Ordinal))
return "model_mismatch";
if (candidate.Values is null || candidate.Values.Length != contract.Dimensions)
return "dimension_mismatch";
double squaredNorm = 0;
foreach (float value in candidate.Values)
{
if (!float.IsFinite(value)) return "non_finite";
squaredNorm += (double)value * value;
}
// This example uses cosine similarity, which needs a nonzero vector.
return squaredNorm == 0 ? "zero_norm" : "PASS";
}
public sealed record EmbeddingContract(string IndexName, string ModelKey,
int Dimensions);
public sealed record Embedding(string ModelKey, float[] Values);
The contract names an index, an embedding configuration, and its dimension. Validate first checks the configuration identity, then the array shape, then each numeric value. It accumulates the squared norm in a double to avoid performing the multiplication in the narrower float type.
This fixture assumes cosine similarity. The zero-vector check belongs to that assumption because a zero vector has no direction to compare. A different distance metric can require a different validation policy.
Which failures does the program catch?
Run dotnet run. The five executed fixture assertions produced:
matching: PASS
wrong_length: dimension_mismatch
wrong_model: model_mismatch
non_finite: non_finite
zero_vector: zero_norm
The wrong-model case is especially useful: it has the correct number of values and would pass a length-only check. The non-finite case demonstrates why a correctly shaped array can still be unusable for a numeric comparison.
An editorial team such as Ranknod could use these distinctions when reviewing an AI search implementation, checking that a claimed embedding upgrade includes compatibility evidence rather than only a successful API response.
The tests establish the behavior of this validator for the supplied fixtures. They do not measure retrieval quality or verify the output of an external embedding service.
Where should validation run?
Apply the same contract on both sides of the index: when documents are embedded for insertion and when a query is embedded for search. Validating queries alone cannot reveal that a previous ingestion job populated the collection with the wrong representation.
Load the contract from a controlled configuration associated with the actual index version. Do not let the request choose its own expected dimension and then compare the vector against that value; that would validate internal agreement rather than compatibility with the destination.
Boundary | Useful check |
Embedding response | Returned length and finite values match the requested representation. |
Index write | Model provenance and dimension match the destination index. |
Query execution | Query representation is compatible with the indexed document representation. |
Configuration change | New model or dimension routes to an intentionally prepared index. |
Keep the validation result separate from the search result. An incompatible vector should stop before a request is submitted, with an error identifying the relevant contract and failure category.
Should you pad or truncate a mismatched vector?
Do not silently resize an arbitrary vector to satisfy the schema. Adding zeros or dropping coordinates changes the representation and can conceal the actual configuration error.
If an embedding model explicitly supports producing a requested lower dimension, use that documented mechanism and evaluate it with a matching index. Treat the output as its own configured representation. An application-side slice of an unsupported vector is not equivalent evidence.
When changing the representation, prepare the corresponding document embeddings and query path together. Keep a known rollback route instead of mixing old and new vectors in an index because their lengths happen to agree.
What should be checked before switching an index?
Validate the new document set and query vectors, then run a fixed collection of reviewed retrieval questions. Inspect which records are returned and whether the relevant evidence remains accessible. Schema compatibility is necessary for this workflow, but a valid vector can still retrieve poor results.
Record the index version, model key, dimension, and failure category in diagnostics. Avoid logging raw document text or full vectors when those values are unnecessary for troubleshooting.
Start with an explicit contract and reject mismatches at the application boundary. That turns a confusing search failure into a specific question: did the model, the vector shape, the numeric values, or the index configuration stop matching?
Join the conversation! Your thoughts help the community grow.