A database-backed API can look simple until the database starts storing data that does not fit neatly into traditional relational columns.

A string, int, DateTime, or decimal is straightforward. JSON documents and vector embeddings are different. JSON introduces semi-structured data, while vectors are usually large numerical arrays used by modern search and AI workloads. Exposing those types through an API can require additional serialization, filtering, and database-specific handling.

Data API Builder 2.1.5 adds support for JSON and vector data types, making it easier to expose databases containing these kinds of values through its generated REST and GraphQL APIs.

The change is particularly relevant for .NET developers building data services around AI applications. A database may now contain traditional business records alongside JSON metadata and vector embeddings, and the API layer needs to preserve those types instead of reducing everything to plain strings.

What Data API Builder Does

Data API Builder, often called DAB, provides an API layer over supported databases without requiring developers to write a complete controller or resolver for every table and view.

Conceptually, the architecture looks like this:

Client Application
       |
       v
REST / GraphQL API
       |
       v
Data API Builder
       |
       v
Database

The configuration describes which database entities should be exposed and how clients can interact with them.

That can remove a significant amount of repetitive API plumbing for applications where the database already represents the core data model.

The important detail is that DAB does not turn the database into an interchangeable storage engine. Database-specific types and capabilities still matter.

JSON and vector support makes that boundary more useful for applications that use modern database features.

Why JSON Support Matters

Relational databases are very good at structured data.

Consider a customer record:

CustomerId
Name
Email
CreatedAt

Each field has a predictable type.

Real applications frequently have additional information that is less predictable:

{
  "theme": "dark",
  "notifications": {
    "email": true,
    "sms": false
  },
  "preferences": {
    "language": "en",
    "timezone": "Asia/Kolkata"
  }
}

Putting every possible preference into separate columns can make the schema unnecessarily wide and difficult to evolve.

A JSON column provides another option.

The database can keep strongly structured fields relational while storing flexible metadata in JSON.

The API layer then needs to preserve the JSON structure when returning data to clients.

Returning this:

{
  "preferences": "{\"language\":\"en\",\"timezone\":\"Asia/Kolkata\"}"
}

is very different from returning:

{
  "preferences": {
    "language": "en",
    "timezone": "Asia/Kolkata"
  }
}

The second representation is what application developers generally expect from a JSON-aware API.

JSON Does Not Mean "Put Everything in One Column"

This is an important design consideration.

JSON is useful when the structure is genuinely variable or when the database has a specific reason to store document-style data.

It is not a replacement for relational modeling.

For example, an order should not necessarily become:

{
  "order": {
    "customer": "...",
    "items": [
      ...
    ],
    "total": ...
  }
}

simply because JSON is available.

If OrderItem needs foreign keys, reporting, constraints, joins, indexing, and independent lifecycle management, a relational table is usually a better model.

JSON is more appropriate for information such as optional metadata, configuration, provider-specific properties, or data whose structure legitimately changes.

The API layer should expose the database model that makes sense rather than encouraging developers to put all application state into JSON.

Why Vector Support Is More Interesting

Vector support is particularly relevant to AI applications.

An embedding model converts text, images, or other information into a numerical representation.

A simplified example might look like:

"How do I reset my password?"
        ↓
Embedding model
        ↓
[0.018, -0.142, 0.731, ...]

The resulting vector can be stored in a database and compared with other vectors.

This allows applications to implement semantic search.

Instead of searching only for exact words:

password reset

the application can search for content that is semantically related to:

I forgot how to access my account

That is one of the foundations of retrieval-augmented generation and other AI search workflows.

What a Vector Column Represents

A vector is essentially an ordered collection of numerical values.

For example:

[0.12, -0.41, 0.88, 0.03]

Real embedding vectors are typically much larger.

The important property is dimensionality.

If an embedding model produces a vector with 1,536 dimensions, the database column needs to be configured to represent vectors of the expected dimensionality.

Conceptually:

Document
   |
   +-- Id
   +-- Content
   +-- Metadata
   +-- Embedding
                |
                +-- 1536 numerical values

The API layer must preserve that type rather than treating the vector as an ordinary text field.

JSON and Vectors Fit Naturally Into AI Data Architectures

A modern AI application may have a database record like this:

Document
├── Id
├── Title
├── Content
├── Metadata       → JSON
└── Embedding     → Vector

The same record can therefore contain:

Structured data
      +
Semi-structured metadata
      +
AI embedding

That is useful for applications that need both traditional application queries and semantic retrieval.

A typical architecture could look like:

User Query
    |
    v
Embedding Model
    |
    v
Query Vector
    |
    v
Database
    |
    +-- Vector similarity search
    |
    +-- Metadata filtering
    |
    v
Relevant Documents
    |
    v
Application / LLM

Data API Builder sits between the application and database when the application needs API access to those records.

A Practical Data Model

Suppose an application stores knowledge-base documents.

A simplified relational model might contain:

Documents
---------
Id
Title
Content
Metadata
Embedding
CreatedAt

The Metadata field could contain:

{
  "department": "engineering",
  "product": "billing",
  "language": "en"
}

The Embedding field contains the numerical representation generated by an embedding model.

The application can then retrieve the document through an API without building a custom endpoint solely to serialize these database types.

That can be useful for internal tools, administrative applications, AI assistants, and data-heavy applications where much of the API behavior is conventional CRUD.

API Design Still Matters

Generated APIs reduce boilerplate, but they do not remove API design decisions.

Suppose a client can retrieve:

Document

Should it receive the complete embedding vector?

Often, the answer is no.

An embedding may contain hundreds or thousands of numbers. Sending it to a browser for every document can increase payload size without providing any value to the client.

The application may need the vector internally for retrieval while exposing only:

{
  "id": 42,
  "title": "Resetting a user password",
  "content": "...",
  "metadata": {
    "department": "support"
  }
}

The fact that a database column is API-accessible does not mean every client should receive it.

This is where authorization and entity-level API configuration remain important.

Vector Data Has Different Performance Characteristics

Vector data should also not be treated like an ordinary column.

A vector can be large.

If an embedding has 1,536 floating-point values and each value consumes four bytes, the raw numerical data alone is roughly:

1536 × 4 = 6144 bytes

That is about 6 KB before considering database storage overhead, indexes, row structure, serialization, and transport.

Multiply that by hundreds of thousands or millions of records and the storage implications become significant.

This is why vector workloads often require deliberate decisions about:

  • Embedding dimensions

  • Storage format

  • Indexing

  • Similarity metric

  • Query patterns

  • Result size

  • API payloads

The API layer does not remove those database-level concerns.

JSON and Vector Types Should Not Be Confused With Strings

A common workaround for unsupported database types is to serialize everything into strings.

For JSON:

"{\"department\":\"engineering\"}"

For vectors:

"[0.12,-0.41,0.88,...]"

This can make serialization possible, but it loses type information.

The client now has to deserialize the string itself.

That creates several problems:

  • More application code

  • More opportunities for parsing errors

  • Less natural API contracts

  • More difficult querying

  • Potentially larger payloads

  • Confusion between textual data and structured data

Native type support is therefore useful because it preserves the semantic meaning of the database value through the API boundary.

Security Considerations

JSON fields deserve the same security attention as ordinary columns.

A JSON object may contain:

{
  "internalRole": "administrator",
  "serviceToken": "...",
  "debug": true
}

If the entire JSON document is exposed through an API, sensitive fields may leak even when the top-level entity looks harmless.

Vector data has a different concern. Embeddings are not generally readable business data in the same way as a name or email address, but they can still represent information derived from sensitive content.

Developers should therefore avoid assuming that "it is only an embedding" means the value can be exposed freely.

API permissions should be designed around what clients actually need.

Common Mistakes

Treating JSON as a Replacement for a Relational Schema

JSON is flexible, but flexibility can become a maintenance problem when the data actually has a stable structure.

If a field needs joins, constraints, reporting, and predictable indexing, a relational column or table may be more appropriate.

Returning Embeddings to Every Client

A vector can be much larger than the rest of a record.

If a client only needs search results, returning the full embedding wastes bandwidth and increases processing requirements.

Keep internal retrieval data internal unless a client genuinely needs it.

Assuming Vector Support Automatically Provides Semantic Search

Supporting a vector data type is not the same thing as implementing a complete vector-search system.

You still need:

Embedding generation
        ↓
Vector storage
        ↓
Similarity search
        ↓
Appropriate indexing
        ↓
Ranking
        ↓
Application logic

The API layer solves only part of that architecture.

Ignoring Database-Specific Behavior

JSON and vector capabilities vary between database engines.

The exact type, indexing behavior, operators, query syntax, and performance characteristics depend on the underlying database.

Data API Builder can expose database data, but it does not erase those differences.

Troubleshooting

When JSON or vector data behaves unexpectedly through an API, isolate the problem into three layers:

Client
  ↓
Data API Builder
  ↓
Database

First verify that the database itself stores and returns the value correctly.

Then determine whether Data API Builder is correctly mapping the database type.

Finally inspect the API response and serialization behavior.

For JSON, pay particular attention to whether the value is returned as structured JSON or as an escaped string.

For vectors, verify dimensionality and the database's expected representation.

If a vector operation fails, check whether the query uses the correct dimensions and database-specific vector semantics before investigating the API layer.

Advantages and Disadvantages

Advantages

Better support for modern database workloads: Applications can expose structured relational data alongside JSON and vector values without treating everything as plain text.

Useful for AI applications: Vector data can participate in architectures that combine application records with embeddings and semantic retrieval.

Less API boilerplate: Teams can expose supported database entities without writing custom serialization code for every property.

More natural data representation: JSON remains structured JSON and vector values retain their database-level meaning instead of being reduced to strings.

Disadvantages

Database behavior still matters: Different databases implement JSON and vector capabilities differently, so developers still need to understand the underlying storage engine.

Large vectors can increase payload and storage costs: Exposing embeddings through an API can produce unexpectedly large responses.

Type support does not solve application architecture: Developers still need to design indexing, retrieval, authorization, and API boundaries correctly.

Flexible JSON can become difficult to maintain: Using JSON for highly structured business data can move complexity from the schema into application code.

When Data API Builder Makes Sense

Data API Builder is particularly useful when the application needs conventional data access and the database already contains a well-defined model.

It can be a good fit for:

Internal applications
Administrative tools
CRUD-heavy applications
Data services
AI applications with database-backed metadata
Applications exposing existing database entities

It becomes less attractive when the API needs substantial domain-specific behavior.

For example, if creating an order requires:

Validate inventory
Reserve stock
Calculate pricing
Authorize payment
Publish event
Update multiple systems

a generated CRUD endpoint should not become the application's domain service simply because it is available.

Generated data APIs are most useful when they match the actual behavior the application needs.

Summary

Data API Builder 2.1.5's JSON and vector data type support is useful because modern applications increasingly store more than traditional relational values.

JSON can hold flexible metadata without forcing every optional property into the relational schema. Vector values can store embeddings used by semantic search and AI retrieval systems. Exposing those values through an API without reducing them to strings makes the application boundary more natural.

The important limitation is that type support is not the same as complete application support. Developers still need to make decisions about schema design, vector dimensions, indexing, payload size, authorization, and database-specific behavior.

For AI-enabled applications in particular, the combination of structured relational data, JSON metadata, and vector embeddings is becoming a practical database pattern. Data API Builder's expanded type support makes it easier to expose that data, but the underlying database and API design decisions still belong to the application team.