Modern data platforms increasingly separate data storage, metadata, and compute.

That separation is one of the reasons Apache Iceberg has become important for lakehouse architectures.

The data can remain in object storage while different engines such as Spark, BigQuery, Trino, and other Iceberg-compatible systems access the same tables.

But this architecture creates another critical component:

The catalog.

The catalog is responsible for knowing what tables exist, where their data lives, what schema they use, and which table snapshot represents the current state.

Google Cloud's Lakehouse runtime catalog is now designed to operate at this layer, with Spanner providing the scalable database foundation underneath it.

The announcement is significant because the catalog is no longer treated as a relatively passive metadata store. At large scale, particularly with AI agents generating many small and concurrent requests, the catalog itself becomes part of the application's runtime infrastructure.

A simplified architecture looks like this:

                 Applications / Agents
                         |
                         v
                 Query / Compute Engines
                  /        |        \
                 /         |         \
             Spark      BigQuery     Trino
                  \        |        /
                   \       |       /
                    v      v      v
                Iceberg REST Catalog
                         |
                         v
                      Spanner
                         |
                         v
                Table Metadata
                         |
                         v
                  Object Storage

The important distinction is that Spanner does not replace the Iceberg data files.

It provides the highly available, strongly consistent metadata foundation behind the managed catalog.

Why the Catalog Matters

An Apache Iceberg table is more than a collection of Parquet files.

The table also has metadata describing:

Schema
Partition specification
Snapshots
Manifest lists
Manifest files
Data file locations
Statistics

The catalog helps engines discover and manage that metadata.

Conceptually:

Query Engine
     |
     v
Catalog
     |
     +--> Where is the table?
     +--> What is its current snapshot?
     +--> What schema does it have?
     +--> Which metadata files describe it?
     |
     v
Object Storage
     |
     v
Data Files

Without a reliable catalog, different engines would have difficulty coordinating access to the same logical table.

The catalog therefore becomes part of the correctness boundary of the lakehouse.

Why Traditional Catalog Architectures Can Become Difficult

A catalog may initially appear simple.

For a small data platform, a relational database or metadata service may be sufficient.

At larger scale, however, the catalog has to deal with:

High request volume
Concurrent table updates
Atomic commits
Metadata growth
High availability
Multiple compute engines
Frequent small queries
Schema changes
Snapshot management

AI agents add another dimension.

An analytical workload might execute a relatively small number of large queries.

An agentic workload can generate many smaller requests:

Agent
 |
 +--> Discover table
 |
 +--> Inspect schema
 |
 +--> Read metadata
 |
 +--> Query data
 |
 +--> Check another table
 |
 +--> Compare results
 |
 +--> Repeat

The number of metadata operations can therefore become significant even when the amount of actual data scanned is relatively small.

Why Spanner Is a Useful Foundation

Spanner is designed for globally distributed, strongly consistent database workloads.

Its architecture combines relational database semantics with horizontal scaling.

That makes it a reasonable foundation for a catalog that needs:

Strong consistency
High availability
Concurrent reads and writes
Horizontal scalability
Transactional metadata updates

The Lakehouse runtime catalog uses Spanner to provide these properties rather than requiring customers to operate a traditional database-backed metastore themselves. Google describes the catalog as serverless and designed for high availability and concurrency.

This is important because catalog failures can affect the entire data platform.

If the catalog becomes unavailable, the underlying Parquet files may still exist, but engines can have difficulty discovering and committing changes to the logical tables.

Apache Iceberg REST Catalog

One of the most important parts of this architecture is the use of the Apache Iceberg REST Catalog specification.

This creates a standardized interface between compute engines and the catalog.

Instead of requiring every engine to implement a proprietary Google-specific integration, an Iceberg-compatible engine can communicate through the standard REST catalog interface.

The architecture becomes:

                 Iceberg-Compatible Engines
                    /       |        \
                   /        |         \
                  v         v          v
               Spark      Trino    Other Engines
                    \       |       /
                     \      |      /
                      v     v     v
                  Iceberg REST API
                         |
                         v
                  Runtime Catalog
                         |
                         v
                       Spanner

The catalog therefore becomes decoupled from the compute engine.

Google's documentation describes the Lakehouse runtime catalog as a managed metadata service that supports the Iceberg REST Catalog API and allows multiple engines to share tables without duplicating the underlying files.

Data Stays in Object Storage

This architecture is important for another reason:

The catalog does not require copying the underlying data into Spanner.

The table definition points to the underlying storage.

For example:

Spanner
  |
  +--> Table metadata
  +--> Catalog state
  +--> Table pointers
  |
  v
Cloud Storage
  |
  +--> Iceberg metadata
  +--> Manifest files
  +--> Parquet data

This is a metadata architecture, not a requirement to move the data into the database.

That means organizations can maintain large datasets in object storage while using a highly available managed service for metadata operations.

Atomic Commits Matter

Consider two writers updating the same Iceberg table.

Writer A
   |
   +--> Commit snapshot A

Writer B
   |
   +--> Commit snapshot B

If both writers operate concurrently, the catalog needs to maintain a consistent view of the table.

An incorrect metadata update could cause engines to see conflicting table states.

The catalog therefore needs transactional behavior around metadata changes.

The Lakehouse runtime catalog uses Spanner's transactional foundation for atomic commits and concurrency control. Google describes its underlying transactions as serializable, meaning committed transaction ordering is preserved for clients.

That is a significant property for production lakehouse workloads.

Strong Consistency Is More Important Than It Looks

Imagine an agent querying a table immediately after a pipeline updates it.

The sequence might be:

Pipeline
   |
   v
Commit new table snapshot
   |
   v
Agent asks for latest data
   |
   v
Catalog
   |
   v
Current snapshot

If metadata visibility were inconsistent, the agent might see an older table state even though the pipeline has already committed its changes.

For data applications, these small timing differences can produce confusing behavior.

Strong consistency reduces this class of problem.

Multiple Engines Can Share the Same Data Estate

A major advantage of Iceberg is interoperability.

An organization might use:

Apache Spark
     |
     +--> ETL

BigQuery
     |
     +--> Analytics

Trino
     |
     +--> Interactive SQL

AI Agent
     |
     +--> Natural-language analytics

All of these workloads can work against Iceberg tables through a shared catalog.

That is very different from maintaining independent copies of the same dataset for each engine.

The architecture becomes:

                 Shared Iceberg Tables
                         |
             +-----------+-----------+
             |           |           |
             v           v           v
           Spark      BigQuery      Trino

This reduces unnecessary data duplication and makes the lakehouse more interoperable.

Read and Write Interoperability

Interoperability is more valuable when it applies to both reads and writes.

For example:

Spark
  |
  +--> Write Iceberg table
          |
          v
      Shared catalog
          |
          +--> BigQuery reads
          |
          +--> Trino reads

A data platform becomes easier to evolve when teams are not forced to use a single compute engine for every operation.

Google describes support for read/write interoperability across supported Iceberg-compatible engines in its Lakehouse runtime catalog architecture.

The Agent-Scale Problem

AI agents change the workload pattern.

A human might execute:

10 queries

during an analysis session.

An agent could potentially perform:

Question
  |
  +--> Discover schema
  +--> Inspect tables
  +--> Search metadata
  +--> Run query
  +--> Validate result
  +--> Query another table
  +--> Compare
  +--> Refine
  +--> Query again

One user question can therefore produce many backend operations.

At organizational scale:

1000 users
    x
many agent steps
    x
many metadata operations

The catalog can become a high-throughput service in its own right.

This is one reason the announcement connects the runtime catalog with agent-scale workloads.

Catalog and Data Are Different Scaling Problems

This distinction is easy to miss.

Suppose an organization has:

100 TB of data

The catalog does not need to store 100 TB of data.

Instead, it stores information describing the datasets.

However, the number of tables, snapshots, manifests, metadata operations, and concurrent clients can still become very large.

Therefore:

Data scale
    !=
Metadata workload

A system can have a relatively moderate amount of data but an extremely high metadata request rate.

Spanner's horizontal scaling model is intended to handle that kind of workload without forcing the organization to manually shard the metadata database.

No Manual Sharding

Traditional databases can eventually reach scaling limits.

Teams may then need to introduce:

Sharding
Read replicas
Connection pooling
Partitioning
Manual capacity planning
Database failover

A distributed database such as Spanner is designed to scale horizontally without requiring application developers to manually implement database sharding.

Google states that Spanner automatically distributes data ranges and can dynamically scale compute and storage based on workload characteristics.

For a managed catalog, this reduces operational responsibility.

Serverless Changes the Operational Model

A traditional metastore can require infrastructure management.

For example:

VM
 |
 +--> Database
 |
 +--> Metastore service
 |
 +--> Connection pool
 |
 +--> Monitoring
 |
 +--> Patching
 |
 +--> Failover

A managed runtime catalog changes that to:

Application
    |
    v
Managed Catalog
    |
    v
Managed Database Infrastructure

The platform team still needs to understand capacity, permissions, data layout, and workload behavior.

But it no longer needs to operate the underlying metastore infrastructure in the same way.

Security and Credential Vending

The catalog is also involved in access control.

A useful architecture is:

Query Engine
      |
      v
Catalog
      |
      +--> Authorize request
      |
      +--> Provide controlled storage access
      |
      v
Object Storage

Credential vending can allow the catalog to provide temporary access information rather than requiring every compute engine to hold long-lived storage credentials.

Google describes support for credential vending and modern authentication mechanisms as part of the runtime catalog architecture.

That can reduce the need to distribute broad object-storage permissions throughout the data platform.

Governance Becomes More Important

A centralized catalog can become more than a table directory.

It can provide a common place to understand:

What datasets exist?
Who owns them?
What do they represent?
Where are they stored?
Which tables are related?
What policies apply?

Google's architecture integrates the runtime catalog with Knowledge Catalog and Cloud IAM to provide metadata, discovery, governance, and trusted context for AI workloads.

This becomes particularly useful when AI agents need to discover data.

An agent should not simply ask:

"What tables exist?"

and receive an uncontrolled list.

It needs metadata that helps determine which datasets are appropriate and authorized for the task.

AI Agents Need Context, Not Just Data

Suppose an agent finds a table:

customer_transactions

The table name alone is not enough.

The agent may need to understand:

What does the table represent?
What time period does it cover?
Which columns are sensitive?
Which business unit owns it?
Can this user access it?
What does each field mean?

Metadata becomes part of the AI context.

That means the catalog can become an important component of agentic data architecture.

Operational and Analytical Data Can Work Together

Another interesting architectural pattern is combining lakehouse data with operational data.

For example:

Iceberg
  |
  +--> Historical taxi trips

Spanner
  |
  +--> Current taxi zones

An application can combine the datasets to answer a question such as:

Which routes are currently associated
with the highest historical congestion?

This creates a bridge between:

Analytical data
       +
Operational data
       +
AI reasoning

The architecture is no longer a simple warehouse.

It becomes a broader data platform where analytical and operational systems can participate in the same workflow.

What This Means for .NET Developers

C# applications generally should not need to know how the catalog stores its internal metadata.

The application can work against a service boundary:

ASP.NET Core
      |
      v
Data Access Service
      |
      v
Iceberg-Compatible Catalog
      |
      v
Data Platform

A clean application abstraction might look like:

public sealed record TableReference(
    string Catalog,
    string Namespace,
    string TableName);

public interface IDataCatalog
{
    Task<TableMetadata> GetTableAsync(
        TableReference table,
        CancellationToken cancellationToken);
}

The application should depend on the catalog contract rather than Spanner internals.

That keeps the .NET service insulated from infrastructure implementation details.

Do Not Put Catalog Logic Into Every Application

A common architectural mistake is allowing each application to maintain its own understanding of table locations.

For example:

var tableLocation =
    "gs://company-data/analytics/orders/";

Hardcoding storage locations into applications creates coupling.

If the table moves, changes format, or is reorganized, every application may need to change.

A catalog exists partly to prevent this.

Instead:

Application
   |
   v
Catalog
   |
   v
Current table metadata

The catalog becomes the source of truth.

Migration From a Traditional Metastore

Organizations already using Hive or other catalog technologies do not necessarily need to redesign the entire lakehouse.

The managed Lakehouse runtime catalog supports both Apache Iceberg REST catalog workflows and Hive catalog functionality, allowing existing architectures to evolve toward a managed metadata layer.

A migration strategy can look like:

Existing Data
     |
     v
Existing Metastore
     |
     v
Catalog Migration
     |
     v
Managed Runtime Catalog
     |
     v
Multiple Query Engines

The important goal is to avoid unnecessary data movement.

The catalog migration should change metadata management without forcing the organization to rewrite petabytes of underlying data.

Important Compatibility Consideration

Open standards do not mean every version and feature is automatically compatible.

For example, the current Iceberg REST catalog endpoint supports Iceberg V2 tables and has V3 support in preview, while Iceberg V1 tables require an upgrade before using that endpoint.

Teams should therefore validate:

Iceberg version
Catalog API compatibility
Engine compatibility
Table features
Authentication
Storage format
Write behavior

before migrating production workloads.

Common Mistakes

Treating the catalog as just a metadata database

The catalog is part of the consistency and coordination layer for the lakehouse.

Assuming object storage is the only scaling problem

Metadata traffic can become a bottleneck independently of data volume.

Hardcoding table locations

Applications should discover table metadata through the catalog.

Giving every engine unrestricted storage access

Use appropriate authorization and credential-vending mechanisms.

Ignoring concurrent writers

Metadata commits need strong concurrency guarantees.

Assuming Iceberg compatibility means universal compatibility

Engine versions and supported Iceberg features still need validation.

Treating AI agents like ordinary query clients

Agents can generate many more metadata requests per user interaction.

Ignoring metadata quality

Poor descriptions and incomplete ownership information make data discovery harder for both humans and agents.

Advantages and Disadvantages

Advantages

Scalable metadata layer: Spanner provides a distributed foundation for high-concurrency catalog workloads.

Strong consistency: Transactional metadata operations help maintain reliable table state.

High availability: The managed architecture reduces the need to operate a separate metastore infrastructure.

Open interoperability: The Iceberg REST Catalog interface allows multiple compatible engines to share data.

Zero-copy architecture: Existing data can remain in object storage instead of being duplicated into the catalog database.

Better foundation for AI agents: High-volume metadata discovery and governed data access become easier to support.

Reduced operational burden: Teams do not need to manage the underlying database and metastore infrastructure themselves.

Disadvantages

More architectural abstraction: Developers need to understand both Iceberg concepts and the managed catalog model.

Vendor-specific capabilities still exist: Open APIs improve interoperability, but not every feature is necessarily portable across platforms.

Migration requires compatibility testing: Existing Iceberg versions, engines, and table features must be validated.

Metadata quality remains a human responsibility: A scalable catalog does not automatically create useful business metadata.

Additional governance design is required: Centralizing metadata makes permissions and ownership models more important.

When This Architecture Makes Sense

A managed Iceberg catalog backed by a distributed database is particularly attractive when an organization has:

Large data estate
Multiple query engines
High concurrent access
Many Iceberg tables
Frequent table updates
AI or agent workloads
Strict availability requirements
Need for open data formats

It may be unnecessary for a small analytics environment containing a few datasets and a single query engine.

The value increases as the data platform becomes more distributed.

A Practical Architecture

A production-oriented architecture can look like:

                         Users / Agents
                               |
                               v
                         Data Applications
                               |
                  +------------+------------+
                  |                         |
                  v                         v
             BigQuery                    Spark
                  |                         |
                  +------------+------------+
                               |
                               v
                    Apache Iceberg REST
                          Catalog
                               |
                               v
                            Spanner
                               |
             +-----------------+-----------------+
             |                 |                 |
             v                 v                 v
        Table Metadata     Commit State      Permissions
                               |
                               v
                         Object Storage
                               |
                               v
                      Parquet / Iceberg Data

The key separation is:

Compute
   !=
Catalog
   !=
Data Storage

Each layer has a distinct responsibility.

What Developers Should Watch

If your team is adopting this architecture, monitor more than query latency.

Useful signals include:

Catalog request latency
Catalog error rate
Concurrent metadata operations
Commit conflicts
Table discovery latency
Storage access failures
Authorization failures
Query engine compatibility
Agent-generated request volume

For AI workloads, also track the number of metadata operations generated by each user interaction.

An agent that executes 30 metadata operations to answer a simple question may indicate an opportunity to improve schema discovery or tool design.

Summary

Spanner becoming the foundation for Google's Lakehouse runtime catalog is important because it addresses a problem that is easy to underestimate:

At large scale, the metadata system is itself a production workload.

Apache Iceberg separates table data from compute engines, but that interoperability depends on a reliable catalog that can coordinate metadata, commits, table discovery, and access.

By using Spanner underneath a managed Iceberg REST catalog, Google Cloud is targeting the combination of:

Open table format
+
Multiple compute engines
+
Strong consistency
+
High availability
+
Horizontal scalability
+
Managed operations
+
AI-agent workloads

For .NET and application developers, the biggest lesson is architectural: applications should depend on the catalog rather than hardcoding storage locations or building their own metadata registry.

As lakehouse environments become more distributed and AI agents generate increasingly frequent data-discovery requests, the catalog is becoming a critical runtime component rather than simply a bookkeeping service.