Apache Iceberg has become an important table format for data lakes because it adds database-like table management to object storage.

Instead of treating a collection of files in Amazon S3 as a simple folder of data, Iceberg maintains metadata about snapshots, schemas, partitions, and table changes. That makes it possible to build data platforms where multiple engines can work with the same tables without each engine inventing its own table-management layer.

Amazon S3 Tables now adds support for Apache Iceberg version 3 table specifications, including newer data types and metadata capabilities. AWS announced the support on September 30, 2026. (aws.amazon.com)

For data engineers, this is more than a version-number change.

Iceberg V3 introduces data types such as nanosecond timestamps and variant data, along with additional table metadata capabilities that are useful for modern analytics and semi-structured data workloads.

The practical question is how these features affect the design of an S3-based data lake.

What Are S3 Tables?

Amazon S3 Tables are designed specifically for tabular analytics workloads.

Traditional S3 storage gives you objects:

bucket/
    data/
        part-001.parquet
        part-002.parquet
        part-003.parquet

That works well for storing files.

But a data warehouse or analytical table needs more information:

Schema
Partitions
Snapshots
Table history
Metadata
Deletes
Updates

Apache Iceberg provides that table layer.

S3 Tables combine S3 storage with table-oriented management optimized for analytics workloads. AWS manages the underlying table infrastructure and automatically handles tasks such as compaction and snapshot management for S3 Tables. (docs.aws.amazon.com)

The result is closer to:

             S3 Tables
                 |
        +--------+--------+
        |                 |
     Storage          Table Metadata
        |                 |
        +--------+--------+
                 |
             Iceberg

This makes S3 a more capable foundation for analytical table workloads.

What Changes With Iceberg V3?

Iceberg V3 extends the table specification with new capabilities while retaining the core ideas that made Iceberg useful.

Some of the notable changes include:

Nanosecond timestamps
Variant data type
New metadata structures
Deletion vectors
Row lineage
Expanded table capabilities

AWS specifically highlights support for new Iceberg V3 data types such as timestamp_ns and variant in S3 Tables. (aws.amazon.com)

For developers, the most interesting changes are probably the new types.

Nanosecond Timestamps

Traditional timestamp precision is often sufficient for business applications.

For high-frequency event streams, observability data, financial events, telemetry, or other workloads where events can occur extremely close together, higher timestamp precision can become useful.

Iceberg V3 introduces nanosecond timestamp support.

Conceptually:

timestamp_ms
2026-09-30 10:15:30.123

timestamp_us
2026-09-30 10:15:30.123456

timestamp_ns
2026-09-30 10:15:30.123456789

The difference may look small, but at large event volumes it can matter.

Imagine an event-processing system recording:

DeviceId
EventType
EventTimestamp
Sequence
Payload

Two events could share the same millisecond timestamp even though they occurred at different points within that millisecond.

Nanosecond precision gives the table a way to preserve more of the original event timing.

That does not automatically make event ordering perfect, though.

If the source system does not provide reliable high-resolution timestamps, storing more precision does not create information that did not exist.

The Variant Data Type

The variant type is arguably more interesting for application developers.

Modern data pipelines frequently receive semi-structured JSON.

Consider an event stream:

{
  "eventType": "purchase",
  "customerId": "C1001",
  "metadata": {
    "campaign": "spring",
    "device": "mobile"
  }
}

Another event may have:

{
  "eventType": "login",
  "customerId": "C1002",
  "metadata": {
    "browser": "Edge",
    "location": {
      "country": "IN"
    }
  }
}

The structure is related but not identical.

Traditional relational modeling would require decisions about columns and nested structures.

A variant type provides a way to represent semi-structured values without forcing every possible field into a rigid schema.

That makes it useful for workloads such as:

Application events
IoT telemetry
API payloads
Logs
Machine-generated metadata
AI-generated records

Iceberg V3's variant type is designed specifically for this kind of semi-structured data. (iceberg.apache.org)

Why Variant Matters for Data Lakes

Data lakes frequently have an uncomfortable middle ground.

The data is structured enough that analysts want to query it, but flexible enough that forcing everything into a fixed relational schema becomes expensive.

For example, an application may generate events like:

{
  "type": "model_request",
  "model": "example-model",
  "metadata": {
    "latency": 120,
    "region": "us-east",
    "tokens": 1450
  }
}

A few weeks later, the application adds:

{
  "type": "model_request",
  "model": "example-model",
  "metadata": {
    "latency": 120,
    "region": "us-east",
    "tokens": 1450,
    "cacheHit": true
  }
}

With a rigid schema, adding fields can require schema evolution.

With a variant field, additional properties can remain part of the semi-structured value.

That does not mean variant should replace normal columns.

Frequently queried fields should generally still have well-defined schema where practical.

Variant is most useful when the structure itself is expected to evolve.

Iceberg Schema Evolution Still Matters

One of Iceberg's strengths is schema evolution.

A table can evolve without rewriting every historical data file simply because the logical schema changed.

For example:

Version 1
CustomerId
EventType
Timestamp

Version 2
CustomerId
EventType
Timestamp
Region

The table metadata tracks the evolution.

Iceberg uses field IDs rather than relying solely on column positions or names, which allows schema changes to be handled more safely than approaches that treat a schema as a fixed file layout.

V3 extends the table specification while preserving this broader approach to table evolution.

The important engineering point is that schema evolution should be treated as part of the data contract, not something that happens accidentally because a new application release added a field.

Deletion Vectors

Iceberg V3 also introduces deletion vectors.

Traditional row-level deletes can require writing new files or tracking delete information separately.

A deletion vector provides metadata describing which rows should be considered deleted without necessarily rewriting the entire underlying data file.

Conceptually:

Data File
+----------------+
| Row 1          |
| Row 2          |
| Row 3  X       |
| Row 4          |
| Row 5  X       |
+----------------+
        |
        v
Deletion Vector
    Row 3
    Row 5

The original file can remain intact while the table metadata identifies the deleted rows.

This can be particularly useful for workloads that perform frequent row-level updates or deletes.

It also creates additional metadata that query engines need to understand correctly.

S3 Tables' support for Iceberg V3 therefore depends not only on storing the data but also on the ecosystem of engines and tools reading those tables.

Row Lineage

Another Iceberg V3 capability is row lineage.

The idea is to make it possible to track information about the origin or evolution of individual rows through table operations.

That can be useful in environments where data quality, auditing, debugging, or downstream lineage matters.

For example:

Raw Event
    |
    v
Cleaned Table
    |
    v
Aggregated Table
    |
    v
Business Report

When a result looks incorrect, understanding how the underlying records moved through the data pipeline can be valuable.

Row lineage does not magically solve data lineage across an entire organization, but it provides more metadata at the table level that downstream systems can use.

S3 Tables and Analytics Engines

The value of Iceberg is largely tied to interoperability.

A table format becomes much more useful when different engines can read the same table.

An architecture might look like:

                   S3 Tables
                       |
          +------------+------------+
          |            |            |
          v            v            v
        Athena       Spark        EMR
          |            |            |
          +------------+------------+
                       |
                       v
                  Analytics

The exact engine support for specific Iceberg V3 features matters, though.

Just because the storage layer supports a new table specification does not mean every query engine in the environment supports every feature immediately.

This is one of the first things teams should verify before upgrading production tables.

Compatibility Is the Main Migration Question

Moving from Iceberg V2 to V3 is not simply a matter of changing a configuration value.

The table specification, writers, readers, catalogs, query engines, and downstream tools all need to be considered.

A typical environment might contain:

S3 Tables
    |
    +---- Athena
    +---- Spark
    +---- ETL jobs
    +---- Data quality tools
    +---- BI tools
    +---- ML pipelines

If one component does not understand a new V3 feature, the team may not be able to upgrade the entire table immediately.

A safer migration approach is:

Inventory Readers/Writers
        ↓
Check Iceberg V3 Support
        ↓
Test New Data Types
        ↓
Create Test Table
        ↓
Validate Queries
        ↓
Validate ETL
        ↓
Validate BI/ML Consumers
        ↓
Upgrade Production

This is particularly important for the variant type.

A table containing variant data is only useful if the tools consuming it can interpret that data correctly.

A Practical Table Example

Imagine an application telemetry table:

events
--------------------------------
event_id
event_time
service_name
event_type
payload

A V3-oriented design might use:

event_id       -> string
event_time     -> timestamp_ns
service_name   -> string
event_type     -> string
payload        -> variant

The structured fields remain easy to query:

SELECT
    service_name,
    event_type,
    event_time
FROM events
WHERE service_name = 'orders';

The flexible payload can retain fields that vary between event types.

That gives the data platform a useful balance:

Stable business fields
        +
Flexible event payload

This is generally more practical than making every field in the payload a first-class column.

When Not to Use Variant

Variant is not an excuse to put an entire relational database into one JSON column.

If the application always queries:

CustomerId
Region
OrderStatus
OrderDate
Revenue

those fields should normally remain explicit columns.

For example:

SELECT
    customer_id,
    region,
    order_status,
    order_date,
    revenue
FROM orders
WHERE region = 'APAC'
  AND order_status = 'Completed';

Putting all five values into an opaque semi-structured field would make common analytical queries harder to understand and potentially harder to optimize.

A good design uses structured columns for stable, frequently queried data and variant for data whose structure legitimately changes.

S3 Tables and Data Maintenance

Another advantage of S3 Tables is that AWS manages several table-maintenance operations.

S3 Tables automatically perform tasks such as compaction and snapshot management according to the table configuration. (docs.aws.amazon.com)

This matters because Iceberg tables can accumulate many small files as data is continuously written.

Without maintenance, a table can end up looking like:

part-001.parquet
part-002.parquet
part-003.parquet
...
part-982341.parquet

Many small files can increase metadata overhead and reduce query efficiency.

Compaction reorganizes data into more useful file layouts.

Managed maintenance reduces the amount of infrastructure work the data platform team needs to build and operate itself.

Common Mistakes

Upgrading Without Checking Readers

A table writer supporting Iceberg V3 does not guarantee that every consumer supports every V3 feature.

Inventory the complete pipeline before migration.

Putting Everything Into Variant

Variant is useful for evolving semi-structured data.

It should not replace a well-designed relational schema for stable analytical dimensions and measures.

Assuming More Timestamp Precision Fixes Ordering

Nanosecond timestamps preserve more precision, but they do not automatically establish event ordering if the source system generates timestamps inconsistently.

If ordering matters, consider an explicit sequence or event identifier as well.

Ignoring Metadata Growth

Iceberg is metadata-driven.

Adding more sophisticated table features means query engines and table-management systems need to handle that metadata correctly.

Monitor table health as part of the data-platform lifecycle.

Treating V3 as a Cosmetic Upgrade

Iceberg V3 introduces meaningful changes to the table specification.

Test the actual workloads and tools instead of assuming that an existing V2 workload will behave identically after an upgrade.

Advantages and Disadvantages

Advantages

Modern data types become available in S3 Tables. Nanosecond timestamps and variant data can better represent high-frequency and semi-structured workloads. (aws.amazon.com)

Semi-structured data can be represented more naturally. The variant type is useful when event payloads or application-generated metadata evolve over time.

Row-level data operations have better options. Features such as deletion vectors can reduce the need to rewrite entire data files for some workloads.

Managed table maintenance reduces operational work. S3 Tables handles maintenance operations such as compaction and snapshot management. (docs.aws.amazon.com)

Iceberg remains an open table format. The table specification is designed for interoperability across data-processing engines rather than tying the data to one query engine.

Disadvantages

V3 compatibility needs to be verified across the entire data platform. A new table feature is only useful if readers and writers can consume it correctly.

More capabilities mean more metadata complexity. Features such as deletion vectors and row lineage add information that engines and operational tooling need to understand.

Variant can be misused. Treating every field as semi-structured data can make analytical models harder to maintain.

Migration requires testing. Existing Iceberg V2 workloads should not be upgraded blindly, particularly when multiple engines and external tools consume the same tables.

When Iceberg V3 Makes Sense

The new capabilities are particularly interesting for workloads involving:

High-frequency event data
IoT telemetry
Application logs
AI and ML datasets
Semi-structured application events
Large analytical tables
Frequent row-level changes
Data lakes shared across multiple engines

For a simple warehouse table with stable columns and predictable workloads, there may be little immediate reason to move to V3.

The value appears when the workload actually benefits from the new table capabilities.

Summary

S3 Tables adding Apache Iceberg V3 support gives AWS data engineers access to newer table-format capabilities without giving up the managed nature of S3 Tables.

The most noticeable changes are support for newer data types such as nanosecond timestamps and variant data, along with Iceberg V3 capabilities such as deletion vectors and row lineage. (aws.amazon.com)

For application and data engineers, the variant type is particularly interesting because modern event streams rarely remain perfectly uniform. It provides a structured way to keep evolving payloads without forcing every possible property into a fixed schema.

But that flexibility needs discipline.

Stable fields that drive common queries should remain explicit. V3 adoption should be tested against every engine, ETL process, BI tool, and ML pipeline that reads or writes the table. And higher timestamp precision should not be confused with guaranteed event ordering.

The broader direction is clear: data lakes are becoming more table-like, and table formats are becoming more capable of handling the kinds of structured and semi-structured data generated by modern applications.

For teams already using S3 Tables and Apache Iceberg, V3 support is worth evaluating when those new capabilities solve a real data-modeling or operational problem. It is not simply a version upgrade to perform for its own sake.