Moving data from one lakehouse format to another can become expensive when it requires rewriting large datasets, building custom migration scripts, or maintaining a separate data-processing application.
Google Cloud Dataflow Job Builder provides a low-code way to import existing Delta Lake tables stored in Cloud Storage into Lakehouse for Apache Iceberg. The workflow uses a visual interface to configure a batch pipeline, read the source table, and write its data into a destination Iceberg table.
This is useful for data engineering teams that want to make existing Delta Lake datasets available through a Google Cloud lakehouse architecture without building an entire Apache Beam pipeline from scratch.
However, the import process still requires careful planning. The source must be a valid Delta Lake table, the destination schema must be compatible, and the pipeline must have the permissions and resources needed to read and write the data.
What Dataflow Job Builder Does
Dataflow is Google Cloud's managed data-processing service built on Apache Beam. Its Job Builder interface allows developers and data engineers to create pipelines by selecting sources, transformations, and destinations in the Google Cloud console.
The Job Builder supports several data sources, including Cloud Storage files, databases, BigQuery, Apache Iceberg, and Delta Lake. It also supports transformations such as filtering, mapping, joining, grouping, and SQL operations.
For Delta Lake imports, the workflow connects a Delta Lake table stored in Cloud Storage to a Lakehouse table using Apache Iceberg.
The process can be summarized as follows:
<box border radius="lg" padding={3} gap={2} align="center"> <box background="surface-secondary" radius="md" padding={3} width="100%" align="center" gap={1}> <icon name="database" size="xl" /> Delta Lake source <text color="secondary" size="sm">Parquet data files and _delta_log/ in Cloud Storage</text> </box> <icon name="arrow-down" size="xl" color="secondary" /> <box background="surface" border radius="md" padding={3} width="100%" align="center" gap={1}> <icon name="workflow" size="xl" /> Dataflow batch pipeline <text color="secondary" size="sm">Read the table and process its records</text> </box> <icon name="arrow-down" size="xl" color="secondary" /> <box background="surface-secondary" radius="md" padding={3} width="100%" align="center" gap={1}> <icon name="database" size="xl" /> Lakehouse destination <text color="secondary" size="sm">Apache Iceberg table available for downstream analytics</text> </box> </box>
The main benefit is that the pipeline can be configured through a visual interface rather than requiring developers to write and maintain custom transformation code for a straightforward import.
It is important to distinguish importing data from moving or converting the underlying storage layout manually. The Job Builder workflow is designed to make Delta Lake data available in the destination Lakehouse table. Teams should still validate the resulting table and understand the storage and metadata behavior of their chosen destination.
Prerequisites Before Starting
Before creating the import pipeline, verify that the source data, destination table, and Google Cloud project are ready.
1. Enable the required APIs
The documented workflow requires the Dataflow, BigQuery, and Lakehouse APIs to be enabled in the Google Cloud project.
You also need the appropriate Identity and Access Management (IAM) permissions to enable services and create the required resources.
Use the principle of least privilege when assigning permissions. The user configuring the job and the service account executing the pipeline may need different permissions.
2. Prepare the Delta Lake source
The source must be a valid Delta Lake table stored in Google Cloud Storage.
Its root directory must contain the Parquet data files and the _delta_log/ transaction log directory. For example:
gs://my-data-bucket/
└── tables/
└── orders/
├── _delta_log/
│ ├── 00000000000000000000.json
│ └── ...
├── part-00000-....parquet
└── part-00001-....parquet
This is an illustrative directory structure. The actual table can contain additional log files, checkpoints, and data files.
The important detail is that the source path must point to the Delta Lake table root, not simply to a directory containing Parquet files.
Reading individual Parquet files without respecting the Delta transaction log can produce an incorrect view of the table. Delta Lake metadata tracks which files belong to the table's committed state, so the import must use a valid table rather than treating the directory as an arbitrary collection of files.
This workflow does not support Amazon S3 as the Delta Lake source. If your source table is stored in S3, you need a separate supported approach to make the data available in the required source environment.
3. Create the destination Lakehouse table
Before importing data, identify the destination Iceberg catalog, namespace, and table.
You can import into a new table or an existing table, but the schema behavior differs:
For a new destination table, Dataflow automatically creates the schema.
For an existing destination table, the schema is not modified. It must correctly map to the source Delta Lake schema.
This distinction is important when importing datasets with evolving schemas, nested structures, decimal types, or nullable fields.
Do not assume that selecting an existing table automatically makes it compatible with the source.
Import a Delta Lake Table Through the Console
Once the prerequisites are satisfied, the import can be configured from the Google Cloud console.
Step 1: Open the Lakehouse catalog
Navigate to the Lakehouse runtime catalog in the Google Cloud console.
Select the catalog and namespace containing the destination table. If the destination table does not exist, create the required destination resources according to your Lakehouse configuration.
Open the table details page for the table you want to populate.
Step 2: Select the Delta Lake import workflow
From the table details page, select Import table, then choose From Delta Lake (Batch).
This opens Dataflow Job Builder with the Delta Lake-to-Lakehouse pipeline blueprint loaded.
The batch designation matters. This workflow is designed for batch imports rather than continuous streaming ingestion.
Review the generated pipeline before running it, especially if the import is part of a repeatable production process.
Step 3: Configure the source table
In the Sources section, expand the ReadFromDeltaLake source configuration.
Enter the Cloud Storage URI of the Delta Lake table root.
For example:
gs://my-data-bucket/tables/orders
Replace the bucket and directory with the actual source table location.
The path should contain both the table's Parquet data and its _delta_log/ directory.
The source configuration also provides an optional field for Hadoop configuration properties. If your environment requires additional properties to access Cloud Storage, configure them there.
For example, the documented configuration can include:
fs.gs.impl=com.google.cloud.hadoop.fs.gcs.GoogleHadoopFileSystem
Only add properties that your environment requires. A Hadoop filesystem setting should not be treated as a replacement for the correct service-account permissions or Cloud Storage access configuration.
After entering the source details, confirm the configuration.
Step 4: Review the destination
Expand the WriteToIceberg sink configuration in the Sinks section.
The import blueprint typically prepopulates information such as the destination Lakehouse table, catalog, and warehouse location.
Verify that the destination points to the intended catalog, namespace, and table before running the job.
For an existing table, confirm that its schema is compatible with the Delta Lake source. For a new table, review the expected schema after the import completes.
A destination mistake can be more costly than a source configuration error because the pipeline might successfully process data into an unintended table.
Step 5: Run the batch pipeline
After reviewing the source and destination, open the Dataflow options section and select Run job.
Dataflow creates the batch job and displays its execution graph. You can monitor the job status and inspect the individual processing stages.
A successful submission does not mean that the import has completed successfully. Wait for the job to reach a terminal status and confirm that it reports Succeeded.
If the job fails, inspect the job logs and worker logs before retrying. Repeatedly launching the same pipeline without understanding the error can waste processing time and complicate troubleshooting.
Validate the Imported Iceberg Table
After the Dataflow job succeeds, verify that the destination table contains the expected data.
The Lakehouse integration documentation describes querying the imported table through BigQuery. The table can be addressed using its project, catalog, namespace, and table identifiers.
For example, the following query illustrates a basic inspection:
SELECT *
FROM `my-project.my-catalog.my-namespace.orders`
LIMIT 10;
Replace the example identifiers with the actual values for your environment. Confirm the identifier format supported by the destination before using the query.
A successful query establishes that the table can be accessed through the query interface, but it does not prove that every record was imported correctly.
For production imports, perform additional checks:
Compare source and destination row counts where a reliable source count is available.
Verify key columns and representative records.
Check null values, numeric precision, timestamps, and nested fields.
Confirm that expected partitions or business date ranges are present.
Compare important aggregates, such as order totals or event counts.
Run representative downstream queries used by reports and applications.
For large datasets, avoid repeatedly scanning the entire table just to validate a migration. Use partition-aware checks and targeted comparisons where possible.
Handle Schema Compatibility Carefully
Schema compatibility is one of the most important considerations in Delta Lake imports.
Suppose the source table contains these fields:
order_id STRING
customer_id STRING
order_total DECIMAL
created_at TIMESTAMP
The destination table must represent these fields compatibly for the import to succeed.
Potential problems include a destination missing a source field, incompatible data types, unexpected nullability, or differences in how nested data is represented.
Existing destination schemas are not automatically modified by this workflow. If the destination table has an incompatible schema, review the mismatch and decide whether to correct the destination or create a suitable new table.
Avoid changing a production schema solely to make an import pass without understanding the downstream impact. A change in numeric precision, timestamp interpretation, or nullable fields can affect reports and applications even if the pipeline finishes successfully.
For repeatable imports, treat the source and destination schemas as part of the migration contract. Validate them before launching the job and record intentional differences.
Common Problems and Troubleshooting
The source table cannot be read
First, verify that the source URI points to the Delta Lake table root. Confirm that the _delta_log/ directory and the required data files exist.
Next, check that the Dataflow worker identity has the required access to the Cloud Storage objects.
A directory containing only Parquet files is not sufficient evidence that it is a valid Delta Lake table.
The job fails with a permission error
Check the permissions of the identity that runs the Dataflow workers, not only the account used to configure the job.
The pipeline needs access to the source data and the destination resources. The required permissions depend on the project's resource configuration.
Review the error logs to determine whether the failure involves Cloud Storage, Dataflow, the destination catalog, or another dependency.
The destination schema is incompatible
For an existing table, compare its schema with the source table's schema and correct the incompatibility before rerunning the import.
Do not expect the import workflow to modify the existing destination schema automatically.
The job succeeds, but the results appear incorrect
Check whether the configured source path identifies the intended table, whether the destination is the expected table, and whether representative records match.
Also verify the source's committed Delta Lake state and investigate any differences in data types or timestamp interpretation.
A successful job status is necessary, but data-quality validation remains the responsibility of the migration workflow.
The pipeline is expensive or takes longer than expected
Large imports require processing resources and can generate substantial data-read and data-write activity. Review the Dataflow job graph, worker logs, dataset size, and processing behavior.
Avoid unnecessary repeated imports when a validation query or smaller representative dataset would answer the immediate question.
For recurring ingestion or highly customized transformations, evaluate whether a reusable Dataflow pipeline or an Apache Beam implementation is more appropriate than repeatedly configuring a one-off import.
When to Use Job Builder Instead of Custom Apache Beam Code
Job Builder is a strong fit when the import follows a supported source-to-destination pattern and the pipeline does not require extensive custom logic.
It offers a visual workflow, supports pipeline configuration through the console, and can save supported pipeline definitions as Apache Beam YAML for reuse.
A custom Apache Beam pipeline may be more suitable when the migration requires specialized transformations, complex validation rules, custom error handling, or source and destination connectors that the Job Builder does not support.
The decision should be based on the complexity of the data workflow, not simply on whether a no-code option exists.
For a straightforward Delta Lake-to-Lakehouse import, Job Builder reduces implementation effort. For a migration with substantial business logic or strict data-quality requirements, additional pipeline code and automated validation may be necessary.
Summary
Google Cloud Dataflow Job Builder provides a low-code way to import Delta Lake tables stored in Cloud Storage into Lakehouse for Apache Iceberg using a batch pipeline.
The workflow starts by selecting the destination table, configuring the Delta Lake source root, reviewing the Iceberg sink, and running the Dataflow job. New destination tables receive an automatically created schema, while existing destination tables must already have a compatible schema.
The main constraints are that the source must be a valid Delta Lake table in Cloud Storage, the workflow requires a batch job, and Amazon S3 is not supported as the source for this import path.
Before relying on the imported data, verify job completion, inspect the destination table, compare representative records, and validate the schema and downstream queries. These checks turn a convenient visual import into a dependable data migration process.
Join the conversation! Your thoughts help the community grow.