Coding agents have become good at writing SQL, Python, PySpark, notebooks, and pipeline code.

The problem is that code generation is only one part of data engineering.

An agent can write:

SELECT *
FROM sales.orders
WHERE status = 'pending';

But it may not know:

  • Which project contains the table.

  • Which dataset is trusted.

  • How the table is partitioned.

  • Which columns contain sensitive data.

  • Whether the query will scan terabytes.

  • Why yesterday's pipeline failed.

  • Which data source should be used instead.

Google Cloud Data Agent Kit is designed to give coding agents access to this missing context.

It provides MCP tools and agent skills that let supported coding agents work with Google Cloud data services directly from an IDE or CLI. The current Data Agent Kit supports services including BigQuery, Bigtable, Spanner, AlloyDB, Cloud SQL, Cloud Storage, Dataflow, Managed Service for Apache Spark, Managed Service for Apache Airflow, and Knowledge Catalog.

The interesting part is not simply that an AI can generate SQL.

It is that the agent can work with real cloud data and real data infrastructure instead of relying on schemas and examples manually pasted into a prompt.

What Is Google Cloud Data Agent Kit?

Data Agent Kit is a collection of MCP tools and agent skills for data engineering and data science workflows.

Google made the kit generally available in September 2026. It is available at no additional charge, although the Google Cloud services used by the agent continue to incur their normal charges.

A simplified architecture looks like this:

Developer
    |
    v
Coding Agent
    |
    v
Data Agent Kit
    |
    +-- MCP Tools
    +-- Data Skills
    |
    v
Google Cloud
    |
    +-- BigQuery
    +-- Cloud Storage
    +-- Spanner
    +-- Bigtable
    +-- Cloud SQL
    +-- Dataflow
    +-- Spark
    +-- Airflow

Instead of manually copying information from Google Cloud into the agent's context, the agent can retrieve relevant information through its tools.

Why Does an Agent Need Access to Real Data?

Consider a developer asking:

Find the reason yesterday's sales pipeline failed.

A generic coding agent does not automatically know:

Which pipeline?
Which project?
Which dataset?
Which execution?
Which logs?
Which source files?

A data-aware agent can investigate the environment.

Conceptually:

User Request
     |
     v
Agent
     |
     +--> Discover Pipeline
     |
     +--> Inspect Recent Run
     |
     +--> Read Logs
     |
     +--> Inspect Code
     |
     +--> Query Data
     |
     v
Root Cause Summary

This is the key difference between generating code and operating with context.

Data Agent Kit Uses MCP

Model Context Protocol provides the connection between the coding agent and Google Cloud data tools.

The architecture looks like:

Coding Agent
     |
     v
MCP Client
     |
     v
Data Agent Kit MCP Tools
     |
     v
Google Cloud APIs
     |
     v
Data Resources

The agent does not need a separate custom integration for every task.

Instead, tools expose structured operations that the agent can invoke.

For example:

list datasets
inspect table schema
run query
inspect job
read pipeline logs
create pipeline
deploy workload

The available operations depend on the supported service and the user's permissions.

Which Coding Agents Can Use It?

Data Agent Kit is designed to work inside the developer environments teams already use.

Current Google documentation lists integrations including:

  • VS Code and VS Code-compatible IDEs

  • Antigravity

  • Cursor

  • Claude Code

  • Codex CLI

  • Gemini CLI

  • Google Cloud Shell

  • Google Cloud Workstations

The IDE extension provides a unified interface, while the plugin brings the tools and skills into supported coding agents.

This matters because developers do not have to move their entire workflow into a new data-specific application.

Installing Data Agent Kit

For VS Code-based environments, the extension can be installed from the Extensions panel.

Search for:

Google Cloud Data Agent Kit

For Codex CLI, Google documents the following installation flow:

codex plugin marketplace add GoogleCloudPlatform/data-agent-kit-plugin
codex plugin add dak@dak-marketplace

For Claude Code:

claude plugin marketplace add https://github.com/GoogleCloudPlatform/data-agent-kit-plugin
claude plugin install dak@dak-marketplace

Google Cloud Shell and Cloud Workstations come with Data Agent Kit preinstalled.

After installation, the agent needs Google Cloud credentials.

For the CLI plugin, Google documents using the Google Cloud CLI and Application Default Credentials:

gcloud init

gcloud auth login

gcloud auth application-default login

The important point is that the agent operates using Google Cloud identity and permissions.

Querying BigQuery With an Agent

Suppose your organization has:

Project:
analytics-prod

Dataset:
sales

Table:
orders

You could ask:

Find the top five products by revenue
for the last 30 days.

The agent can discover the relevant data and construct a query.

Conceptually:

Natural Language
       |
       v
Agent
       |
       v
Discover Schema
       |
       v
Generate SQL
       |
       v
Run Query
       |
       v
Analyze Results
       |
       v
Answer

A generated query might look like:

SELECT
    product_id,
    SUM(revenue) AS total_revenue
FROM `analytics-prod.sales.orders`
WHERE order_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
GROUP BY product_id
ORDER BY total_revenue DESC
LIMIT 5;

The important difference is that the agent can work with the actual environment rather than assuming a schema.

Why Schema Discovery Matters

Without schema access, an agent might guess:

customer_id
customerId
client_id
customer_key

Only one may exist.

With Data Agent Kit, the agent can inspect the schema first.

Agent
  |
  v
Inspect Table
  |
  +-- Column names
  +-- Data types
  +-- Partitioning
  +-- Metadata
  |
  v
Generate Query

This reduces one of the most common problems with AI-generated SQL: making assumptions about the database.

Query Cost Matters

Writing correct SQL is not enough.

Consider a partitioned BigQuery table containing several terabytes.

An agent could generate:

SELECT *
FROM `analytics-prod.sales.events`;

The query may be syntactically correct.

It may also scan far more data than necessary.

Google's Data Agent Kit skills are designed to guide agents toward cost-aware patterns, including checking partition information and using a dry run before executing BigQuery queries.

A better workflow is:

Question
   |
   v
Inspect Table
   |
   v
Check Partitioning
   |
   v
Build Filtered Query
   |
   v
Dry Run
   |
   v
Review Cost
   |
   v
Execute

This is much closer to how an experienced data engineer works.

A Safer BigQuery Workflow

Instead of allowing the agent to immediately execute every query, use:

1. Understand request
2. Discover schema
3. Check partitioning
4. Generate SQL
5. Validate SQL
6. Estimate cost
7. Execute
8. Analyze results

For example:

Developer:
Analyze monthly revenue by region.

Agent:
I found sales.orders partitioned by order_date.

I will restrict the query to the requested month
before executing it.

That is better than blindly running a full-table scan.

Working With Real Results

The agent can use query results as context for the next step.

For example:

Developer:
Which region has the highest return rate?

Agent:
I will calculate returns divided by completed orders
for each region.

The workflow becomes:

Question
   |
   v
SQL
   |
   v
Query Result
   |
   v
Agent Analysis
   |
   v
Answer

The agent can then follow up:

Why is that region different?

and generate another query.

This creates an iterative data-analysis workflow.

Querying More Than BigQuery

Data Agent Kit is not limited to BigQuery.

Google's documentation lists supported query sources including:

  • BigQuery

  • Bigtable

  • Spanner

  • AlloyDB for PostgreSQL

  • Cloud SQL for MySQL

  • Cloud SQL for PostgreSQL

The specific authentication and capabilities vary by service.

This makes the architecture useful for organizations that have a mixed Google Cloud data estate.

For example:

                   Coding Agent
                        |
                        v
                 Data Agent Kit
                        |
        +---------------+---------------+
        |               |               |
        v               v               v
    BigQuery         Spanner        Cloud SQL
        |               |               |
        +---------------+---------------+
                        |
                        v
                    Analysis

The agent can work across data sources rather than treating every database as an isolated system.

Building Data Pipelines

Data Agent Kit also goes beyond querying.

Google documents workflows for creating and deploying data engineering pipelines, including pipelines using Managed Service for Apache Spark. It can also support orchestration workflows.

A developer might ask:

Create a pipeline that reads raw sales
files from Cloud Storage, transforms them,
and writes the cleaned data to BigQuery.

The agent can work through a workflow such as:

Cloud Storage
     |
     v
Raw Data
     |
     v
Transformation
     |
     v
Quality Checks
     |
     v
BigQuery

The agent can generate pipeline code, inspect configuration, and help troubleshoot execution failures.

This is where Data Agent Kit becomes more than an SQL assistant.

Debugging Failed Pipelines

Imagine a scheduled pipeline failed overnight.

The old workflow might be:

Open Cloud Console
      |
      v
Find Pipeline
      |
      v
Find Failed Run
      |
      v
Open Logs
      |
      v
Find Error
      |
      v
Open Source Code
      |
      v
Fix

A coding agent with data tools can bring those steps into one workflow.

Developer
   |
   v
"Why did the pipeline fail last night?"
   |
   v
Agent
   |
   +-- Find pipeline
   +-- Inspect failed run
   +-- Read logs
   +-- Inspect source
   +-- Identify likely cause
   |
   v
Suggested Fix

The agent can then help modify the pipeline code.

That is a meaningful productivity improvement because the agent has access to both the code and the environment where the code runs.

Working With Data Lineage

Knowing that a table exists is not always enough.

A developer may need to know:

Where did this table come from?
What pipeline writes it?
Which reports depend on it?
Which source table feeds the pipeline?

Knowledge Catalog can provide data discovery, quality, and lineage information within the broader Data Agent Kit workflow.

A conceptual workflow is:

Business Table
      |
      v
Lineage
      |
      +---- Source Table
      |
      +---- Transformation
      |
      +---- Pipeline
      |
      +---- Downstream Dataset

This gives the agent more context when investigating data problems.

From Data Question to Visualization

A useful workflow does not have to stop at SQL.

Consider:

"What happened to sales in Europe?"

The agent could:

1. Find the relevant dataset.
2. Inspect the schema.
3. Query sales by month.
4. Compare regions.
5. Identify an unusual trend.
6. Create a visualization.

The process becomes:

Natural Language
       |
       v
Data Discovery
       |
       v
SQL
       |
       v
Results
       |
       v
Analysis
       |
       v
Visualization

This shortens the path between a business question and a data-backed answer.

Agent Skills Are Important

MCP tools provide capabilities.

Skills provide guidance about how those capabilities should be used.

For example:

MCP Tool
   |
   v
"Run BigQuery query"

does not necessarily tell the agent:

Check partitioning first.
Avoid SELECT *.
Use a dry run.
Limit the result set.

A skill can provide those instructions.

Google describes Data Agent Kit as combining MCP tools with Google-authored skills that encode data best practices.

The architecture therefore becomes:

Coding Agent
      |
      +---- MCP Tools
      |
      +---- Data Skills
      |
      v
Google Cloud Data

This combination is important.

A powerful tool without good operating guidance can still produce poor results.

Security and IAM

Giving an AI agent access to real data creates obvious security questions.

The agent should not receive unrestricted permissions.

Data Agent Kit uses the authenticated user or an impersonated service account, so existing Google Cloud authorization controls continue to apply. Google also documents governance options such as IAM, Model Armor, VPC Service Controls, and Principal Access Boundary policies.

The basic model is:

Developer Identity
       |
       v
Google Cloud IAM
       |
       v
Data Agent Kit
       |
       v
Authorized Resources

If the user cannot access a dataset, the agent should not receive elevated permissions simply because it is an AI agent.

Row-Level and Column-Level Security

Enterprise data often has more granular controls.

A user may be allowed to query a table but only see certain rows or columns.

For example:

Sales Table
   |
   +-- Region
   +-- Customer
   +-- Revenue
   +-- Internal Cost

A regional manager may be allowed to see:

Region = North

but not:

Region = South

Likewise, sensitive columns may be restricted.

Because the agent operates using the user's identity or an impersonated identity, existing data-access policies can remain part of the security boundary.

This is far safer than creating a separate privileged AI account with access to everything.

Agent Security Is Still Your Responsibility

Even with IAM, the agent itself is another layer that needs governance.

Potential risks include:

  • Destructive SQL

  • Expensive queries

  • Sensitive data exposure

  • Accidental writes

  • Prompt injection

  • Incorrect transformations

  • Unintended pipeline deployments

A production workflow should distinguish between:

Read
  |
  v
Analyze

and:

Modify
  |
  v
Deploy

The second category deserves stronger controls.

Use Human Approval for High-Impact Operations

A useful architecture is:

Developer Request
      |
      v
Agent Analysis
      |
      v
Proposed Change
      |
      v
Human Review
      |
      v
Approved
      |
      v
Deployment

For example:

Agent:
"I found that the pipeline scans the entire
events table. I propose adding a partition filter."

Developer:
"Apply the change."

Agent:
"Updated pipeline code and prepared the change
for review."

The final deployment can still pass through GitHub Actions or another approved CI/CD process.

Google's pipeline documentation supports deployment workflows that can use GitHub Actions.

Common Mistakes

Giving the Agent Owner-Level Access

Do not solve permission problems by giving the agent excessive IAM privileges.

Start with the minimum required roles.

Running Queries Without Checking Cost

A correct query can still be an expensive query.

For large BigQuery tables, inspect partitioning and estimate query cost before execution.

Trusting Generated SQL

AI-generated SQL should be reviewed like any other generated code.

Check joins, filters, aggregations, and data semantics.

Allowing Direct Production Deployment

Keep production changes behind established review and deployment controls.

Ignoring Sensitive Data

An agent that can query a database can potentially expose the returned information in its response.

Treat agent output as another data-access surface.

Treating MCP as a Security Boundary by Itself

MCP connects the agent to tools.

It does not replace IAM, data governance, network controls, or application authorization.

Troubleshooting

The Agent Cannot Find a Dataset

Check:

  1. Active Google Cloud project.

  2. Authentication.

  3. IAM permissions.

  4. Dataset location.

  5. Service availability.

The agent may be authenticated correctly but still lack permission to access the resource.

Query Generation Is Wrong

Ask the agent to inspect the schema first.

For example:

Before writing the query, inspect the table schema
and partitioning information.

This reduces assumptions.

Query Is Too Expensive

Check:

Partition Filter
Column Selection
Join Conditions
Result Limit

Then use a dry run where supported before executing the query.

Pipeline Debugging Is Incomplete

Ask the agent to inspect both:

Pipeline Source Code
+
Execution Logs

Looking at only the code may miss runtime configuration or data-related failures.

The Agent Cannot Deploy a Pipeline

Check the deployment path and required permissions.

Google documents that automated deployment of orchestration pipelines currently works with GitHub Actions, so deployment architecture matters.

Data Agent Kit vs Traditional Data Tools

Area

Traditional Workflow

Data Agent Kit

SQL generation

Developer writes SQL

Agent can generate SQL

Schema discovery

Manual

Agent can inspect

Query execution

Separate console/tool

Agent workflow

Pipeline development

IDE + cloud console

Coding agent + cloud tools

Debugging

Logs and code separately

Agent can connect both

Data lineage

Separate tools

Integrated workflow

Best-practice guidance

Team knowledge

Agent skills

Human control

High

Configurable

Automation

Script-based

Natural-language + tools

Data Agent Kit does not eliminate traditional data tools.

It changes how developers interact with them.

Best Practices

  1. Use least-privilege Google Cloud IAM permissions.

  2. Let the agent inspect schemas before generating SQL.

  3. Check partitioning before querying large BigQuery tables.

  4. Use dry runs or other cost checks before expensive queries.

  5. Review generated SQL before running high-impact operations.

  6. Keep production writes and deployments behind human approval.

  7. Use the developer's identity or a narrowly scoped service account rather than a global administrator account.

  8. Protect sensitive query results from unnecessary exposure.

  9. Use lineage and data-quality information when investigating data problems.

  10. Keep pipeline source code in version control.

  11. Use CI/CD for production deployment rather than allowing unrestricted agent deployment.

  12. Test agent workflows against representative datasets before using them in production.

When Should You Use Data Agent Kit?

Data Agent Kit is a strong fit when developers regularly move between:

IDE
 |
 +-- SQL
 +-- Data
 +-- Pipelines
 +-- Logs
 +-- Cloud Resources
 +-- ML

It is particularly useful for:

  • BigQuery analysis

  • Data pipeline development

  • Pipeline troubleshooting

  • Data discovery

  • Schema exploration

  • Data transformation

  • Spark development

  • Data quality investigation

  • ML workflows

  • Cross-service data engineering

It may be unnecessary when the task is a simple SQL query that a developer can already run safely and efficiently.

Summary

Google Cloud Data Agent Kit gives supported coding agents access to Google Cloud data tools and data-engineering skills through MCP.

It can help agents discover schemas, run queries, analyze data, build pipelines, inspect execution information, work with data services, and assist with ML and data workflows.

Its biggest advantage is context.

The agent can work with the actual cloud environment instead of relying entirely on information pasted into a prompt.

For production use, the important safeguards are IAM, least privilege, query-cost controls, sensitive-data protection, human review, version control, and controlled deployment.

The real shift is simple:

AI agents are moving from writing data code to working with the data systems that code operates on.

That is where tools such as Data Agent Kit become much more interesting for developers.