Coding agents have become good at writing SQL, Python, PySpark, notebooks, and pipeline code.
The problem is that code generation is only one part of data engineering.
An agent can write:
SELECT *
FROM sales.orders
WHERE status = 'pending';But it may not know:
Which project contains the table.
Which dataset is trusted.
How the table is partitioned.
Which columns contain sensitive data.
Whether the query will scan terabytes.
Why yesterday's pipeline failed.
Which data source should be used instead.
Google Cloud Data Agent Kit is designed to give coding agents access to this missing context.
It provides MCP tools and agent skills that let supported coding agents work with Google Cloud data services directly from an IDE or CLI. The current Data Agent Kit supports services including BigQuery, Bigtable, Spanner, AlloyDB, Cloud SQL, Cloud Storage, Dataflow, Managed Service for Apache Spark, Managed Service for Apache Airflow, and Knowledge Catalog.
The interesting part is not simply that an AI can generate SQL.
It is that the agent can work with real cloud data and real data infrastructure instead of relying on schemas and examples manually pasted into a prompt.
What Is Google Cloud Data Agent Kit?
Data Agent Kit is a collection of MCP tools and agent skills for data engineering and data science workflows.
Google made the kit generally available in September 2026. It is available at no additional charge, although the Google Cloud services used by the agent continue to incur their normal charges.
A simplified architecture looks like this:
Developer
|
v
Coding Agent
|
v
Data Agent Kit
|
+-- MCP Tools
+-- Data Skills
|
v
Google Cloud
|
+-- BigQuery
+-- Cloud Storage
+-- Spanner
+-- Bigtable
+-- Cloud SQL
+-- Dataflow
+-- Spark
+-- AirflowInstead of manually copying information from Google Cloud into the agent's context, the agent can retrieve relevant information through its tools.
Why Does an Agent Need Access to Real Data?
Consider a developer asking:
Find the reason yesterday's sales pipeline failed.A generic coding agent does not automatically know:
Which pipeline?
Which project?
Which dataset?
Which execution?
Which logs?
Which source files?A data-aware agent can investigate the environment.
Conceptually:
User Request
|
v
Agent
|
+--> Discover Pipeline
|
+--> Inspect Recent Run
|
+--> Read Logs
|
+--> Inspect Code
|
+--> Query Data
|
v
Root Cause SummaryThis is the key difference between generating code and operating with context.
Data Agent Kit Uses MCP
Model Context Protocol provides the connection between the coding agent and Google Cloud data tools.
The architecture looks like:
Coding Agent
|
v
MCP Client
|
v
Data Agent Kit MCP Tools
|
v
Google Cloud APIs
|
v
Data ResourcesThe agent does not need a separate custom integration for every task.
Instead, tools expose structured operations that the agent can invoke.
For example:
list datasets
inspect table schema
run query
inspect job
read pipeline logs
create pipeline
deploy workloadThe available operations depend on the supported service and the user's permissions.
Which Coding Agents Can Use It?
Data Agent Kit is designed to work inside the developer environments teams already use.
Current Google documentation lists integrations including:
VS Code and VS Code-compatible IDEs
Antigravity
Cursor
Claude Code
Codex CLI
Gemini CLI
Google Cloud Shell
Google Cloud Workstations
The IDE extension provides a unified interface, while the plugin brings the tools and skills into supported coding agents.
This matters because developers do not have to move their entire workflow into a new data-specific application.
Installing Data Agent Kit
For VS Code-based environments, the extension can be installed from the Extensions panel.
Search for:
Google Cloud Data Agent KitFor Codex CLI, Google documents the following installation flow:
codex plugin marketplace add GoogleCloudPlatform/data-agent-kit-plugin
codex plugin add dak@dak-marketplaceFor Claude Code:
claude plugin marketplace add https://github.com/GoogleCloudPlatform/data-agent-kit-plugin
claude plugin install dak@dak-marketplaceGoogle Cloud Shell and Cloud Workstations come with Data Agent Kit preinstalled.
After installation, the agent needs Google Cloud credentials.
For the CLI plugin, Google documents using the Google Cloud CLI and Application Default Credentials:
gcloud init
gcloud auth login
gcloud auth application-default loginThe important point is that the agent operates using Google Cloud identity and permissions.
Querying BigQuery With an Agent
Suppose your organization has:
Project:
analytics-prod
Dataset:
sales
Table:
ordersYou could ask:
Find the top five products by revenue
for the last 30 days.The agent can discover the relevant data and construct a query.
Conceptually:
Natural Language
|
v
Agent
|
v
Discover Schema
|
v
Generate SQL
|
v
Run Query
|
v
Analyze Results
|
v
AnswerA generated query might look like:
SELECT
product_id,
SUM(revenue) AS total_revenue
FROM `analytics-prod.sales.orders`
WHERE order_date >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
GROUP BY product_id
ORDER BY total_revenue DESC
LIMIT 5;The important difference is that the agent can work with the actual environment rather than assuming a schema.
Why Schema Discovery Matters
Without schema access, an agent might guess:
customer_id
customerId
client_id
customer_keyOnly one may exist.
With Data Agent Kit, the agent can inspect the schema first.
Agent
|
v
Inspect Table
|
+-- Column names
+-- Data types
+-- Partitioning
+-- Metadata
|
v
Generate QueryThis reduces one of the most common problems with AI-generated SQL: making assumptions about the database.
Query Cost Matters
Writing correct SQL is not enough.
Consider a partitioned BigQuery table containing several terabytes.
An agent could generate:
SELECT *
FROM `analytics-prod.sales.events`;The query may be syntactically correct.
It may also scan far more data than necessary.
Google's Data Agent Kit skills are designed to guide agents toward cost-aware patterns, including checking partition information and using a dry run before executing BigQuery queries.
A better workflow is:
Question
|
v
Inspect Table
|
v
Check Partitioning
|
v
Build Filtered Query
|
v
Dry Run
|
v
Review Cost
|
v
ExecuteThis is much closer to how an experienced data engineer works.
A Safer BigQuery Workflow
Instead of allowing the agent to immediately execute every query, use:
1. Understand request
2. Discover schema
3. Check partitioning
4. Generate SQL
5. Validate SQL
6. Estimate cost
7. Execute
8. Analyze resultsFor example:
Developer:
Analyze monthly revenue by region.
Agent:
I found sales.orders partitioned by order_date.
I will restrict the query to the requested month
before executing it.That is better than blindly running a full-table scan.
Working With Real Results
The agent can use query results as context for the next step.
For example:
Developer:
Which region has the highest return rate?
Agent:
I will calculate returns divided by completed orders
for each region.The workflow becomes:
Question
|
v
SQL
|
v
Query Result
|
v
Agent Analysis
|
v
AnswerThe agent can then follow up:
Why is that region different?and generate another query.
This creates an iterative data-analysis workflow.
Querying More Than BigQuery
Data Agent Kit is not limited to BigQuery.
Google's documentation lists supported query sources including:
BigQuery
Bigtable
Spanner
AlloyDB for PostgreSQL
Cloud SQL for MySQL
Cloud SQL for PostgreSQL
The specific authentication and capabilities vary by service.
This makes the architecture useful for organizations that have a mixed Google Cloud data estate.
For example:
Coding Agent
|
v
Data Agent Kit
|
+---------------+---------------+
| | |
v v v
BigQuery Spanner Cloud SQL
| | |
+---------------+---------------+
|
v
AnalysisThe agent can work across data sources rather than treating every database as an isolated system.
Building Data Pipelines
Data Agent Kit also goes beyond querying.
Google documents workflows for creating and deploying data engineering pipelines, including pipelines using Managed Service for Apache Spark. It can also support orchestration workflows.
A developer might ask:
Create a pipeline that reads raw sales
files from Cloud Storage, transforms them,
and writes the cleaned data to BigQuery.The agent can work through a workflow such as:
Cloud Storage
|
v
Raw Data
|
v
Transformation
|
v
Quality Checks
|
v
BigQueryThe agent can generate pipeline code, inspect configuration, and help troubleshoot execution failures.
This is where Data Agent Kit becomes more than an SQL assistant.
Debugging Failed Pipelines
Imagine a scheduled pipeline failed overnight.
The old workflow might be:
Open Cloud Console
|
v
Find Pipeline
|
v
Find Failed Run
|
v
Open Logs
|
v
Find Error
|
v
Open Source Code
|
v
FixA coding agent with data tools can bring those steps into one workflow.
Developer
|
v
"Why did the pipeline fail last night?"
|
v
Agent
|
+-- Find pipeline
+-- Inspect failed run
+-- Read logs
+-- Inspect source
+-- Identify likely cause
|
v
Suggested FixThe agent can then help modify the pipeline code.
That is a meaningful productivity improvement because the agent has access to both the code and the environment where the code runs.
Working With Data Lineage
Knowing that a table exists is not always enough.
A developer may need to know:
Where did this table come from?
What pipeline writes it?
Which reports depend on it?
Which source table feeds the pipeline?Knowledge Catalog can provide data discovery, quality, and lineage information within the broader Data Agent Kit workflow.
A conceptual workflow is:
Business Table
|
v
Lineage
|
+---- Source Table
|
+---- Transformation
|
+---- Pipeline
|
+---- Downstream DatasetThis gives the agent more context when investigating data problems.
From Data Question to Visualization
A useful workflow does not have to stop at SQL.
Consider:
"What happened to sales in Europe?"The agent could:
1. Find the relevant dataset.
2. Inspect the schema.
3. Query sales by month.
4. Compare regions.
5. Identify an unusual trend.
6. Create a visualization.The process becomes:
Natural Language
|
v
Data Discovery
|
v
SQL
|
v
Results
|
v
Analysis
|
v
VisualizationThis shortens the path between a business question and a data-backed answer.
Agent Skills Are Important
MCP tools provide capabilities.
Skills provide guidance about how those capabilities should be used.
For example:
MCP Tool
|
v
"Run BigQuery query"does not necessarily tell the agent:
Check partitioning first.
Avoid SELECT *.
Use a dry run.
Limit the result set.A skill can provide those instructions.
Google describes Data Agent Kit as combining MCP tools with Google-authored skills that encode data best practices.
The architecture therefore becomes:
Coding Agent
|
+---- MCP Tools
|
+---- Data Skills
|
v
Google Cloud DataThis combination is important.
A powerful tool without good operating guidance can still produce poor results.
Security and IAM
Giving an AI agent access to real data creates obvious security questions.
The agent should not receive unrestricted permissions.
Data Agent Kit uses the authenticated user or an impersonated service account, so existing Google Cloud authorization controls continue to apply. Google also documents governance options such as IAM, Model Armor, VPC Service Controls, and Principal Access Boundary policies.
The basic model is:
Developer Identity
|
v
Google Cloud IAM
|
v
Data Agent Kit
|
v
Authorized ResourcesIf the user cannot access a dataset, the agent should not receive elevated permissions simply because it is an AI agent.
Row-Level and Column-Level Security
Enterprise data often has more granular controls.
A user may be allowed to query a table but only see certain rows or columns.
For example:
Sales Table
|
+-- Region
+-- Customer
+-- Revenue
+-- Internal CostA regional manager may be allowed to see:
Region = Northbut not:
Region = SouthLikewise, sensitive columns may be restricted.
Because the agent operates using the user's identity or an impersonated identity, existing data-access policies can remain part of the security boundary.
This is far safer than creating a separate privileged AI account with access to everything.
Agent Security Is Still Your Responsibility
Even with IAM, the agent itself is another layer that needs governance.
Potential risks include:
Destructive SQL
Expensive queries
Sensitive data exposure
Accidental writes
Prompt injection
Incorrect transformations
Unintended pipeline deployments
A production workflow should distinguish between:
Read
|
v
Analyzeand:
Modify
|
v
DeployThe second category deserves stronger controls.
Use Human Approval for High-Impact Operations
A useful architecture is:
Developer Request
|
v
Agent Analysis
|
v
Proposed Change
|
v
Human Review
|
v
Approved
|
v
DeploymentFor example:
Agent:
"I found that the pipeline scans the entire
events table. I propose adding a partition filter."
Developer:
"Apply the change."
Agent:
"Updated pipeline code and prepared the change
for review."The final deployment can still pass through GitHub Actions or another approved CI/CD process.
Google's pipeline documentation supports deployment workflows that can use GitHub Actions.
Common Mistakes
Giving the Agent Owner-Level Access
Do not solve permission problems by giving the agent excessive IAM privileges.
Start with the minimum required roles.
Running Queries Without Checking Cost
A correct query can still be an expensive query.
For large BigQuery tables, inspect partitioning and estimate query cost before execution.
Trusting Generated SQL
AI-generated SQL should be reviewed like any other generated code.
Check joins, filters, aggregations, and data semantics.
Allowing Direct Production Deployment
Keep production changes behind established review and deployment controls.
Ignoring Sensitive Data
An agent that can query a database can potentially expose the returned information in its response.
Treat agent output as another data-access surface.
Treating MCP as a Security Boundary by Itself
MCP connects the agent to tools.
It does not replace IAM, data governance, network controls, or application authorization.
Troubleshooting
The Agent Cannot Find a Dataset
Check:
Active Google Cloud project.
Authentication.
IAM permissions.
Dataset location.
Service availability.
The agent may be authenticated correctly but still lack permission to access the resource.
Query Generation Is Wrong
Ask the agent to inspect the schema first.
For example:
Before writing the query, inspect the table schema
and partitioning information.This reduces assumptions.
Query Is Too Expensive
Check:
Partition Filter
Column Selection
Join Conditions
Result LimitThen use a dry run where supported before executing the query.
Pipeline Debugging Is Incomplete
Ask the agent to inspect both:
Pipeline Source Code
+
Execution LogsLooking at only the code may miss runtime configuration or data-related failures.
The Agent Cannot Deploy a Pipeline
Check the deployment path and required permissions.
Google documents that automated deployment of orchestration pipelines currently works with GitHub Actions, so deployment architecture matters.
Data Agent Kit vs Traditional Data Tools
Area | Traditional Workflow | Data Agent Kit |
|---|---|---|
SQL generation | Developer writes SQL | Agent can generate SQL |
Schema discovery | Manual | Agent can inspect |
Query execution | Separate console/tool | Agent workflow |
Pipeline development | IDE + cloud console | Coding agent + cloud tools |
Debugging | Logs and code separately | Agent can connect both |
Data lineage | Separate tools | Integrated workflow |
Best-practice guidance | Team knowledge | Agent skills |
Human control | High | Configurable |
Automation | Script-based | Natural-language + tools |
Data Agent Kit does not eliminate traditional data tools.
It changes how developers interact with them.
Best Practices
Use least-privilege Google Cloud IAM permissions.
Let the agent inspect schemas before generating SQL.
Check partitioning before querying large BigQuery tables.
Use dry runs or other cost checks before expensive queries.
Review generated SQL before running high-impact operations.
Keep production writes and deployments behind human approval.
Use the developer's identity or a narrowly scoped service account rather than a global administrator account.
Protect sensitive query results from unnecessary exposure.
Use lineage and data-quality information when investigating data problems.
Keep pipeline source code in version control.
Use CI/CD for production deployment rather than allowing unrestricted agent deployment.
Test agent workflows against representative datasets before using them in production.
When Should You Use Data Agent Kit?
Data Agent Kit is a strong fit when developers regularly move between:
IDE
|
+-- SQL
+-- Data
+-- Pipelines
+-- Logs
+-- Cloud Resources
+-- MLIt is particularly useful for:
BigQuery analysis
Data pipeline development
Pipeline troubleshooting
Data discovery
Schema exploration
Data transformation
Spark development
Data quality investigation
ML workflows
Cross-service data engineering
It may be unnecessary when the task is a simple SQL query that a developer can already run safely and efficiently.
Summary
Google Cloud Data Agent Kit gives supported coding agents access to Google Cloud data tools and data-engineering skills through MCP.
It can help agents discover schemas, run queries, analyze data, build pipelines, inspect execution information, work with data services, and assist with ML and data workflows.
Its biggest advantage is context.
The agent can work with the actual cloud environment instead of relying entirely on information pasted into a prompt.
For production use, the important safeguards are IAM, least privilege, query-cost controls, sensitive-data protection, human review, version control, and controlled deployment.
The real shift is simple:
AI agents are moving from writing data code to working with the data systems that code operates on.
That is where tools such as Data Agent Kit become much more interesting for developers.

Join the conversation! Your thoughts help the community grow.