AI agents are becoming a new way for people to work with data.
Instead of opening a dashboard, writing SQL, or navigating a catalog, a user can ask a question in natural language:
Show me the revenue for my region for the last quarter.The agent can translate that request into a data query, call the appropriate tools, retrieve the results, and return an answer.
The difficult part starts when different users have different permissions.
A finance manager may be allowed to see revenue for several regions. A regional manager may only see data for one region. Another employee may not be allowed to see customer-level information at all.
With a traditional application, these rules can be enforced through the application's identity and authorization layer. With an AI data agent, there is another component in the middle: the agent itself.
AWS has now documented an identity-aware pattern for AI data agents using Amazon Bedrock AgentCore, AWS IAM Identity Center, AWS Lake Formation, AWS Lambda, and Trusted Identity Propagation. The approach allows Lake Formation to evaluate the permissions of the actual user rather than simply seeing the IAM role used by the agent. AWS published the pattern on October 6, 2026.
That distinction is important because it lets organizations keep data authorization in the data-governance layer instead of duplicating access-control logic inside every AI agent.
The Problem With a Shared Agent Role
Consider a data agent used by 500 employees.
The basic architecture might look like this:
User
|
v
AI Agent
|
v
IAM Role
|
v
Lake Formation
|
v
DataThe agent needs permissions to query the data.
So the agent runs using an IAM role with access to the required tables.
The problem is that Lake Formation sees the agent's role.
It does not automatically know which human initiated the request.
That creates an authorization problem.
Suppose Alice is allowed to see:
Sales
Marketing
Financewhile Bob is allowed to see only:
SalesIf both requests reach Lake Formation through the same agent role, the data layer sees essentially the same principal.
The application then has two choices.
It can give the shared agent role broad access and implement authorization inside the application, or it can restrict the role so heavily that the agent becomes less useful.
Neither approach is particularly attractive.
The first moves authorization logic away from the data-governance layer.
The second limits self-service analytics.
The Identity-Aware Approach
AWS's new pattern changes the request flow.
Instead of asking Lake Formation to authorize the agent's identity, the architecture propagates the user's identity through the agent workflow.
The simplified architecture becomes:
User
|
v
Identity Provider
|
v
AI Agent
|
v
Tool / Lambda
|
v
Trusted Identity Propagation
|
v
Lake Formation
|
v
User's Data PermissionsThe key idea is that the AI agent does not become the authorization subject.
The actual user remains the authorization subject.
AWS calls this Trusted Identity Propagation, or TIP. AWS IAM Identity Center provides the identity context, while Lake Formation evaluates the resulting user identity against its existing data permissions.
This means organizations can keep their existing Lake Formation governance model instead of creating a separate permission system specifically for AI agents.
Lake Formation Already Provides Fine-Grained Data Controls
AWS Lake Formation is designed to govern access to data lakes and the AWS Glue Data Catalog.
Its permission model can control access at several levels, including:
Catalog
Database
Table
Column
Row
CellLake Formation can also use data filters and tag-based access control for more complex governance requirements.
For example, an organization could have a table containing:
EmployeeId
Department
Region
CustomerName
Revenue
SalaryOne user might receive access to:
Region
Revenuewhile another user may receive:
Region
Revenue
CustomerNameThe agent should not have to reproduce those rules in its own code.
Ideally, the query reaches the governed data layer with the user's identity attached, and Lake Formation makes the authorization decision.
That is what the new pattern is designed to achieve.
How the Request Flows
The full request path is more involved than simply passing a user ID through an API.
AWS's reference architecture uses an OpenID Connect identity provider, Amazon Bedrock AgentCore Runtime, AgentCore Gateway, Lambda, IAM Identity Center, and Lake Formation.
A simplified version looks like this:
User
|
| OIDC authentication
v
Identity Provider
|
| access_token + id_token
v
AgentCore Runtime
|
| identity headers
v
AI Agent
|
| MCP connection
v
AgentCore Gateway
|
v
Lambda
|
| identity token exchange
v
IAM Identity Center
|
| identity context
v
Lake Formation
|
v
Athena
|
v
Governed DataThere is an important security detail here.
The user's identity token should not become part of the model's context.
The foundation model should not receive the token as a tool argument or prompt value.
Instead, AWS's pattern transports the identity information through HTTP headers outside the model's reasoning path.
That is a much cleaner boundary.
Why the Token Should Stay Outside the Model
An AI model processes context.
Anything placed inside that context can potentially become part of the model's input, traces, tool arguments, or other agent state depending on the architecture.
An identity token is not something the model needs to reason about.
The model needs to know:
Query the sales data.It does not need to know:
Here is the user's OAuth identity token.The AWS pattern therefore keeps the identity token at the transport layer.
Conceptually:
HTTP Layer
--------------------------------
Authorization: access_token
Custom identity header: id_token
--------------------------------
|
v
AI Agent
|
v
Model ContextThe model receives the user's request and tool definitions, but the identity token remains outside that reasoning context. AWS explicitly describes this as a security property of the architecture.
This separation becomes especially important when the agent can invoke multiple tools.
AgentCore Runtime Carries the Identity
The first important hop is Bedrock AgentCore Runtime.
AWS's pattern configures a request-header allow list so the required headers can reach the agent container. The standard authorization header carries the access token, while a custom header carries the identity token.
Conceptually, the configuration looks like:
requestHeaderConfiguration:
requestHeaderAllowlist:
- Authorization
- X-Amzn-Bedrock-AgentCore-Runtime-Custom-IdTokenThe runtime can authenticate the incoming request while also making the identity information available to the agent's surrounding infrastructure.
The agent itself does not need to inspect the token and decide what the user can access.
That responsibility stays downstream.
MCP Gateway Carries the Identity Further
The agent can connect to an AgentCore Gateway using Model Context Protocol.
The identity information is propagated along that connection rather than being included in the tool's business arguments.
The request flow becomes:
Agent
|
| MCP
v
AgentCore Gateway
|
v
Lambda ToolThe gateway can be configured to allow the required identity header to reach the Lambda target. AWS documents this as another part of the identity chain.
This is an important design pattern for agent systems.
Authentication information should travel through trusted infrastructure boundaries without becoming application-level data that the model has to manipulate.
Lambda Performs the Identity Exchange
The Lambda function is where the identity context becomes useful for the data query.
The function receives the propagated identity information, validates it, and performs a server-side token exchange.
The resulting identity context is then used when assuming an IAM role for the data operation. AWS calls this role the Trusted Identity Propagation role, or TIP role.
The interesting part is that this role does not carry the user's Lake Formation data grants.
Instead, it provides the AWS permissions required to call services such as Athena and Lake Formation.
The propagated user identity determines what data the user can actually access.
Conceptually:
TIP Role
|
+-- Can call required AWS APIs
|
+-- Does NOT define user's data grants
|
v
Propagated User Identity
|
v
Lake Formation
|
v
Data PermissionsThat separation is what makes the architecture useful.
Lake Formation Makes the Final Decision
Suppose the organization has already granted a user SELECT permission on a table.
The agent does not need a second authorization system.
The Lambda function executes the query using the propagated identity context.
Lake Formation evaluates the user's grants and applies the existing policies.
For example:
Alice
|
+-- Sales table: SELECT
+-- Finance table: SELECT
+-- Salary column: DENYWhen Alice asks:
Show me sales by region.the query can return the permitted data.
If she asks for salary information that she is not authorized to access, Lake Formation remains the enforcement point.
AWS specifically notes that existing Lake Formation grants can continue to be used without modifying them to account for the AI agent.
That is one of the strongest parts of the design.
Row-Level and Column-Level Security Still Matter
The value becomes even clearer when an organization uses fine-grained policies.
Imagine a customer table:
CustomerId
CustomerName
Region
Revenue
Email
PhoneA regional manager might be allowed to see only customers from their region.
A support employee might see:
CustomerId
CustomerName
Regionbut not:
Revenue
Email
PhoneThe agent should not need to implement those restrictions itself.
Lake Formation supports fine-grained controls, including row- and cell-level security, so the data layer can continue enforcing the organization's governance rules.
This reduces the amount of security-sensitive logic that needs to be embedded in agent code.
CloudTrail Provides the Audit Trail
Authorization is only half of the problem.
Organizations also need to know who accessed what.
That becomes difficult if every query appears to originate from a shared agent role.
The identity-aware pattern addresses this by allowing CloudTrail to record the human identity associated with the assumed role session.
AWS's example uses an onBehalfOf field in CloudTrail to identify the actual IAM Identity Center user responsible for the access.
The resulting audit trail can conceptually look like:
Time
|
User: Alice
|
Agent: Data Assistant
|
Tool: Athena Query
|
Table: Sales
|
Result: AuthorizedThat is much more useful during a security investigation than:
User: DataAgentRolefor every request.
Existing Lake Formation Policies Can Be Reused
One of the biggest practical advantages is that organizations do not need to create a separate permission model for AI agents.
If Lake Formation already contains policies such as:
Finance Group -> Finance tables
Sales Group -> Sales tables
HR Group -> HR tablesthe AI agent can use the same governance model.
AWS documents IAM Identity Center integration with Lake Formation specifically for granting permissions to users and groups on Data Catalog resources.
This is preferable to building something like:
if (user.Department == "Finance")
{
// Allow finance query
}inside every AI application.
That type of application-level authorization becomes difficult to maintain as the number of datasets, teams, and agents grows.
A Practical Agent Architecture
A production data agent could be structured like this:
User
|
v
Identity Provider
|
v
Web Application
|
v
Bedrock AgentCore Runtime
|
v
AI Agent
|
MCP Gateway
|
v
Lambda
|
Token Exchange
|
v
Identity Context
|
v
Athena / Glue
|
v
Lake Formation
|
v
S3 Data LakeThe model's responsibility is primarily reasoning.
The application's responsibility is orchestration.
The identity infrastructure is responsible for authentication and propagation.
Lake Formation remains responsible for data authorization.
That separation makes the overall architecture easier to reason about.
What Developers Should Not Do
There are several tempting shortcuts that should be avoided.
Do Not Put Authorization Logic in the Prompt
This is not a secure authorization mechanism:
You are allowed to show Finance data
but never show HR data.A prompt is not an access-control boundary.
Authorization should happen at the data layer.
Do Not Give Every Agent Full Data Access
Giving an agent unrestricted access and relying on the model to decide what information it should reveal is fundamentally unsafe.
The model should not be the final authorization authority.
Do Not Pass Tokens as Tool Arguments
Avoid schemas such as:
{
"query": "...",
"accessToken": "..."
}The token is authentication material, not business data.
Keeping it outside the model and tool argument path reduces unnecessary exposure.
Do Not Rebuild Lake Formation Permissions in Application Code
If Lake Formation already knows which users can access which tables, columns, or rows, use that governance layer.
Duplicating those policies in every agent creates another source of truth.
Advantages and Disadvantages
Advantages
Existing data governance can be reused. AI agents can work with Lake Formation permissions without creating a separate authorization model specifically for AI applications.
The real user can remain the authorization subject. Lake Formation evaluates the propagated user identity rather than simply trusting the agent's service role.
Identity stays outside model context. The identity token is propagated through the transport layer rather than being exposed as model input or a tool argument.
Auditing becomes more meaningful. CloudTrail can record the human user associated with the data access through the onBehalfOf identity information.
Fine-grained governance remains possible. Existing Lake Formation controls for tables, columns, rows, and cells can continue to enforce data access.
Disadvantages
The architecture is more complicated. OIDC, IAM Identity Center, AgentCore Runtime, AgentCore Gateway, Lambda, token exchange, and Lake Formation all have to be configured correctly.
Identity propagation must be protected across every hop. A mistake in header forwarding or token handling can break the trust chain.
The model still needs careful tool design. Keeping identity outside model context does not automatically make an agent safe. Tool permissions and query construction still need review.
The solution is closely tied to AWS services. Organizations using another identity platform or data-governance stack may need a different implementation.
When This Pattern Makes Sense
This architecture is particularly useful when an organization is building AI agents that query governed enterprise data.
Good candidates include:
Enterprise data assistants
Natural-language analytics
Financial reporting agents
Customer-support data agents
Internal business intelligence agents
Multi-tenant analytics systems
AI applications using Lake Formation-governed dataIt is less useful for a small application where every user has identical access to a small dataset.
In that situation, the additional identity propagation infrastructure may not justify its complexity.
The pattern becomes valuable when the organization already has meaningful data-governance requirements and wants AI agents to operate within those same boundaries.
Summary
AI data agents create a security problem that traditional applications do not always have.
The agent needs access to the data, but the data layer still needs to know which human is making the request.
AWS's identity-aware pattern for Lake Formation addresses that problem by propagating the user's identity through Bedrock AgentCore, MCP-based tool access, Lambda, and AWS IAM Identity Center until the request reaches the governed data layer.
Lake Formation can then evaluate the user's existing permissions rather than trusting a shared agent role.
The architecture also keeps the identity token outside the foundation model's context and provides an audit path through CloudTrail.
That separation is the important part.
The model decides how to solve the user's request. The agent decides which tools to use. The identity system establishes who the user is. Lake Formation decides what data that user is allowed to access.
For enterprise AI agents, keeping those responsibilities separate is much safer than asking the model or application code to become the final authority on data access.

Join the conversation! Your thoughts help the community grow.