
Where Security Belongs in the AI Agent Stack: A Practical Guide for Developers and Architects
An AI agent can understand an assignment, write a convincing plan, and select a useful tool. It can still attempt an action that it should never be allowed to perform.
For developers and architects, this creates a fundamental question: Where should the final permission decision happen?
NVIDIA addresses this in its article, Where Security Fits in an AI Agent Stack. Its central distinction is between components that guide an agent’s behavior and systems that independently enforce its authority. Models and harnesses help determine what an agent attempts; the execution environment must determine what it is permitted to do. NVIDIA Technical Blog
That distinction has practical consequences for anyone building coding assistants, customer-service agents, enterprise automation, or systems that coordinate multiple agents.
The examples below translate it into implementation guidance. They are illustrative designs, not claims about deployments NVIDIA has tested.
A Prompt Is Not an Access-Control Policy
Imagine giving an agent this task:
Investigate why a customer’s order has not shipped.
The agent might need to read order status, inspect shipment events, and prepare a response. It does not necessarily need permission to change the delivery address, issue a refund, or export customer records.
A task description establishes an objective. It does not define all the permissions required to pursue that objective safely.
The problem becomes more serious when the agent reads external material. A support ticket, document, or tool response might contain instructions telling it to upload diagnostic information or use another service.
OWASP identifies this as an indirect prompt-injection risk: instructions embedded in external content can redirect model behavior. That content must not become an authorization source simply because the agent encounters it. OWASP Gen AI Security Project
A secure implementation should still ask:
Who requested the task?
Which customer records can that person access?
Which operations are permitted?
Where may information be sent?
Which actions require separate approval?
Those decisions need enforcement outside the model’s interpretation of the text.
Understand Which Component Owns Each Responsibility
An agent application usually includes several distinct responsibilities:
Component | Responsibility | Architecture question |
|---|---|---|
Model | Interprets context and proposes responses or actions | What happens when its proposal is wrong? |
Harness | Manages execution, tools, context, and sessions | Can another execution path bypass its checks? |
Orchestration | Coordinates tasks and agents | How is authority delegated? |
Runtime | Restricts the execution environment | Which operations are actually contained? |
Business services | Apply domain-specific rules | Is this action valid for this user and resource? |
Inference infrastructure | Serves and routes model requests | How are access, availability, and data handling managed? |
This table is a practical review framework. Products may combine responsibilities, but an architecture review should still identify where each decision is enforced.
For example, a network restriction can prevent an agent from reaching an unapproved destination. It cannot determine whether a refund amount is appropriate under the company’s return policy. That rule belongs in the service performing the refund.
Design Tools Around Specific Business Operations
A common source of risk is giving an agent a broad tool when its task requires a narrow capability.
Consider these two designs:
Broad tool:
ExecuteSql(sqlText)
Task-specific tools:
GetOrderStatus(orderId)
GetReturnEligibility(orderId)
CreateRefundProposal(orderId, reason)The second design creates clearer opportunities to enforce ownership, validate inputs, and apply business rules.
OWASP’s excessive-agency guidance recommends limiting functionality, permissions, and autonomy to what the application requires. It also recommends independent approval for consequential actions. OWASP Gen AI Security Project
For an ASP.NET Core application, an order tool should resolve the tenant from verified identity and confirm that the order belongs to that tenant. A model-supplied tenantId should not establish authority.
The distinction is straightforward: valid JSON means a request can be parsed. It does not mean the request is authorized.
Runtime Controls Must Cover Generated Code
Agents may use more than the tools developers initially expose. A coding agent can launch a shell, create a script, or invoke a different HTTP client.
That makes enforcement coverage important.
OpenShell’s security-policy documentation describes controls for filesystem access, unprivileged processes, and network requests. Network decisions can consider destinations, ports, calling binaries, and configured application-level rules. GitHub
Suppose a repository agent has a read-only API tool. Test whether it can perform an equivalent write through a shell command or generated script.
If the restriction exists only inside the original tool wrapper, the apparent boundary may disappear when the agent chooses another execution path.
A useful security test therefore examines the effect, not just the preferred interface.
Protect Credentials and Limit Their Authority
OpenShell documents a credential-proxy pattern in which the agent workload uses a placeholder, while infrastructure supplies the actual credential for an authorized request. This keeps reusable secrets outside the workload in that configured flow. developer.nvidia.com
For architects, this raises two separate questions:
Can the agent obtain or redirect the secret?
What can the secret authorize at the receiving service?
Solving the first does not automatically solve the second.
A reporting agent should not receive the effective authority of a database administrator merely because its credential is hidden. Service-side permissions must still match the intended task.
Credential revocation also needs testing during execution. A team should know whether revocation affects active sessions, queued requests, delegated work, and cached credentials.
Make Approvals Specific
“Approve this agent” is not a sufficiently precise authorization for a deployment or payment.
An approval should identify the action, target, important parameters, expiration, and approving identity.
For a deployment, bind approval to an immutable artifact identifier and a target environment. If the artifact changes after review, the earlier approval should not silently authorize the replacement.
For a refund, approval should identify the order and amount. Increasing the amount or changing the order should require a new decision.
The execution service should validate the approval immediately before the action. It should also protect against replay so that retries do not repeat an external effect.
These are application responsibilities. A general sandbox cannot infer every organization’s financial or operational rules.
Multi-Agent Systems Need Explicit Delegation
Consider a workflow with three agents:
An investigation agent reads logs.
A coding agent prepares a patch.
A release agent proposes deployment.
Their access should remain distinct.
Record the parent task, child identity, permitted resources, and expiration when delegating work. A child agent should not acquire broader permissions merely by claiming that its parent approved an action.
Shared memory needs similar treatment. An agent-generated note can provide context, but it should not become a permission grant when another agent reads it.
A practical test is to ask whether several individually restricted agents can combine their access to perform an operation that none should be able to authorize.
Audit What Happened, Not Just What the Agent Said
An agent’s transcript may explain its stated intention. It does not prove what a downstream system actually did.
A useful audit record connects:
Evidence | What it establishes |
|---|---|
User and task identity | Who initiated the work |
Agent and parent task | How work was delegated |
Requested operation | What the agent attempted |
Policy version and decision | Why access was allowed or denied |
Approval reference | Who authorized a sensitive operation |
Downstream result | What actually happened |
Recovery record | How failures were handled |
Protect these records from modification by the agent. Apply redaction and retention rules so that logging does not create another collection of exposed credentials or sensitive customer data.
Stopping an Agent Does Not Undo Its Work
Terminating an agent process does not reverse an email, payment, deployment, or external job it already initiated.
Classify operations by reversibility before increasing autonomy.
A draft can usually be deleted. A published message may require a correction. A payment may require a separately authorized reversal. A database rollback cannot undo every external consequence.
Test the entire recovery path: the main process, subprocesses, queues, delegated agents, credentials, and downstream services.
For availability-critical systems, work with the responsible operators to define a safe degraded mode. An abrupt shutdown can itself create an operational problem.
What Teams Should Demonstrate Before Production
A deployment review should ask for evidence that:
Tools reject actions outside the user’s authority.
Generated code cannot bypass the normal tool restrictions.
External documents cannot grant new permissions.
Changed action parameters invalidate earlier approvals.
Delegated agents remain within their assigned access.
Revocation affects ongoing work as designed.
Audit records match downstream outcomes.
Retries and recovery do not duplicate consequential actions.
Start with a limited workflow and expand only after testing its boundaries.
The lasting architectural value is separation: models and workflows can improve without gaining the ability to redefine their own permissions. That allows teams to increase capability while retaining control over the systems those agents can affect.

Join the conversation! Your thoughts help the community grow.