AI agents are becoming capable of performing tasks that previously required direct developer involvement. An agent can generate code, execute scripts, process files, query services, and combine multiple tools to complete a request.

That capability creates a security problem that traditional application architectures do not always have to address:

What happens when an AI agent needs to execute code that cannot be fully trusted?

Running that code directly inside the application's process or on the application server can expose sensitive resources. A malicious prompt, unexpected model output, compromised dependency, or incorrectly generated command could potentially affect files, credentials, network resources, or other application components.

The safer approach is to treat agent-generated code as untrusted input and execute it inside an isolated sandbox.

This article explains how AI agent sandboxing works, why it matters, how to design the execution boundary, and what developers should consider before using sandboxed code execution in production.

Why AI-Generated Code Should Be Treated as Untrusted

AI models generate code based on context and instructions. They do not provide a security guarantee for the generated program.

Consider an agent that receives this request:

Analyze this data file and write a Python script
to calculate the required statistics.

The generated program may be completely reasonable. However, the application should not assume that every generated instruction is safe.

For example, the code could potentially attempt to:

open("/some/system/file")

or make an unexpected network request:

import requests

requests.get("https://example.com")

The problem is not limited to malicious intent. Generated code can contain accidental behavior as well.

A coding agent might:

  • Read files outside the intended workspace

  • Consume excessive CPU or memory

  • Create large files

  • Run indefinitely

  • Access network resources

  • Install packages

  • Execute operating-system commands

  • Interact with services that were never part of the original task

Therefore, the correct security model is:

AI-generated code
       |
       v
Treat as untrusted
       |
       v
Validate request
       |
       v
Execute in sandbox
       |
       v
Return controlled result

What Is an AI Agent Sandbox?

An AI agent sandbox is an isolated execution environment where potentially unsafe operations can be performed without giving the workload unrestricted access to the host application or server.

A simplified architecture looks like this:

+----------------------+
|       User           |
+----------+-----------+
           |
           v
+----------------------+
|     AI Agent         |
+----------+-----------+
           |
           v
+----------------------+
| Execution Policy     |
| - Authorization      |
| - Validation         |
| - Limits             |
+----------+-----------+
           |
           v
+----------------------+
|     Sandbox          |
|                      |
| Generated Code       |
| Temporary Files      |
| Limited Runtime      |
+----------+-----------+
           |
           v
+----------------------+
| Controlled Result    |
+----------------------+

The sandbox becomes a security boundary between the trusted application and the untrusted workload.

What Should Be Isolated?

A useful sandbox should restrict several resources.

Resource

Example Control

File system

Temporary workspace only

Network

Disabled or allowlisted

CPU

Resource limit

Memory

Resource limit

Processes

Restrict child processes

Runtime

Execution timeout

Packages

Controlled dependencies

Credentials

No inherited secrets

Storage

Size and lifetime limits

The goal is not simply to isolate the generated code from the application process.

The goal is to limit what the generated code can do even if it behaves unexpectedly.

Keep the Application and Sandbox Separate

One of the most important architectural decisions is to avoid executing untrusted code directly inside the main application.

For example, this approach is risky:

public async Task<string> ExecuteAgentCodeAsync(string code)
{
    // Execute generated code directly in application process
    return await RunCodeAsync(code);
}

The application and generated workload share the same runtime boundary.

A safer design introduces a separate execution service:

public interface ICodeExecutionService
{
    Task<ExecutionResult> ExecuteAsync(
        ExecutionRequest request,
        CancellationToken cancellationToken);
}

The application sends an execution request to an isolated environment.

This gives the application a clear boundary:

Application
     |
     | ExecutionRequest
     v
Execution Service
     |
     v
Sandbox

The exact sandbox technology can vary depending on the workload and security requirements.

Validate the Execution Request

The sandbox should not be the only security control.

The application should validate requests before execution.

For example:

public sealed record ExecutionRequest(
    string Language,
    string Code,
    long InputSize);

A simple validation layer could look like:

public bool IsAllowed(ExecutionRequest request)
{
    if (string.IsNullOrWhiteSpace(request.Code))
        return false;

    if (request.Code.Length > 100_000)
        return false;

    return request.Language switch
    {
        "python" => true,
        "javascript" => true,
        _ => false
    };
}

This example is intentionally simple. String-based checks should not be treated as a complete security mechanism.

The important design principle is that the application should establish policy before the workload reaches the execution environment.

Do Not Rely on Code Filtering Alone

A common mistake is attempting to make generated code safe by searching for dangerous strings.

For example:

if (code.Contains("Process.Start"))
{
    return false;
}

This is not a reliable security boundary.

Programming languages provide many ways to access functionality indirectly, and code can be transformed or obfuscated.

Code filtering may be useful as an additional validation layer, but it should not replace process or environment isolation.

The stronger boundary is:

Validation
    +
Isolation
    +
Resource limits
    +
Identity controls
    +
Network controls

Restrict File System Access

A sandbox should have access only to the files required by the task.

Suppose the user uploads:

sales.csv

The execution environment might receive:

/workspace/input/sales.csv

The generated program should not have unrestricted access to the host file system.

The workflow can be:

User File
   |
   v
Temporary Workspace
   |
   v
Sandbox Execution
   |
   v
Output Artifact
   |
   v
Workspace Cleanup

After execution, temporary files should be cleaned up according to the application's retention requirements.

This prevents one task from accidentally exposing data belonging to another task.

Restrict Network Access

Network access deserves special attention.

If generated code can freely access the internet or internal services, the sandbox becomes much more powerful than necessary.

For example, a data-processing task may not need network access at all.

A safer default is:

No network access

If network access is required, define what is allowed.

For example:

Sandbox
   |
   +--> Approved API
   |
   +--> Approved Storage

rather than:

Sandbox
   |
   +--> Any Internet Address
   +--> Internal Network
   +--> Cloud Services

The exact implementation depends on the hosting environment, but the security principle remains the same: make network access explicit rather than automatic.

Never Inherit Application Credentials

The sandbox should not automatically receive the application's credentials.

Avoid designs where an execution environment can access:

Database credentials
Cloud credentials
API keys
Signing keys
Deployment credentials
Service-account tokens

Instead, use narrowly scoped identities where access is genuinely required.

For example:

Application Identity
        |
        +--> Task-specific authorization
                     |
                     v
                 Sandbox

The sandbox should receive the minimum permissions needed for the specific workload.

In many cases, the sandbox does not need credentials at all.

Add Execution Timeouts

Untrusted code can run indefinitely.

For example:

while True:
    pass

Without a timeout or resource control, this could consume resources continuously.

In a .NET application, cancellation can provide an application-level execution boundary:

using var timeout =
    CancellationTokenSource.CreateLinkedTokenSource(
        cancellationToken);

timeout.CancelAfter(TimeSpan.FromSeconds(30));

var result = await executionService.ExecuteAsync(
    request,
    timeout.Token);

The sandbox itself should also enforce an execution limit.

Application cancellation should not be the only protection.

Control CPU and Memory

Timeouts alone are not sufficient.

Consider code that completes within the timeout but allocates excessive memory.

For example:

data = []

while True:
    data.append("large-data-block")

The process could consume substantial memory before the timeout occurs.

A production sandbox should therefore have limits for:

  • CPU

  • Memory

  • Process count

  • File size

  • Disk usage

  • Execution duration

  • Concurrent executions

These controls belong to the execution environment rather than the AI model.

Isolate Parallel Tasks

Multiple agent tasks may run simultaneously.

For example:

                Agent
                  |
        +---------+---------+
        |         |         |
        v         v         v
    Sandbox A Sandbox B Sandbox C

Each task should have its own workspace and execution boundary.

Do not allow one task to reuse another task's temporary files.

A useful identifier can be generated for each execution:

var executionId = Guid.NewGuid().ToString("N");

The identifier can be used for:

  • Temporary workspace

  • Logging

  • Execution status

  • Output artifacts

  • Cleanup

This makes concurrent execution easier to manage.

Handle Errors Without Exposing the Host

The sandbox should return controlled execution results.

For example:

public sealed record ExecutionResult(
    bool Success,
    string? Output,
    string? Error,
    TimeSpan Duration);

Avoid returning raw infrastructure details to end users.

For example, an internal error might contain information about:

Host paths
Container configuration
Internal service names
Credentials
Network topology

The user-facing application should expose an appropriate error while keeping sensitive infrastructure details in protected logs.

Logging and Auditing

Sandbox execution should be observable.

At minimum, record:

Execution ID
User or application identity
Task type
Start time
End time
Execution status
Resource usage where available
Failure reason

For example:

logger.LogInformation(
    "Sandbox execution {ExecutionId} completed with status {Status}",
    executionId,
    result.Success ? "Success" : "Failed");

Do not automatically log the complete generated code or sensitive input data.

Logging should provide enough information to investigate failures without creating another data-exposure problem.

Common Mistakes

Running Generated Code Inside the Web Server

This removes a critical isolation boundary.

Use a separate execution environment.

Assuming the AI Model Will Behave Safely

Model instructions are not a substitute for infrastructure security.

Enforce security outside the model.

Using Only String-Based Code Blocking

Searching for dangerous keywords is not a complete security mechanism.

Use actual execution isolation.

Giving the Sandbox Full Network Access

Most workloads do not need unrestricted network access.

Start with no access and explicitly enable required destinations.

Sharing Temporary Directories

Parallel tasks should have isolated workspaces.

Passing Application Credentials to the Sandbox

Use scoped access only when required.

Omitting Resource Limits

A sandbox without CPU, memory, storage, and execution limits can still become a resource-exhaustion problem.

Advantages and Disadvantages

Advantages

AI agent sandboxing provides:

  • Stronger isolation for untrusted workloads

  • Reduced risk to the main application

  • Controlled file-system access

  • Resource boundaries

  • Independent task execution

  • Better support for parallel workloads

  • Clearer security boundaries

  • Easier cleanup of temporary environments

Disadvantages

Sandboxing also introduces additional engineering work:

  • More infrastructure components

  • Additional execution latency

  • Resource-management complexity

  • More complicated monitoring

  • More difficult debugging

  • Additional networking and identity configuration

  • Potentially higher infrastructure costs

The right architecture depends on what the agent needs to execute and how sensitive the surrounding application is.

A Practical Secure Architecture

A production-oriented design can follow this pattern:

+-------------------+
|       User        |
+---------+---------+
          |
          v
+-------------------+
|    AI Agent       |
+---------+---------+
          |
          v
+-------------------+
| Policy Gateway    |
|                   |
| Authorization     |
| Input Validation  |
| Rate Limits       |
+---------+---------+
          |
          v
+-------------------+
| Execution Queue   |
+---------+---------+
          |
     +----+----+
     |         |
     v         v
+---------+ +---------+
|Sandbox A| |Sandbox B|
+----+----+ +----+----+
     |           |
     +-----+-----+
           |
           v
    Result Collector
           |
           v
       AI Agent

This architecture separates responsibilities.

The AI agent decides what it wants to accomplish.

The policy layer decides whether the requested operation is allowed.

The sandbox performs the isolated work.

The result collector returns a controlled result.

That separation is important because the AI model should not be the final authority over security-sensitive operations.

Best Practices Checklist

Before running AI-generated code in production, verify:

[ ] Generated code is treated as untrusted
[ ] Execution happens outside the main application process
[ ] File-system access is restricted
[ ] Network access is disabled or explicitly allowlisted
[ ] Application credentials are not inherited
[ ] CPU and memory limits are configured
[ ] Execution timeouts are enforced
[ ] Input and output sizes are limited
[ ] Each task has an isolated workspace
[ ] Parallel tasks cannot access each other's data
[ ] Execution events are logged appropriately
[ ] Sensitive data is excluded from logs
[ ] Failed tasks are cleaned up
[ ] Security policy is enforced outside the AI model

Summary

AI agent sandboxing is an important architectural pattern whenever an agent needs to execute generated code or process potentially untrusted workloads.

A secure implementation separates the agent from the execution environment, validates requests, restricts file and network access, protects credentials, enforces resource limits, isolates parallel tasks, applies execution timeouts, and maintains appropriate observability.

Most importantly, security should be enforced by the application and infrastructure rather than relying on the AI model to make the correct security decision every time.