Introduction
When adding large language models (LLMs) to a .NET application, developers usually focus on correctness, latency, reliability, and security. One concern that can easily be overlooked is cost as a failure mode.
Token-based APIs can generate unexpected costs when an application retries requests repeatedly, processes large documents, or allows an AI agent to make more model calls than expected. Monitoring and billing reports can show what happened after the fact, but they do not prevent an application from exceeding a spending limit.
A more reliable approach is to enforce spending limits before an API request is executed.
This article explains how pre-flight budget enforcement can be used with .NET applications and OpenAI-compatible LLM gateways to place explicit cost boundaries around applications, services, and AI agents.
The Problem With Post-Request Cost Monitoring
Consider a background service that sends documents to an LLM for processing.
A typical workflow might look like this:
Application
↓
LLM API
↓
Token processing
↓
Response
↓
Usage recorded
If the application accidentally retries a request several times, the requests may already have consumed tokens before the application detects the problem.
The same issue can occur with AI agents.
An agent might:
Make more tool calls than expected.
Repeat an operation after a failed response.
Process unexpectedly large inputs.
Enter an inefficient retry loop.
Generate substantially more output than anticipated.
A monitoring system can identify these events, but the cost has already been incurred.
The key question is therefore:
Can the application determine whether a request is allowed to spend money before sending it to the model?
This is where pre-flight budget enforcement becomes useful.
What Is Pre-Flight Budget Enforcement?
Pre-flight budget enforcement places a spending check in front of the model request.
The simplified flow becomes:
Application
↓
Budget Check
↓
Request Allowed?
↙ ↘
No Yes
↓ ↓
Reject LLM API
↓
Response
Before the request reaches the model, the gateway or API control layer checks whether the request is within the configured budget.
Budgets can be defined at different levels, depending on the infrastructure:
Organization
Team
Application
Service
API key
AI agent
If the request would exceed the applicable budget, the request can be rejected immediately.
For example:
Monthly Budget: $100
Current Usage: $96
Estimated Request: $7
Available Budget: $4
Result:
Request rejected
The important distinction is that the request is rejected before model execution.
Why API-Level Budget Enforcement Matters
A budget implemented only inside application code can be useful, but it may not protect every path through which an application accesses an LLM.
For example:
Web Application ─────┐
│
Background Worker ───┼──→ LLM Gateway ─→ Model Providers
│
AI Agent ────────────┘
A centralized gateway can enforce the same financial boundary for all of these clients.
This is particularly useful when multiple applications or services share access to the same collection of models.
The gateway becomes the enforcement point:
Application
↓
Gateway
↓
Budget Validation
↓
Model Provider
The application does not need to implement an independent spending-control mechanism for every model provider.
Handling Budget Errors in .NET
From a .NET application's perspective, a budget rejection is simply an HTTP failure response.
For example, a gateway can use an HTTP status such as 402 Payment Required to indicate that a configured spending limit has been reached.
The application can handle that condition explicitly:
try
{
ChatCompletion completion = client.CompleteChat(
"Summarize this quarter's support tickets..."
);
Console.WriteLine(completion.Content[0].Text);
}
catch (ClientResultException ex) when (ex.Status == 402)
{
// The configured spending limit was reached.
Console.WriteLine(
"The request was rejected because the configured LLM budget was exceeded."
);
// Queue the work, notify an administrator, or degrade gracefully.
}
catch (ClientResultException ex)
{
Console.WriteLine($"Request failed: {ex.Status}");
}
The important part is to treat a budget rejection differently from a transient network failure.
A 429 response may indicate that the application should wait and retry because of rate limiting.
A 402 budget response can instead mean:
Stop spending.
Automatically retrying the request could defeat the purpose of the budget.
Using an OpenAI-Compatible Endpoint
One advantage of an OpenAI-compatible gateway is that existing .NET applications can often continue using the same SDK patterns.
The endpoint changes, while the application-level request structure remains similar.
For example:
using OpenAI;
using OpenAI.Chat;
var client = new ChatClient(
model: "your-model",
credential: new ApiKeyCredential(
Environment.GetEnvironmentVariable("LLM_API_KEY")!
),
options: new OpenAIClientOptions
{
Endpoint = new Uri("https://your-gateway.example.com/v1")
});
ChatCompletion completion = client.CompleteChat(
"Summarize this quarter's support tickets..."
);
Console.WriteLine(completion.Content[0].Text);
The exact configuration depends on the gateway being used, but the architectural idea remains the same:
.NET Application
↓
OpenAI-Compatible Gateway
↓
Budget Enforcement
↓
Model Provider
This can also make model changes easier because the application does not necessarily need separate provider-specific credentials for every model.
Per-Application and Per-Service Budgets
A single organization-wide budget is often not enough.
Suppose an organization has three applications:
Customer Support
Document Processing
AI Agent
If they all share one budget, a runaway process in one application could consume resources intended for the others.
A more controlled structure is:
Organization Budget
│
├── Customer Support Budget
│
├── Document Processing Budget
│
└── AI Agent Budget
The same principle can be applied at the API-key level.
For example:
summarizer-key
Budget: $50/month
document-worker-key
Budget: $100/month
agent-key
Budget: $25/month
If the agent unexpectedly generates a large number of requests, its budget can be exhausted without necessarily affecting unrelated applications.
This provides a form of financial blast-radius control.
Mapping Budgets to ASP.NET Core Clients
In an ASP.NET Core application, separate API keys can be associated with different named HTTP clients.
For example:
builder.Services.AddHttpClient("summarizer", client =>
{
client.BaseAddress = new Uri(
"https://your-gateway.example.com/v1"
);
client.DefaultRequestHeaders.Authorization =
new AuthenticationHeaderValue(
"Bearer",
builder.Configuration["LLM:SummarizerKey"]
);
});
Another service can use a different key:
builder.Services.AddHttpClient("documentProcessor", client =>
{
client.BaseAddress = new Uri(
"https://your-gateway.example.com/v1"
);
client.DefaultRequestHeaders.Authorization =
new AuthenticationHeaderValue(
"Bearer",
builder.Configuration["LLM:DocumentProcessorKey"]
);
});
The infrastructure can then associate each key with its own budget.
This creates a useful relationship:
.NET Service
↓
Named Client
↓
API Key
↓
Budget
The application architecture and the financial boundary can therefore be aligned.
Budget Enforcement for AI Agents
Budget controls become particularly important when working with AI agents.
A conventional request might have a relatively predictable cost:
Request → Model → Response
An agent can behave differently:
User Request
↓
Agent
↓
Tool Call
↓
Model
↓
Tool Call
↓
Model
↓
Tool Call
↓
Model
↓
Final Response
The number of model calls may depend on the task.
Instead of attempting to predict the exact cost of every possible agent execution, an application can establish a maximum spending boundary.
For example:
Agent Budget: $2
Request 1 → $0.40
Request 2 → $0.35
Request 3 → $0.50
Request 4 → $0.45
Total → $1.70
Remaining → $0.30
Once the configured limit is reached, subsequent requests can be rejected.
This does not guarantee that an agent will behave efficiently. It ensures that inefficient behavior has a defined financial boundary.
Logging Cost and Usage Information
Budget enforcement answers:
Can this request execute?
Usage tracking answers:
Where is the money being spent?
For production .NET applications, cost information can be recorded alongside existing application telemetry.
For example:
logger.LogInformation(
"LLM request completed. CorrelationId={CorrelationId}, Cost={Cost}",
correlationId,
requestCost
);
When combined with:
Correlation IDs
Request IDs
Service names
API keys
User or tenant identifiers
Model names
cost information becomes easier to trace back to the originating workload.
A useful production record might therefore look like:
CorrelationId: 8f42...
Service: DocumentProcessor
Model: Model-A
Input Tokens: 4200
Output Tokens: 850
Cost: $0.08
This makes cost part of normal application observability rather than something reviewed only during billing reconciliation.
Budget Enforcement and Retry Policies
Retry policies need special consideration when budget enforcement is involved.
A transient error may be appropriate for retry:
Timeout
↓
Wait
↓
Retry
A rate-limit response may also be retriable:
429
↓
Backoff
↓
Retry
A budget rejection is different:
402
↓
Budget exhausted
↓
Do not repeatedly retry
Automatically retrying a request that has already been rejected because of a budget limit generally does not solve the underlying problem.
Instead, the application can:
Queue the work.
Notify an administrator.
Degrade to a lower-cost workflow.
Wait until the budget becomes available.
Ask for an explicit budget increase.
The exact behavior depends on the application's requirements.
What About Development and Testing?
Budget controls are useful outside production as well.
Development and test environments can accidentally generate significant usage through:
Automated integration tests.
Load tests.
Large test datasets.
Agent experiments.
Repeated local executions.
CI/CD pipelines.
Separate development keys with small budgets can provide an additional safety boundary.
For example:
Production Key
Budget: $500
Development Key
Budget: $10
CI Key
Budget: $5
If a test unexpectedly makes hundreds of requests, the budget boundary can stop the workload rather than allowing it to continue indefinitely.
Choosing the Right Budget Strategy
Not every application needs the same level of enforcement.
For a small application with predictable usage, a single organization-level budget may be sufficient.
For larger systems, consider separating budgets by:
Environment
Team
Application
Service
Customer
API key
Agent
A useful hierarchy might look like:
Organization
│
├── Team A
│ ├── Service A1
│ └── Service A2
│
└── Team B
├── Service B1
└── AI Agent
The goal is not to create as many budgets as possible. The goal is to establish boundaries that correspond to meaningful ownership and financial responsibility.
Conclusion
LLM cost should be treated as an operational failure mode rather than something discovered only after usage has already occurred.
Post-request monitoring remains important for understanding usage, but it does not prevent runaway spending. Pre-flight budget enforcement adds a control point before the model request executes, allowing applications to reject requests that would exceed defined spending limits.
For .NET applications, the implementation can remain straightforward. An OpenAI-compatible gateway can sit between the application and model providers, while HTTP error handling, named clients, logging, and existing observability infrastructure can handle budget-related failures.
For applications using AI agents or other unpredictable workloads, per-service or per-key budgets can provide an additional financial boundary. The result is a system where cost is not merely measured after the fact but becomes part of the request path itself.

Join the conversation! Your thoughts help the community grow.