Suppose an AI workflow drafts an answer, checks the answer, and then rewrites it. Each call looks reasonable on its own. Together, they consume more tokens than the workflow was supposed to use. By the time the application adds up the usage reports, the calls have already happened.
Control a workflow's token use by reserving an allowance before each model call, accounting for outstanding reservations, and reconciling each reservation against reported usage. Block the next call when its allowance does not fit. If usage is unknown, preserve that uncertainty and stop admitting further calls until the workflow can be reviewed.
This tutorial implements that accounting boundary in C#. Its protection depends on valid per-call bounds and every call passing through the ledger. It cannot make an inaccurate token estimate into a guaranteed provider limit.
What does a token budget control?
A workflow token budget is an application-defined allowance shared by the model calls needed to complete one task. Here, the allowance covers an agreed total of input and output tokens across calls.
An output setting addresses a smaller scope. Microsoft's documentation defines ChatOptions.MaxOutputTokens as a limit on the generated chat response. It does not represent the total allowance for a multi-call workflow.
Keep three quantities separate:
Quantity |
Meaning in this example |
Reported |
Tokens accounted for from completed usage reports. |
Held |
Tokens reserved for calls whose accounting is not complete. |
Available |
The remaining allowance after reported and held tokens are deducted, with a minimum of zero. |
A token budget also differs from a currency budget. If you need a monetary limit, define the applicable model rates and billing categories separately rather than multiplying every token by an assumed universal price.
What must be known before reserving a call?
The reservation combines an input upper bound and an output cap. The input bound must cover the request actually sent: instructions, conversation history, retrieved passages, tool definitions, and other provider-counted material. A count of the latest user message alone is insufficient when the application sends more context.
Use the model's documented counting method or a verified bound. A rough character-to-token estimate may help planning, but it should not be presented as a hard limit. Apply the output cap to the actual request as well; recording a number in this ledger does not configure the provider.
The example uses one token accounting convention. When integrating a provider, establish how cached input, reasoning, and other usage categories relate to the reported total before combining them. Microsoft's UsageDetails exposes several categories; blindly adding all available fields can mix overlapping quantities.
Build the ledger
Use a .NET 8 console project with C# 12. The example has no external packages, network calls, or API keys. It was compiled and executed with .NET SDK 8.0.414 and runtime 8.0.20 on Linux x64; these identify the test environment, not a recommendation to deploy an older patch.
Create the project:
dotnet new console --name TokenBudgetDemo --framework net8.0
cd TokenBudgetDemo
Replace Program.cs with the following code. Reserve creates a unique accounting entry. Settle records actual usage, or preserves the reservation and halts new admissions if usage is missing.
using System;
using System.Collections.Generic;
var budget = new TokenBudget(3_000);
// Synthetic counts: no model request is made in this example.
Guid draft = budget.Reserve(inputUpperBound: 1_200, outputCap: 400);
budget.Settle(draft, actualTotal: 1_300);
Show("After draft", budget);
Guid review = budget.Reserve(inputUpperBound: 700, outputCap: 500);
Show("Review reserved", budget);
try
{
budget.Reserve(inputUpperBound: 400, outputCap: 200);
}
catch (InvalidOperationException)
{
Console.WriteLine("Next call: BLOCKED");
}
// A missing usage report is unknown, not zero.
budget.Settle(review, actualTotal: null);
Show("Usage missing", budget);
static void Show(string label, TokenBudget budget)
{
BudgetState state = budget.Snapshot();
Console.WriteLine($"{label}: reported={state.Reported}, " +
$"held={state.Held}, available={state.Available}, " +
$"halted={state.Halted}");
}
public readonly record struct BudgetState(
long Reported, long Held, long Available, bool Halted);
public sealed class TokenBudget
{
private readonly object gate = new();
private readonly Dictionary<Guid, long> pending = new();
private readonly long limit;
private long reported;
private long held;
private bool halted;
public TokenBudget(long limit)
{
if (limit <= 0)
throw new ArgumentOutOfRangeException(nameof(limit));
this.limit = limit;
}
public Guid Reserve(long inputUpperBound, long outputCap)
{
if (inputUpperBound < 0)
throw new ArgumentOutOfRangeException(nameof(inputUpperBound));
if (outputCap <= 0)
throw new ArgumentOutOfRangeException(nameof(outputCap));
long requested = checked(inputUpperBound + outputCap);
lock (gate)
{
if (halted || requested > Available())
throw new InvalidOperationException("Token budget blocked.");
Guid id = Guid.NewGuid();
pending.Add(id, requested);
held += requested;
return id;
}
}
public void Settle(Guid id, long? actualTotal)
{
if (actualTotal is < 0)
throw new ArgumentOutOfRangeException(nameof(actualTotal));
lock (gate)
{
if (!pending.TryGetValue(id, out long reserved))
throw new InvalidOperationException("Unknown or settled call.");
if (actualTotal is null)
{
halted = true;
return; // Keep the reservation until usage is reconciled.
}
long updated = checked(reported + actualTotal.Value);
pending.Remove(id);
held -= reserved;
reported = updated;
halted |= actualTotal.Value > reserved;
}
}
public BudgetState Snapshot()
{
lock (gate)
return new(reported, held, Available(), halted);
}
// Called only while holding gate. Clamping avoids overflow after an overrun.
private long Available() =>
Math.Max(0, limit - Math.Min(limit, reported) - held);
}
The shared lock makes checking availability and recording a reservation one atomic operation within this object. Without that boundary, two callers could both observe the same balance and spend it twice. Microsoft's C# lock reference documents this mutual-exclusion behavior.
The lock covers accounting operations only. Perform the asynchronous provider request after reserving and outside the lock. All calls in this workflow must share the same TokenBudget instance.
What should the example print?
Run dotnet run. The executed example produced:
After draft: reported=1300, held=0, available=1700, halted=False
Review reserved: reported=1300, held=1200, available=500, halted=False
Next call: BLOCKED
Usage missing: reported=1300, held=1200, available=500, halted=True
The draft reserves 1,600 tokens and reports 1,300, releasing 300. The review then holds 1,200 of the remaining 1,700. A third call requiring 600 cannot fit into the available 500.
When the review's usage is missing, the ledger retains its 1,200-token hold and halts further admissions. The displayed 500 is an accounting balance, not permission to continue while Halted is true.
An editorial team such as Ranknod could apply this pattern to a hypothetical AI drafting-and-review workflow, deciding in advance which optional revision to skip when the remaining allowance is too small.
Where does the provider call belong?
Integrate the ledger at the boundary that actually dispatches each model request:
Construct the full request and determine its input bound and output cap.
Reserve the combined allowance before sending anything.
Send the request with the corresponding provider output setting.
Normalize the complete usage report and settle that reservation once.
If dispatch may have occurred but usage is unavailable, settle with
nulland stop the workflow's new calls.
Place this boundary inside any retry mechanism that can issue another provider request. One reservation around a wrapper that silently retries does not account for every attempt. Likewise, an agent library that makes several model calls internally needs accounting at those internal call boundaries.
Use an actual total of zero only when the relevant accounting contract confirms zero usage, including a request known never to have been sent. A timeout or cancellation alone does not establish that.
Which failure cases should you check?
The console walkthrough above was compiled and executed. Use the following additional cases to validate the ledger when adapting it to your application; the table gives expected behavior, not additional measured results:
Check |
Expected behavior |
Reservation and reconciliation |
Unused reserved tokens return to availability. |
Duplicate or unknown settlement |
Rejected without changing the balance. |
Missing usage and later reconciliation |
Reservation stays held; a later report updates totals while the halt remains set. |
Usage above the reservation |
Full reported usage is recorded, and new admissions are halted. |
Concurrent reservations |
Of 100 attempts to reserve 100 tokens against a 1,000-token budget, exactly 10 should succeed when none is settled during the check. |
Confirmed unsent request |
Settling with zero releases its reservation. |
Invalid inputs and reservation overflow |
Rejected before the relevant accounting state changes. |
These checks concern local accounting with synthetic inputs. Tokenizer accuracy, provider usage mapping, and billing reconciliation need separate validation with the integration you actually use.
What are the implementation limits?
The halt is deliberately persistent in this sample. Reconciling a missing report does not automatically restart an interrupted workflow. A production recovery path should review the outstanding calls and persist the decision to resume.
If actual usage exceeds its reservation, the ledger can record the overrun and block future work; it cannot undo already consumed tokens. Requests already admitted in parallel may also finish. Stronger admission guarantees require verified upper bounds and enforcement at every model-call boundary.
Finally, this is an in-memory ledger for one process. A restart loses its state, and a second process has a separate lock. A distributed application needs durable reservation records, atomic balance updates, and deduplication of usage events.
Define the stop behavior before connecting the first model: preserve completed work, report which stage stopped, and require an explicit decision before increasing the allowance. That gives the workflow a clear boundary the application can explain and test.
Join the conversation! Your thoughts help the community grow.