Introduction
A user closes an AI request after waiting too long. A background log records “model timeout.” Later, the team raises the timeout, even though some of those requests ended because the caller had already left. Two different events have been given the same name, and an infrastructure change is being used to solve a diagnosis problem.
Keep caller cancellation and the application's time budget as separate signals, link them for the AI operation, and classify the result from the original signals. This lets a .NET application request cancellation through one token while preserving a useful local explanation of why it stopped waiting.
The example below uses a cooperative fake model call. It does not contact an AI provider, cancel remote billing, or prove that a provider stopped generation. Those are separate behaviors that must be verified for the actual integration.
What does cancellation mean in .NET?
.NET uses cooperative cancellation: a caller signals a token and participating operations observe it. Passing a token to an operation matters only when that operation responds appropriately.
This distinction becomes important with AI clients. A token may stop a local HTTP wait while the remote system continues working. A wrapper cannot establish what the provider did after the connection was interrupted. Record the remote outcome as unknown unless the provider exposes a documented way to determine it.
An application also needs to distinguish a request that its user canceled from one that exceeded its own waiting budget. Neither event, by itself, proves that the model produced an incorrect answer.
Set up the demonstration
Use the .NET 10 SDK and a console project. No external package or API credential is required:
dotnet new console --framework net10.0 --name AiCancellationDemo
cd AiCancellationDemo
Replace Program.cs with the following code. Code status: illustrative and untested. It has not been compiled in this drafting environment; the expected results below must be verified on the stated SDK.
using System;
using System.Threading;
using System.Threading.Tasks;
Console.WriteLine((await RunAsync(
token => FakeModelAsync(10, token),
TimeSpan.FromSeconds(5),
CancellationToken.None)).Status);
Console.WriteLine((await RunAsync(
token => FakeModelAsync(10_000, token),
TimeSpan.FromMilliseconds(25),
CancellationToken.None)).Status);
using var caller = new CancellationTokenSource();
caller.Cancel();
Console.WriteLine((await RunAsync(
token => FakeModelAsync(10, token),
TimeSpan.FromSeconds(5),
caller.Token)).Status);
static async Task<string> FakeModelAsync(
int delayMs,
CancellationToken token)
{
await Task.Delay(delayMs, token);
return "Synthetic AI answer";
}
static async Task<CallOutcome> RunAsync(
Func<CancellationToken, Task<string>> invoke,
TimeSpan budget,
CancellationToken callerToken)
{
ArgumentNullException.ThrowIfNull(invoke);
if (budget <= TimeSpan.Zero)
throw new ArgumentOutOfRangeException(nameof(budget));
using var deadline = new CancellationTokenSource();
deadline.CancelAfter(budget);
using var linked = CancellationTokenSource.CreateLinkedTokenSource(
callerToken,
deadline.Token);
try
{
linked.Token.ThrowIfCancellationRequested();
string answer = await invoke(linked.Token);
linked.Token.ThrowIfCancellationRequested();
return new CallOutcome(
CallStatus.Completed,
answer);
}
catch (OperationCanceledException)
when (callerToken.IsCancellationRequested)
{
return new CallOutcome(
CallStatus.CallerCancelled,
null);
}
catch (OperationCanceledException)
when (deadline.IsCancellationRequested)
{
return new CallOutcome(
CallStatus.DeadlineExceeded,
null);
}
}
enum CallStatus
{
Completed,
CallerCancelled,
DeadlineExceeded
}
record CallOutcome(CallStatus Status, string? Answer);
The first case gives a short synthetic operation a generous budget. The second deliberately waits longer than its budget. The third uses an already-canceled caller token, making it clear that the wrapper should not start useful work.
CallOutcome separates a successful answer from cancellation status. It avoids passing a partial or missing answer onward as a completed result. An application can map that outcome to its own user interface or job state without rewriting the cancellation logic.
Why use two token sources?
The deadline source owns the application's local time budget. The caller token belongs to the request or job owner. CreateLinkedTokenSource produces a token that is signaled when either source is signaled.
The wrapper passes that linked token to the operation. It also checks before and after awaiting the result. The first check avoids starting an operation after cancellation is already known. The second prevents a late result from being labeled successful when cancellation was signaled during the wait but the operation returned normally.
The two catch filters inspect the original signals. Caller cancellation is checked first, so it wins if both are already signaled when the exception is handled. That is an explicit classification policy, not a precise reconstruction of which event happened first.
For a technical explanation prepared by Ranknod, that local-versus-remote distinction should stay visible: a canceled wait is not evidence of a canceled inference.
An OperationCanceledException with neither original source signaled is allowed to propagate. The wrapper does not relabel an unexplained provider or adapter cancellation as an application deadline. Other faults also propagate for the caller to handle.
What should the program print?
Run dotnet run. Under normal scheduling, the expected output is:
Completed
DeadlineExceeded
CallerCancelled
These are expected demonstration outcomes, not observed results from this draft. The timing case depends on the scheduler, so a serious test suite should avoid using a tiny wall-clock race as its main proof.
Add a controllable fake operation to test four boundaries: success before either signal, an already-canceled caller, a deadline while the operation is pending, and an unrelated exception. For the simultaneous-signal case, assert the documented caller-first policy rather than guessing which cancellation “really” won.
Also test an operation that ignores the token. The wrapper will keep awaiting it. If it eventually returns after cancellation, the post-await check classifies the canceled result. If it never returns, this implementation never regains control. That limitation is essential to the design.
Is this a hard timeout?
No. A cancellation request is cooperative. This wrapper does not forcibly terminate an arbitrary operation at an exact elapsed time. If the caller must stop waiting independently of cooperation, design that waiting policy separately, including how the still-running task is observed and cleaned up.
Avoid disposing resources that an abandoned operation still needs without an explicit ownership strategy. Avoid treating “we stopped awaiting” as “nothing else can happen.” Both assumptions can create difficult background failures.
For a real AI SDK, inspect which methods accept a cancellation token and test the exact path you use, including streaming if relevant. Where an HTTP client's own timeout is involved, document how its exception is distinguished from the application budget. Do not rely on a generic catch block to infer that distinction from an error message.
What belongs in the operational record?
Record a request identifier, the local outcome, elapsed time, configured budget, and any permitted provider request identifier. Keep raw prompts, credentials, and private retrieved passages out of general-purpose logs.
A request that times out locally may require provider-specific reconciliation before a retry. A user who deliberately canceled may not want a retry at all. Those are product decisions that should follow a clear outcome rather than be hidden inside this wrapper.
Useful cancellation handling gives the next component enough information to act honestly. The user can leave, the application can enforce its waiting policy, and the operator can investigate the right cause without turning every interrupted AI request into the same incident.
Summary
Caller cancellation and an application's own time budget represent different signals and should be tracked separately. By linking the cancellation tokens for the operation while retaining the original sources for classification, a .NET application can distinguish completed requests, caller cancellations, and local deadline expirations without incorrectly assuming what happened on the remote AI provider.
Join the conversation! Your thoughts help the community grow.