Why Does an AI Feature Slow Down When Several People Use It?

An AI feature may work comfortably during a demonstration with a small number of users. The situation changes when many requests arrive at the same time.

A burst of requests can start hundreds of asynchronous operations. Each operation may spend most of its time waiting for a remote AI service, but the application still needs to retain request state, manage responses, handle cancellations, and eventually return results to users.

Users may see a spinner without knowing whether their request has started, is waiting for capacity, or has already been canceled.

A practical way to control this workload is to use a bounded admission queue with a fixed number of workers.

The queue limits how much work can wait. The workers limit how many operations can execute at the same time.

You also need an explicit policy for what happens when the queue is full and when a caller cancels.

This approach controls local application concurrency. It does not automatically enforce an AI provider's request-per-minute, token, or billing limits.

The following console example makes these boundaries visible before connecting a real AI model.

Which Limit Solves Which Problem?

Different controls solve different problems.

Control

Question It Answers

What It Does Not Establish

Queue capacity

How many accepted jobs may wait here?

How quickly the provider accepts requests

Worker count

How many jobs may execute at once here?

A requests-per-minute limit

Admission policy

What happens when no waiting slot is available?

Whether another application instance has capacity

Rate limiter

How much traffic may start within a time window?

How much queued work the application can retain

A semaphore can restrict active work while many callers continue waiting to acquire it. That may be acceptable when admission is bounded elsewhere.

However, if every incoming request creates another waiting task without a limit, a semaphore alone does not bound the backlog.

For this example, a full queue rejects a new job immediately. An application could instead allow callers to wait for capacity, but that waiting period should also have a deadline and an explicit limit.

What Do You Need to Run the Example?

Use the .NET 10 SDK and create a console project:

dotnet new console --framework net10.0 --name AiQueueDemo
cd AiQueueDemo

Replace Program.cs with the following code.

The example uses a short, cancellable delay as a fake AI operation. This makes it possible to inspect the concurrency mechanism independently of an external AI provider.

The code is illustrative. The important properties are the queue capacity, worker limit, cancellation behavior, and admission behavior.

using System;
using System.Linq;
using System.Threading;
using System.Threading.Channels;
using System.Threading.Tasks;

const int capacity = 3;
const int workerCount = 2;

var queue = Channel.CreateBounded<AiJob>(
    new BoundedChannelOptions(capacity)
    {
        FullMode = BoundedChannelFullMode.Wait,
        SingleWriter = true,
        SingleReader = false
    });

using var shutdown = new CancellationTokenSource();
using var canceledCaller = new CancellationTokenSource();

canceledCaller.Cancel();

int active = 0;
int peak = 0;
int completed = 0;
int canceled = 0;
int rejected = 0;

// Prefill before starting workers so admission results are deterministic.
var jobs = new[]
{
    new AiJob("q1", CancellationToken.None),
    new AiJob("q2", CancellationToken.None),
    new AiJob("q3", canceledCaller.Token),
    new AiJob("q4", CancellationToken.None)
};

foreach (var job in jobs)
{
    if (!queue.Writer.TryWrite(job))
    {
        rejected++;
    }
}

queue.Writer.Complete();

await Task.WhenAll(
    Enumerable.Range(0, workerCount)
        .Select(_ => ConsumeAsync()));

Console.WriteLine(
    $"completed={completed}; canceled={canceled}; rejected={rejected}");

Console.WriteLine(
    $"peak active={peak}; limit respected={peak <= workerCount}");

async Task ConsumeAsync()
{
    await foreach (var job in queue.Reader.ReadAllAsync(shutdown.Token))
    {
        if (job.Caller.IsCancellationRequested)
        {
            Interlocked.Increment(ref canceled);
            continue;
        }

        using var linked = CancellationTokenSource.CreateLinkedTokenSource(
            job.Caller,
            shutdown.Token);

        int now = Interlocked.Increment(ref active);
        UpdatePeak(now);

        try
        {
            await Task.Delay(100, linked.Token);

            linked.Token.ThrowIfCancellationRequested();

            Interlocked.Increment(ref completed);
        }
        catch (OperationCanceledException)
            when (linked.IsCancellationRequested)
        {
            Interlocked.Increment(ref canceled);
        }
        finally
        {
            Interlocked.Decrement(ref active);
        }
    }
}

void UpdatePeak(int candidate)
{
    int observed;

    do
    {
        observed = Volatile.Read(ref peak);

        if (candidate <= observed)
        {
            return;
        }
    }
    while (Interlocked.CompareExchange(
        ref peak,
        candidate,
        observed) != observed);
}

record AiJob(string Id, CancellationToken Caller);

The example uses .NET's System.Threading.Channels API to create a bounded producer-consumer queue.

The channel stores waiting jobs, while two consumers perform the work.

Why Use TryWrite with Wait Mode?

The channel's full mode and the writer method serve different purposes.

BoundedChannelFullMode.Wait defines the behavior of the bounded channel when it reaches capacity.

TryWrite attempts to add an item immediately and returns whether the operation succeeded. It does not wait for capacity to become available.

This example deliberately fills the queue before starting the workers.

With a capacity of three:

Job 1 → Accepted
Job 2 → Accepted
Job 3 → Accepted
Job 4 → Rejected

The fourth job cannot enter because all three waiting slots are occupied.

Calling:

queue.Writer.Complete();

announces that no additional jobs will be written. It does not discard jobs that were already accepted.

Two consumers then read the queue.

Each consumer waits for its current operation to finish before reading another job. Therefore, the design allows at most two active operations at a time.

This relationship is important.

Starting an unawaited operation inside the worker loop could allow the worker to continue reading additional jobs and would break the intended concurrency limit.

How Does the Concurrency Limit Work?

The important relationship is:

                 Bounded Queue
                Capacity = 3
                      |
                      ↓
          ┌─────────────────────┐
          │       Waiting       │
          │  q1  q2  q3        │
          └─────────────────────┘
                 ↓       ↓
              Worker   Worker
                 1        2
                 ↓        ↓
              AI Call  AI Call

The queue controls waiting work.

The workers control active work.

With:

const int capacity = 3;
const int workerCount = 2;

the application can have up to three accepted jobs waiting in the queue and two jobs executing concurrently.

These are separate limits.

A larger queue does not increase execution capacity.

More workers increase execution concurrency, but they may also increase pressure on the downstream AI service.

What Should the Output Show?

Run the application:

dotnet run

The expected final counters are:

completed=2; canceled=1; rejected=1

The second line reports the measured peak and whether it remained at or below two:

peak active=2; limit respected=True

The exact scheduling order is not important. The important property is that the peak number of active operations never exceeds the configured worker count.

The third job has already been canceled before it reaches the worker.

It can still occupy a queue slot because cancellation does not automatically remove an item from a Channel<T>.

The worker discovers the cancellation when it dequeues the job and skips execution.

This behavior matters in larger systems because canceled jobs can temporarily occupy queue capacity.

What Happens When an Active Request Is Canceled?

The example also demonstrates cancellation during an active operation.

The worker creates a linked cancellation token:

using var linked = CancellationTokenSource.CreateLinkedTokenSource(
    job.Caller,
    shutdown.Token);

The token combines:

  • The caller's cancellation request.

  • The application's shutdown request.

The fake AI operation then uses the linked token:

await Task.Delay(100, linked.Token);

If cancellation occurs while the operation is running, the delay throws OperationCanceledException.

The worker handles that cancellation and decrements the active-operation counter in the finally block.

That cleanup is important.

Without reliable cleanup, an application could incorrectly believe that workers are still occupied even after an operation has ended.

What Checks Matter Before Connecting a Real AI Model?

Before connecting an actual AI provider, verify the behavior of the concurrency mechanism.

1. Verify Queue Capacity

Submit more jobs than the queue can hold.

Confirm that jobs beyond the configured capacity are rejected rather than silently accumulating.

2. Verify the Worker Limit

Generate sustained input and track the number of active operations.

The measured peak should never exceed the configured worker count.

3. Verify Queued Cancellation

Cancel a request while it is still waiting.

Confirm that the job is skipped when the worker dequeues it.

4. Verify Active Cancellation

Cancel a request while its operation is running.

Confirm that the operation observes cancellation and that the active count returns to its previous value.

5. Verify Shutdown

Test application shutdown while jobs are waiting and while operations are active.

A production application needs an explicit policy for:

  • Draining accepted work.

  • Canceling active work.

  • Handling jobs that remain after a shutdown deadline.

  • Reporting failures from worker tasks.

What Needs to Change in a Production AI Application?

A bounded queue and worker pool are only one part of production AI request management.

Limit Payload Size

Queue length alone does not necessarily bound memory usage.

One request might contain a small prompt while another might contain a very large document.

Therefore, consider admission limits for:

  • Prompt size

  • Document size

  • Number of attached files

  • Estimated token count

  • Serialized request size

Consider Multiple Application Instances

The worker limit is local to one application process.

For example, if an application runs five instances and each instance has:

workerCount = 10

the theoretical aggregate concurrency can be much higher than ten.

Instance 1 → 10 workers
Instance 2 → 10 workers
Instance 3 → 10 workers
Instance 4 → 10 workers
Instance 5 → 10 workers
                  ↓
            Potentially 50
          concurrent operations

Therefore, provider quotas and shared workload limits may require a distributed coordination mechanism or centralized rate-limiting strategy.

Handle Provider Rate Limits Separately

A worker limit does not establish a provider's requests-per-minute or token-per-minute limit.

For example:

Local Worker Limit
        ↓
   2 concurrent calls

Provider Limit
        ↓
   60 requests/minute

These controls solve different problems.

An application may need both.

Bound Retries

Retries also consume capacity.

A rate-limit response should not cause an unlimited number of new tasks or immediate retry storms.

Use:

  • A bounded retry count

  • Appropriate retry delays

  • Exponential backoff where appropriate

  • Provider-specific rate-limit guidance

  • Cancellation support

The goal is to prevent retries from turning temporary provider pressure into additional application pressure.

What Should You Measure?

A production queue should expose enough telemetry to explain where users are spending time.

Useful metrics include:

  • Queue depth

  • Queue wait time

  • Admission rejections

  • Active worker count

  • Peak concurrency

  • Execution duration

  • Cancellation count

  • Failure count

  • Retry count

  • Provider rate-limit responses

  • Token consumption where available

These measurements help distinguish different problems.

For example:

High queue wait
      +
Low worker utilization
      ↓
Investigate scheduling or dependency delays

Whereas:

High queue wait
      +
Workers continuously busy
      ↓
Execution capacity may be too low

The solution should be based on observed workload rather than simply increasing the queue size.

Important Production Considerations

The console application intentionally keeps the example small. A hosted production service needs additional decisions.

Graceful Shutdown

The application should define what happens to:

  • Jobs waiting in the queue

  • Active AI calls

  • Requests that have already canceled

  • Jobs that exceed the shutdown deadline

A shutdown token alone does not define the complete shutdown policy.

Worker Exceptions

The sample handles expected cancellation but does not implement a general worker-failure strategy.

Production systems should explicitly decide whether an unexpected exception should:

  • Fail the individual job and continue processing.

  • Restart the worker.

  • Stop the worker service.

  • Trigger an alert.

Unexpected failures should not be treated as successful cancellations.

Fairness

A bounded queue can still create fairness problems.

For example, one user or tenant could potentially consume most of the available queue capacity.

Depending on the application, consider:

  • Per-user limits

  • Per-tenant limits

  • Priority policies

  • Separate queues

  • Fair scheduling

User Feedback

A queue changes the user experience.

Instead of displaying an indefinite spinner, the application can communicate meaningful states such as:

Request accepted
      ↓
Waiting for capacity
      ↓
Processing
      ↓
Completed

If the queue is full:

Request rejected
      ↓
Retry later

This makes the system's behavior more predictable for users.

Queue Limits Are Not AI Provider Limits

This distinction is particularly important when building AI applications.

A local bounded queue answers:

How much work should this application accept and retain?

A worker limit answers:

How many AI operations should this application execute concurrently?

A rate limiter answers:

How quickly should requests be started?

A provider quota answers:

How much usage does the external AI service allow?

These controls can work together:

                 Incoming Requests
                         |
                         ↓
                Admission Control
                         |
                         ↓
                  Bounded Queue
                         |
                         ↓
                 Worker Pool
                         |
                         ↓
                  Rate Limiter
                         |
                         ↓
                   AI Provider

Each layer controls a different part of the workload.

Conclusion

AI applications can become difficult to operate when many users generate requests simultaneously. Simply creating asynchronous tasks does not provide workload control. A large number of waiting operations can still consume memory and make the system difficult to reason about.

A bounded queue limits accepted waiting work, while a fixed worker count limits active operations. Cancellation handling prevents abandoned requests from unnecessarily consuming execution capacity, while explicit admission behavior gives the application a predictable response when capacity is exhausted.

However, these local controls do not replace provider rate limits, token budgets, distributed coordination, payload limits, or retry policies.

The key design principle is to separate these concerns:

Queue capacity controls waiting work.

Worker count controls concurrent execution.

Rate limiting controls request frequency.

Provider quotas control external service usage.

Once these boundaries are explicit, an AI application becomes easier to test, measure, and operate under real-world concurrency.