An AI-provider circuit breaker temporarily blocks new calls after a configured pattern of failures, then permits a recovery probe. In .NET, a shared Polly resilience pipeline can implement that behavior around the provider call. The breaker must persist across requests; creating a new breaker for every call discards the failure history it needs.
Suppose an assistant keeps sending generation requests while its provider is returning errors. Users wait, infrastructure remains busy, and repeated attempts obscure the fact that the dependency is unhealthy. A breaker gives the application a defined unavailable state and a controlled route back to service.
This tutorial uses a scripted HTTP handler, so no external provider is contacted. It tests opening, blocking, a failed recovery probe, and successful recovery.
How is a circuit breaker different from a retry or timeout?
These mechanisms solve different problems.
Mechanism |
Purpose |
Question it answers |
Timeout |
Bound a call's duration |
How long may this attempt take? |
Retry |
Make another attempt after a selected failure |
Should this operation be tried again? |
Circuit breaker |
Stop new attempts while a dependency appears unhealthy |
Should this call reach the dependency now? |
Polly's circuit-breaker documentation describes failure sampling, an open period, and recovery probes. A breaker does not itself retry a failed operation.
The example deliberately adds no retry strategy. That makes provider-call counts and state changes easy to interpret. A production combination needs a deliberate order and a clear definition of whether the breaker observes individual attempts or final operation outcomes.
What do you need to run the example?
The code was compiled with SDK 8.0.414 and C# 12, then executed on .NET 8.0.20 with Polly.Core 8.6.4 on Linux. These are the tested versions, not a claim that they are the newest releases.
Use a .NET 8 console project with a Polly.Core package reference pinned to 8.6.4. Replace Program.cs with the code below and run the project. The package's NuGet entry identifies the version used.
The fake handler returns four scripted statuses: 503, 503, 503, and 200. The extra application calls made while the breaker is open should never reach that handler.
Build one pipeline and reuse it
using System;
using System.Collections.Generic;
using System.Net;
using System.Net.Http;
using System.Threading;
using System.Threading.Tasks;
using Polly;
using Polly.CircuitBreaker;
var state = new CircuitBreakerStateProvider();
var transitions = new List<string>();
var pipeline = new ResiliencePipelineBuilder<HttpResponseMessage>()
.AddCircuitBreaker(new CircuitBreakerStrategyOptions<HttpResponseMessage>
{
FailureRatio = 1.0,
MinimumThroughput = 2,
SamplingDuration = TimeSpan.FromSeconds(10),
BreakDuration = TimeSpan.FromSeconds(1),
ShouldHandle = new PredicateBuilder<HttpResponseMessage>()
.Handle<HttpRequestException>()
.HandleResult(r => (int)r.StatusCode >= 500),
StateProvider = state,
OnOpened = _ => { transitions.Add("Open"); return default; },
OnHalfOpened = _ => { transitions.Add("HalfOpen"); return default; },
OnClosed = _ => { transitions.Add("Closed"); return default; }
}).Build();
using var handler = new ScriptedHandler(503, 503, 503, 200);
using var client = new HttpClient(handler);
var observed = new List<string>();
async Task Call(string label)
{
try
{
using var response = await pipeline.ExecuteAsync(async token =>
{
using var request = new HttpRequestMessage(
HttpMethod.Post, "https://provider.invalid/generate");
request.Content = new StringContent("{\"input\":\"synthetic\"}");
return await client.SendAsync(request, token);
}, CancellationToken.None);
observed.Add($"{label}: {(int)response.StatusCode}; {state.CircuitState}");
}
catch (BrokenCircuitException)
{
observed.Add($"{label}: blocked; {state.CircuitState}");
}
}
await Call("first");
await Call("second");
await Call("while-open");
if (handler.Calls != 2) throw new Exception("Open circuit reached provider");
await Task.Delay(1200);
await Call("failed-probe");
await Call("blocked-again");
if (handler.Calls != 3) throw new Exception("Failed probe did not reopen");
await Task.Delay(1200);
await Call("successful-probe");
string events = string.Join(",", transitions);
if (events != "Open,HalfOpen,Open,HalfOpen,Closed" || handler.Calls != 4)
throw new Exception("Unexpected state transitions");
foreach (string line in observed) Console.WriteLine(line);
Console.WriteLine("transitions:events");Console.WriteLine("provider calls: {handler.Calls}");
sealed class ScriptedHandler(params int[] statuses) : HttpMessageHandler
{
public int Calls { get; private set; }
protected override Task<HttpResponseMessage> SendAsync(
HttpRequestMessage request, CancellationToken cancellationToken)
{
cancellationToken.ThrowIfCancellationRequested();
return Task.FromResult(new HttpResponseMessage((HttpStatusCode)statuses[Calls++]));
}
}
The pipeline counts HTTP 5xx results and HttpRequestException failures. In this fixture, two failures within the ten-second sampling window meet the configured failure ratio and minimum throughput. The one-second break duration keeps the demonstration short; it is a test setting, not a production recommendation.
The callback creates a fresh request and returns the response. The caller disposes the returned response. When the circuit blocks execution, Polly throws BrokenCircuitException before the callback reaches the handler.
The provider.invalid address is intercepted by ScriptedHandler. Do not replace the handler and assume this URL is a real service endpoint.
Why test a failed recovery probe?
Recovery is not simply a timer expiring. After the break duration, a permitted probe must establish whether the dependency has recovered. The example's first probe fails, so the circuit opens again. Its next probe succeeds, so normal execution can resume.
An engineering note prepared for Ranknod could use this sequence to explain the difference between “waiting long enough” and “observing recovery.” The distinction matters because an unavailable dependency can remain unavailable after a pause.
The transition callbacks record the half-open state, which may be too brief to observe by inspecting the state only after the call completes. An after-call state of Closed does not mean the probe skipped the half-open transition.
What did the executed example produce?
first: 503; Closed
second: 503; Open
while-open: blocked; Open
failed-probe: 503; Open
blocked-again: blocked; Open
successful-probe: 200; Closed
transitions: Open,HalfOpen,Open,HalfOpen,Closed
provider calls: 4
The assertions verify that blocked calls do not increment the provider count and that the expected transitions occur. This is evidence about the local fixture, not about a real provider's recovery time or availability.
The short delays make the sequential demonstration readable. For a larger test suite, use a supported controllable clock or testing facility for the selected library version, and add concurrent-call cases. Neither high concurrency nor distributed behavior was measured here.
Which failures should count in a real AI application?
Count failures that match the failure domain you intend to protect. A provider outage and an invalid user request should not automatically have the same effect on a shared circuit.
This example excludes 4xx results. An actual integration may treat 429 rate limits separately by honoring provider guidance and controlling admission. Do not copy the predicate without reviewing the provider's error contract, quota scope, and retry instructions.
An authentication error may require a configuration correction rather than repeated probes. Caller cancellation should also be distinguished from evidence that the provider is unhealthy.
Microsoft's HTTP resilience guidance documents combinations of retries, timeouts, and breakers. It also highlights the risk of retrying operations whose effects can be duplicated. For AI calls, review both SDK-level and application-level retries before adding another layer.
Where should the breaker live?
Reuse a pipeline for a defined provider failure domain, such as a deployment and credential or quota scope. Avoid a global circuit that lets an isolated failure unnecessarily block unrelated routes. Also avoid unbounded creation of circuits from arbitrary user-controlled keys.
A process-local pipeline protects that process. Several application instances normally maintain separate local histories unless you deliberately introduce coordination. The example makes no claim of cluster-wide protection.
Open-circuit handling belongs in the user flow. Return a clear temporary-unavailability result, preserve safe work where appropriate, and provide the next action. If you use a fallback model, assess its capabilities and data handling explicitly; a fallback response should not silently imply identical behavior.
What should you verify before deployment?
Verify pipeline lifetime, failure classification, cancellation, timeout behavior, and response disposal in the actual integration. Test that a blocked call performs no provider work and that a failed probe leaves the unavailable state intact.
Monitor state transitions and blocked-call counts without logging prompts or credentials. A circuit that never opens may have an unreachable minimum-throughput setting; a circuit that repeatedly reopens may indicate a persistent dependency or configuration problem.
The implementation is useful when it makes failure behavior predictable. The goal is a controlled decision about new calls while the provider is unhealthy, followed by recovery supported by an observed probe.
Join the conversation! Your thoughts help the community grow.