Introduction
AI features in web applications often start with a simple request-response pattern.
A user enters a question, the application sends it to an AI model, waits for the response, and then displays the result.
That approach works for short requests, but it becomes less useful when an AI agent needs several seconds to reason, call tools, retrieve information, or generate a long response.
Blazor applications can provide a better experience by streaming AI output into the interface as it becomes available.
With the current AI capabilities available across the .NET ecosystem, developers can build Blazor interfaces that display streamed model responses, tool activity, status updates, and completed results without waiting for the entire operation to finish.
This article explains how to structure a streaming AI experience in Blazor, how components should handle incremental responses, how tool results can be displayed, and what to consider for production applications.
Why Stream AI Responses?
Consider a traditional AI request:
User
|
v
Blazor
|
v
AI Service
|
| 10 seconds
v
Complete Response
|
v
Blazor UI
The user sees nothing during those ten seconds.
Streaming changes the interaction:
User
|
v
Blazor
|
v
AI Service
|
+---- "Let me check..."
|
+---- "I found..."
|
+---- "The results show..."
|
v
Complete Response
The UI can render each piece as it arrives.
This is particularly useful for:
AI chat
Document analysis
Coding assistants
Research assistants
Agent workflows
Long responses
Tool-enabled agents
What Are Blazor AI Components?
A Blazor AI component is a UI component responsible for presenting an AI interaction.
A simple component might contain:
Chat Component
|
+---- Message List
|
+---- Input
|
+---- Send Button
|
+---- Streaming Response
|
+---- Tool Activity
The component should not need to understand the internal implementation of the AI model.
Instead, it receives structured application events and updates its state.
Basic Blazor Chat Component
A simple Blazor component could start with:
<div class="chat">
@foreach (var message in Messages)
{
<div class="message">
<strong>@message.Role</strong>
<p>@message.Content</p>
</div>
}
</div>
<input @bind="UserMessage" />
<button @onclick="SendMessageAsync">Send</button>
The component maintains a collection of messages:
private readonly List<ChatMessage> Messages = [];
private string UserMessage = string.Empty;
public record ChatMessage(
string Role,
string Content);
The next step is supporting incremental responses.
Streaming Instead of Waiting
A streaming service might expose an asynchronous sequence of response fragments.
Conceptually:
await foreach (var chunk in agent.StreamAsync(
UserMessage,
cancellationToken))
{
CurrentResponse += chunk;
await InvokeAsync(StateHasChanged);
}
The important difference is that the UI receives data continuously rather than waiting for one final response.
The exact streaming API depends on the AI framework being used.
Updating the UI Incrementally
Blazor components maintain UI state.
When streaming data arrives, update that state and request a render.
For example:
private string CurrentResponse = string.Empty;
private async Task AppendResponseAsync(string chunk)
{
CurrentResponse += chunk;
await InvokeAsync(StateHasChanged);
}
The UI can display:
Here is
then:
Here is the answer
and eventually:
Here is the answer to your question...
This creates the familiar streaming-chat experience.
Why InvokeAsync Matters
AI responses may arrive from an asynchronous operation that is not executing directly inside the Blazor component's normal UI event flow.
Calling:
await InvokeAsync(StateHasChanged);
helps marshal the UI update through the appropriate Blazor synchronization context.
This is especially important in Blazor Server applications where UI updates occur through an active circuit.
Streaming a Complete Message
A useful pattern is to create the assistant message before streaming begins.
private async Task SendMessageAsync()
{
if (string.IsNullOrWhiteSpace(UserMessage))
return;
Messages.Add(
new ChatMessage("User", UserMessage));
var assistantIndex = Messages.Count;
Messages.Add(
new ChatMessage("Assistant", string.Empty));
var prompt = UserMessage;
UserMessage = string.Empty;
await foreach (var chunk in agent.StreamAsync(prompt))
{
var current = Messages[assistantIndex];
Messages[assistantIndex] = current with
{
Content = current.Content + chunk
};
await InvokeAsync(StateHasChanged);
}
}
The assistant message starts empty and is progressively populated.
This avoids creating a separate UI element for every token or response fragment.
Why Not Create One Component Per Chunk?
Suppose the model produces:
Hello
,
how
can
I
help
?
Creating a separate UI element for every fragment would create unnecessary rendering work.
Instead, maintain one logical assistant message:
Hello, how can I help?
and append incoming fragments to that message.
This produces a cleaner component state model.
Tool-Enabled AI Responses
AI agents often need tools.
For example:
User
|
v
Agent
|
+---- Search Orders
|
+---- Query Database
|
+---- Call API
|
v
Final Response
The UI should distinguish between model-generated text and tool activity.
For example:
Assistant
Checking your recent orders...
Tool
Searching order history
Tool
3 matching orders found
Assistant
You placed three orders this month.
This provides useful transparency without exposing private model reasoning.
Representing Tool Events
Create an application-level event model:
public abstract record AgentEvent;
public record TextDelta(string Text)
: AgentEvent;
public record ToolStarted(string Name)
: AgentEvent;
public record ToolCompleted(
string Name,
string Result)
: AgentEvent;
The Blazor component can then handle each event type.
await foreach (var item in agent.RunAsync(
prompt,
cancellationToken))
{
switch (item)
{
case TextDelta text:
AppendText(text.Text);
break;
case ToolStarted tool:
AddToolActivity(tool.Name);
break;
case ToolCompleted tool:
CompleteToolActivity(tool.Name);
break;
}
await InvokeAsync(StateHasChanged);
}
This gives the UI a structured representation of the agent's activity.
Tool Activity UI
A tool activity component might display:
Searching customer orders...
When the operation completes:
Searching customer orders
Completed
The UI can then collapse or retain the activity depending on the application design.
For enterprise applications, tool visibility can also help users understand why an answer took longer than expected.
Avoid Showing Private Model Reasoning
There is an important distinction between:
Application-level progress
Tool execution
Private model reasoning
A UI can safely display:
Searching the order database...
without displaying private internal reasoning such as hidden deliberation.
The frontend should receive only information intentionally designed for the user.
Streaming With Cancellation
AI operations can be expensive.
Users should be able to stop a request.
Use a CancellationTokenSource:
private CancellationTokenSource? _cancellation;
private async Task SendMessageAsync()
{
_cancellation?.Cancel();
_cancellation = new CancellationTokenSource();
try
{
await StreamResponseAsync(
_cancellation.Token);
}
catch (OperationCanceledException)
{
// Request was cancelled.
}
}
The cancel button can call:
private void Cancel()
{
_cancellation?.Cancel();
}
The cancellation token must also reach the model and tool calls for cancellation to be effective.
Handling Errors During Streaming
Streaming creates a different error model from ordinary APIs.
A failure can occur after the user has already received part of a response.
For example:
Assistant:
The analysis shows that...
Tool call
Database query
ERROR
Database unavailable
The UI should not simply erase the partial response.
Instead, preserve useful output and display an appropriate error state.
For example:
The analysis shows that...
Unable to complete the database lookup.
Please try again.
This provides a much better user experience.
Disabling Duplicate Requests
A common UI problem is allowing users to click Send multiple times while a request is already running.
Track the request state:
private bool IsRunning;
Then:
<button
@onclick="SendMessageAsync"
disabled="@IsRunning">
@(IsRunning ? "Working..." : "Send")
</button>
This prevents accidental concurrent requests.
Managing Conversation History
An AI chat component usually needs conversation state.
For example:
private readonly List<ChatMessage> Messages = [];
The backend should not blindly trust the complete conversation sent by the browser.
For sensitive applications, maintain server-side conversation state or validate the messages and permissions associated with each request.
The architecture might be:
Blazor
|
v
ASP.NET Core
|
+---- Conversation State
|
+---- Agent
|
+---- Tools
Rendering Markdown Safely
AI responses frequently contain Markdown.
For example:
## Results
The API returned three orders.
Rendering Markdown improves readability, but generated content should not be inserted into the DOM as raw HTML without appropriate sanitization.
The application should use a trusted Markdown-rendering approach and sanitize HTML when HTML output is permitted.
Never assume AI-generated content is safe simply because it came from your own model.
Long Responses and Rendering Performance
A long AI response can produce many UI updates.
Updating the component for every tiny fragment may cause unnecessary rendering overhead.
A practical strategy is to buffer small fragments.
For example:
Incoming chunks
|
v
Buffer
|
| periodic flush
v
Blazor UI
Instead of rendering for every fragment, update the UI at a controlled interval.
The ideal frequency depends on response rate and application complexity.
Blazor Server vs Blazor WebAssembly
The streaming architecture can differ depending on the Blazor hosting model.
Blazor Server
The component runs on the server and communicates with the browser through the Blazor connection.
Advantages include:
Direct server-side access
Easier access to backend services
Smaller client payload
But network connectivity affects the interactive circuit.
Blazor WebAssembly
The application runs in the browser and communicates with backend APIs.
A typical architecture is:
Blazor WebAssembly
|
| HTTP / Streaming
v
ASP.NET Core API
|
v
AI Agent
This can provide a clean separation between UI and AI infrastructure.
The right choice depends on the application's architecture, security requirements, and deployment model.
Authentication and Authorization
AI applications often access private business data.
For example:
User
|
v
Blazor
|
v
Agent
|
+---- Customer Data
+---- Orders
+---- Documents
The agent should operate within the user's authorization boundary.
Do not rely solely on instructions such as:
Only show the user's own orders.
Authorization must be enforced by the application and underlying tools.
For example:
User Identity
|
v
Authorization
|
v
Tool
|
v
Authorized Data
Observability
Streaming AI applications need telemetry beyond normal HTTP logging.
Track:
Request duration
Time to first token or response fragment
Total response duration
Tool execution time
Model errors
Cancellation rate
Connection failures
Token consumption where available
A useful metric is time to first response.
A request that takes 15 seconds but starts streaming after 500 milliseconds can feel much more responsive than one that remains blank for 10 seconds.
Common Mistakes
Waiting for the Complete Response
This removes much of the benefit of streaming.
Updating the UI for Every Tiny Fragment
Excessive rendering can reduce performance.
Ignoring Cancellation
Users should be able to stop long-running requests.
Exposing Internal Agent Information
Only stream user-relevant events.
Trusting Client-Side Authorization
Protect sensitive tools and data on the server.
Rendering AI HTML Without Sanitization
Generated content must be treated as untrusted input.
Allowing Multiple Concurrent Requests
Track the active operation and prevent accidental duplicate submissions.
Troubleshooting
Response Appears Only at the End
Check whether the backend is actually streaming or buffering the entire response before returning it.
Also inspect proxy and hosting configuration for response buffering.
UI Does Not Refresh
Verify that asynchronous updates use:
await InvokeAsync(StateHasChanged);
when required by the Blazor execution context.
Tool Results Do Not Appear
Check:
Tool event generation
Event serialization
Frontend event handling
Component state updates
Cancel Does Not Stop the Agent
Ensure the same CancellationToken is passed through the entire operation.
UI Becomes Slow During Long Responses
Reduce render frequency and buffer incoming fragments.
Best Practices
Stream AI output instead of waiting for the complete response.
Maintain one logical assistant message while it is being generated.
Represent tool calls as structured UI events.
Keep private model reasoning out of the UI.
Support cancellation.
Prevent duplicate requests.
Handle partial responses gracefully.
Sanitize generated HTML.
Enforce authorization on the server.
Monitor time to first response.
Buffer high-frequency response fragments when necessary.
Keep conversation state controlled and validated.
Test interrupted network connections.
Test long responses and concurrent users.
Keep model and tool failures visible through meaningful application states.
Advantages and Disadvantages
Advantages
More responsive AI interfaces
Better user experience for long responses
Natural support for agent workflows
Tool activity can be displayed in real time
Users receive feedback while work is in progress
Cancellation can stop expensive operations
Fits well with Blazor's component model
Disadvantages
More complex state management
Requires careful cancellation handling
Network failures become more visible
Excessive UI updates can affect performance
Security is more complicated when agents access user data
Streaming infrastructure can require additional testing
When Should You Use Streaming AI in Blazor?
Streaming is particularly useful when:
AI responses are long
Model latency is noticeable
Agents call external tools
Users need progress feedback
Operations may take several seconds
The application behaves like an AI assistant
For a simple AI operation that consistently returns a short response quickly, traditional request-response communication may be sufficient.
Summary
Blazor provides a strong component model for building interactive AI experiences, while streaming allows those components to display responses as they are generated.
A well-designed implementation maintains a logical assistant message, appends response fragments incrementally, represents tool calls as structured events, supports cancellation, and handles failures without losing useful partial output.
The backend should remain responsible for authentication, authorization, tool permissions, and access to sensitive data. The frontend should focus on presenting safe, meaningful application events.
For developers building AI assistants with .NET, combining Blazor components with streaming AI responses can turn a basic chatbot into a responsive application that gives users useful feedback throughout the entire agent workflow.

Join the conversation! Your thoughts help the community grow.