Introduction

AI features in web applications often start with a simple request-response pattern.

A user enters a question, the application sends it to an AI model, waits for the response, and then displays the result.

That approach works for short requests, but it becomes less useful when an AI agent needs several seconds to reason, call tools, retrieve information, or generate a long response.

Blazor applications can provide a better experience by streaming AI output into the interface as it becomes available.

With the current AI capabilities available across the .NET ecosystem, developers can build Blazor interfaces that display streamed model responses, tool activity, status updates, and completed results without waiting for the entire operation to finish.

This article explains how to structure a streaming AI experience in Blazor, how components should handle incremental responses, how tool results can be displayed, and what to consider for production applications.

Why Stream AI Responses?

Consider a traditional AI request:

User
 |
 v
Blazor
 |
 v
AI Service
 |
 | 10 seconds
 v
Complete Response
 |
 v
Blazor UI

The user sees nothing during those ten seconds.

Streaming changes the interaction:

User
 |
 v
Blazor
 |
 v
AI Service
 |
 +---- "Let me check..."
 |
 +---- "I found..."
 |
 +---- "The results show..."
 |
 v
Complete Response

The UI can render each piece as it arrives.

This is particularly useful for:

What Are Blazor AI Components?

A Blazor AI component is a UI component responsible for presenting an AI interaction.

A simple component might contain:

Chat Component
   |
   +---- Message List
   |
   +---- Input
   |
   +---- Send Button
   |
   +---- Streaming Response
   |
   +---- Tool Activity

The component should not need to understand the internal implementation of the AI model.

Instead, it receives structured application events and updates its state.

Basic Blazor Chat Component

A simple Blazor component could start with:

<div class="chat">
    @foreach (var message in Messages)
    {
        <div class="message">
            <strong>@message.Role</strong>
            <p>@message.Content</p>
        </div>
    }
</div>

<input @bind="UserMessage" />
<button @onclick="SendMessageAsync">Send</button>

The component maintains a collection of messages:

private readonly List<ChatMessage> Messages = [];

private string UserMessage = string.Empty;

public record ChatMessage(
    string Role,
    string Content);

The next step is supporting incremental responses.

Streaming Instead of Waiting

A streaming service might expose an asynchronous sequence of response fragments.

Conceptually:

await foreach (var chunk in agent.StreamAsync(
    UserMessage,
    cancellationToken))
{
    CurrentResponse += chunk;
    await InvokeAsync(StateHasChanged);
}

The important difference is that the UI receives data continuously rather than waiting for one final response.

The exact streaming API depends on the AI framework being used.

Updating the UI Incrementally

Blazor components maintain UI state.

When streaming data arrives, update that state and request a render.

For example:

private string CurrentResponse = string.Empty;

private async Task AppendResponseAsync(string chunk)
{
    CurrentResponse += chunk;

    await InvokeAsync(StateHasChanged);
}

The UI can display:

Here is

then:

Here is the answer

and eventually:

Here is the answer to your question...

This creates the familiar streaming-chat experience.

Why InvokeAsync Matters

AI responses may arrive from an asynchronous operation that is not executing directly inside the Blazor component's normal UI event flow.

Calling:

await InvokeAsync(StateHasChanged);

helps marshal the UI update through the appropriate Blazor synchronization context.

This is especially important in Blazor Server applications where UI updates occur through an active circuit.

Streaming a Complete Message

A useful pattern is to create the assistant message before streaming begins.

private async Task SendMessageAsync()
{
    if (string.IsNullOrWhiteSpace(UserMessage))
        return;

    Messages.Add(
        new ChatMessage("User", UserMessage));

    var assistantIndex = Messages.Count;

    Messages.Add(
        new ChatMessage("Assistant", string.Empty));

    var prompt = UserMessage;
    UserMessage = string.Empty;

    await foreach (var chunk in agent.StreamAsync(prompt))
    {
        var current = Messages[assistantIndex];

        Messages[assistantIndex] = current with
        {
            Content = current.Content + chunk
        };

        await InvokeAsync(StateHasChanged);
    }
}

The assistant message starts empty and is progressively populated.

This avoids creating a separate UI element for every token or response fragment.

Why Not Create One Component Per Chunk?

Suppose the model produces:

Hello
,
 how
 can
 I
 help
?

Creating a separate UI element for every fragment would create unnecessary rendering work.

Instead, maintain one logical assistant message:

Hello, how can I help?

and append incoming fragments to that message.

This produces a cleaner component state model.

Tool-Enabled AI Responses

AI agents often need tools.

For example:

User
 |
 v
Agent
 |
 +---- Search Orders
 |
 +---- Query Database
 |
 +---- Call API
 |
 v
Final Response

The UI should distinguish between model-generated text and tool activity.

For example:

Assistant
Checking your recent orders...

Tool
Searching order history

Tool
3 matching orders found

Assistant
You placed three orders this month.

This provides useful transparency without exposing private model reasoning.

Representing Tool Events

Create an application-level event model:

public abstract record AgentEvent;

public record TextDelta(string Text)
    : AgentEvent;

public record ToolStarted(string Name)
    : AgentEvent;

public record ToolCompleted(
    string Name,
    string Result)
    : AgentEvent;

The Blazor component can then handle each event type.

await foreach (var item in agent.RunAsync(
    prompt,
    cancellationToken))
{
    switch (item)
    {
        case TextDelta text:
            AppendText(text.Text);
            break;

        case ToolStarted tool:
            AddToolActivity(tool.Name);
            break;

        case ToolCompleted tool:
            CompleteToolActivity(tool.Name);
            break;
    }

    await InvokeAsync(StateHasChanged);
}

This gives the UI a structured representation of the agent's activity.

Tool Activity UI

A tool activity component might display:

Searching customer orders...

When the operation completes:

Searching customer orders
Completed

The UI can then collapse or retain the activity depending on the application design.

For enterprise applications, tool visibility can also help users understand why an answer took longer than expected.

Avoid Showing Private Model Reasoning

There is an important distinction between:

A UI can safely display:

Searching the order database...

without displaying private internal reasoning such as hidden deliberation.

The frontend should receive only information intentionally designed for the user.

Streaming With Cancellation

AI operations can be expensive.

Users should be able to stop a request.

Use a CancellationTokenSource:

private CancellationTokenSource? _cancellation;

private async Task SendMessageAsync()
{
    _cancellation?.Cancel();
    _cancellation = new CancellationTokenSource();

    try
    {
        await StreamResponseAsync(
            _cancellation.Token);
    }
    catch (OperationCanceledException)
    {
        // Request was cancelled.
    }
}

The cancel button can call:

private void Cancel()
{
    _cancellation?.Cancel();
}

The cancellation token must also reach the model and tool calls for cancellation to be effective.

Handling Errors During Streaming

Streaming creates a different error model from ordinary APIs.

A failure can occur after the user has already received part of a response.

For example:

Assistant:
The analysis shows that...

Tool call
Database query

ERROR
Database unavailable

The UI should not simply erase the partial response.

Instead, preserve useful output and display an appropriate error state.

For example:

The analysis shows that...

Unable to complete the database lookup.
Please try again.

This provides a much better user experience.

Disabling Duplicate Requests

A common UI problem is allowing users to click Send multiple times while a request is already running.

Track the request state:

private bool IsRunning;

Then:

<button
    @onclick="SendMessageAsync"
    disabled="@IsRunning">
    @(IsRunning ? "Working..." : "Send")
</button>

This prevents accidental concurrent requests.

Managing Conversation History

An AI chat component usually needs conversation state.

For example:

private readonly List<ChatMessage> Messages = [];

The backend should not blindly trust the complete conversation sent by the browser.

For sensitive applications, maintain server-side conversation state or validate the messages and permissions associated with each request.

The architecture might be:

Blazor
  |
  v
ASP.NET Core
  |
  +---- Conversation State
  |
  +---- Agent
  |
  +---- Tools

Rendering Markdown Safely

AI responses frequently contain Markdown.

For example:

## Results

The API returned three orders.

Rendering Markdown improves readability, but generated content should not be inserted into the DOM as raw HTML without appropriate sanitization.

The application should use a trusted Markdown-rendering approach and sanitize HTML when HTML output is permitted.

Never assume AI-generated content is safe simply because it came from your own model.

Long Responses and Rendering Performance

A long AI response can produce many UI updates.

Updating the component for every tiny fragment may cause unnecessary rendering overhead.

A practical strategy is to buffer small fragments.

For example:

Incoming chunks
    |
    v
Buffer
    |
    | periodic flush
    v
Blazor UI

Instead of rendering for every fragment, update the UI at a controlled interval.

The ideal frequency depends on response rate and application complexity.

Blazor Server vs Blazor WebAssembly

The streaming architecture can differ depending on the Blazor hosting model.

Blazor Server

The component runs on the server and communicates with the browser through the Blazor connection.

Advantages include:

But network connectivity affects the interactive circuit.

Blazor WebAssembly

The application runs in the browser and communicates with backend APIs.

A typical architecture is:

Blazor WebAssembly
       |
       | HTTP / Streaming
       v
ASP.NET Core API
       |
       v
AI Agent

This can provide a clean separation between UI and AI infrastructure.

The right choice depends on the application's architecture, security requirements, and deployment model.

Authentication and Authorization

AI applications often access private business data.

For example:

User
 |
 v
Blazor
 |
 v
Agent
 |
 +---- Customer Data
 +---- Orders
 +---- Documents

The agent should operate within the user's authorization boundary.

Do not rely solely on instructions such as:

Only show the user's own orders.

Authorization must be enforced by the application and underlying tools.

For example:

User Identity
      |
      v
Authorization
      |
      v
Tool
      |
      v
Authorized Data

Observability

Streaming AI applications need telemetry beyond normal HTTP logging.

Track:

A useful metric is time to first response.

A request that takes 15 seconds but starts streaming after 500 milliseconds can feel much more responsive than one that remains blank for 10 seconds.

Common Mistakes

Waiting for the Complete Response

This removes much of the benefit of streaming.

Updating the UI for Every Tiny Fragment

Excessive rendering can reduce performance.

Ignoring Cancellation

Users should be able to stop long-running requests.

Exposing Internal Agent Information

Only stream user-relevant events.

Trusting Client-Side Authorization

Protect sensitive tools and data on the server.

Rendering AI HTML Without Sanitization

Generated content must be treated as untrusted input.

Allowing Multiple Concurrent Requests

Track the active operation and prevent accidental duplicate submissions.

Troubleshooting

Response Appears Only at the End

Check whether the backend is actually streaming or buffering the entire response before returning it.

Also inspect proxy and hosting configuration for response buffering.

UI Does Not Refresh

Verify that asynchronous updates use:

await InvokeAsync(StateHasChanged);

when required by the Blazor execution context.

Tool Results Do Not Appear

Check:

  1. Tool event generation

  2. Event serialization

  3. Frontend event handling

  4. Component state updates

Cancel Does Not Stop the Agent

Ensure the same CancellationToken is passed through the entire operation.

UI Becomes Slow During Long Responses

Reduce render frequency and buffer incoming fragments.

Best Practices

  1. Stream AI output instead of waiting for the complete response.

  2. Maintain one logical assistant message while it is being generated.

  3. Represent tool calls as structured UI events.

  4. Keep private model reasoning out of the UI.

  5. Support cancellation.

  6. Prevent duplicate requests.

  7. Handle partial responses gracefully.

  8. Sanitize generated HTML.

  9. Enforce authorization on the server.

  10. Monitor time to first response.

  11. Buffer high-frequency response fragments when necessary.

  12. Keep conversation state controlled and validated.

  13. Test interrupted network connections.

  14. Test long responses and concurrent users.

  15. Keep model and tool failures visible through meaningful application states.

Advantages and Disadvantages

Advantages

Disadvantages

When Should You Use Streaming AI in Blazor?

Streaming is particularly useful when:

For a simple AI operation that consistently returns a short response quickly, traditional request-response communication may be sufficient.

Summary

Blazor provides a strong component model for building interactive AI experiences, while streaming allows those components to display responses as they are generated.

A well-designed implementation maintains a logical assistant message, appends response fragments incrementally, represents tool calls as structured events, supports cancellation, and handles failures without losing useful partial output.

The backend should remain responsible for authentication, authorization, tool permissions, and access to sensitive data. The frontend should focus on presenting safe, meaningful application events.

For developers building AI assistants with .NET, combining Blazor components with streaming AI responses can turn a basic chatbot into a responsive application that gives users useful feedback throughout the entire agent workflow.