Building an AI-powered feature in .NET can start with a single model request. The complexity increases when the application needs to call tools, preserve conversation history, handle failures, enforce execution limits, and return a reliable result to the user.

Without a consistent execution layer, each application ends up implementing these concerns differently. One service may invoke tools manually, another may keep conversation history in memory, and a third may retry model requests without limiting how much work an agent can perform.

Microsoft.Extensions.AI provides a common set of abstractions for working with AI services in .NET. Its IChatClient interface, function-calling support, and composable middleware make it possible to build a reusable agent harness without tying the application to a single model provider.

An agent harness is the runtime layer that coordinates model calls, tool execution, conversation state, cancellation, and error handling. It is not the model itself, and it does not automatically make an application autonomous or reliable. Those behaviors depend on the execution policies and tools that the application exposes.

This article builds a practical foundation for a .NET agent harness using Microsoft.Extensions.AI, with bounded tool execution and explicit control over the workflow.

What an Agent Harness Should Handle

A model client sends messages to an AI service and receives responses. An agent harness adds the application logic required to turn those exchanges into a controlled workflow.

A useful harness should handle several responsibilities:

The harness should also keep authorization outside the model's control. A model can suggest calling a function, but the application must decide whether that function is available and whether the current user is permitted to execute it.

For a small application, these responsibilities can fit into a lightweight service. Larger applications may need persistent sessions, distributed execution, human approval, and more sophisticated orchestration.

Understand the Microsoft.Extensions.AI Abstractions

The central abstraction is IChatClient. It provides a common interface for requesting model responses and supports both complete responses and streaming output.

A provider-specific implementation can be wrapped in a pipeline that adds functionality such as automatic function invocation, caching, or telemetry.

The main components are:

Component

Responsibility

IChatClient

Sends messages to an AI model and receives responses

ChatMessage

Represents a message in the conversation

ChatOptions

Configures request options and available tools

AIFunction

Represents an invokable function exposed to the model

AIFunctionFactory

Creates AI functions from .NET methods

FunctionInvokingChatClient

Executes supported function calls and continues the model interaction

ChatClientBuilder

Composes chat-client middleware into a pipeline

These components are available through the Microsoft.Extensions.AI libraries, although provider implementations may require additional NuGet packages.

The distinction between a chat client and a harness is important. IChatClient supplies the model interaction; the harness defines how the application uses that interaction to perform work.

Step 1: Create the .NET Project

Start with a console application:

dotnet new console -n AgentHarness
cd AgentHarness

Add Microsoft.Extensions.AI and a provider implementation. This example uses OllamaSharp to connect to a locally running Ollama server.

dotnet add package Microsoft.Extensions.AI
dotnet add package OllamaSharp

Make sure Ollama is installed and running, and that the selected model is available locally. The example uses llama3.1; replace it with a model that supports tool calling in your environment.

Using a local provider keeps the example independent of a hosted API account. The same harness design can work with another provider that implements IChatClient.

For a production application, pin package versions through your project's dependency-management process and verify compatibility between the provider package and Microsoft.Extensions.AI.

Step 2: Create a Tool the Agent Can Invoke

An agent becomes useful when it can retrieve information or perform a well-defined operation through application code.

For this example, define a simple function that retrieves the current service status from an application-controlled source. To keep the sample self-contained, the function returns a fixed demonstration result.

static string GetServiceStatus(string serviceName)
{
    var services = new Dictionary<string, string>(
        StringComparer.OrdinalIgnoreCase)
    {
        ["orders"] = "Healthy",
        ["payments"] = "Degraded",
        ["inventory"] = "Healthy"
    };

    return services.TryGetValue(serviceName, out var status)
        ? $"{serviceName}: {status}"
        : $"Unknown service: {serviceName}";
}

The function accepts a service name and returns the corresponding status.

In a real system, this function could query an internal health API, read monitoring data, or retrieve an authorized record from a database.

The important design decision is to expose a narrow operation instead of giving the model unrestricted access to infrastructure. A read-only status function is easier to reason about than a general-purpose function that can execute arbitrary commands.

The function should also validate its inputs and apply the caller's authorization rules. Tool definitions are not security boundaries by themselves.

Step 3: Build the Chat-Client Pipeline

Next, create the provider client and configure automatic function invocation.

using Microsoft.Extensions.AI;
using OllamaSharp;

IChatClient client = new ChatClientBuilder(
    new OllamaApiClient(
        new Uri("http://localhost:11434/"),
        "llama3.1"))
    .UseFunctionInvocation()
    .Build();

The provider-specific OllamaApiClient implements the chat-client abstraction. UseFunctionInvocation() adds the function-invocation component to the pipeline.

When the model requests an available function, the invocation component can execute it, send the result back to the model, and continue the interaction. This removes the need to manually implement every function-call round trip.

The model does not execute the C# function itself. It produces a structured request that the client pipeline resolves against the functions provided by the application.

This distinction matters because the application controls which functions are registered and how they execute.

Step 4: Register the Available Tool

Create an AIFunction from the method and pass it through ChatOptions.

var serviceStatusTool =
    AIFunctionFactory.Create(GetServiceStatus);

var options = new ChatOptions
{
    Tools = [serviceStatusTool]
};

The function factory uses the method's metadata to describe the function to the model. A useful method name, clear parameter names, and meaningful descriptions help the model decide when the tool is appropriate.

Only register tools that the current workflow needs. Providing every application function to every agent increases the number of available actions and can make tool selection less predictable.

For sensitive operations, add explicit approval or authorization checks before executing the function. A model-generated request should never be treated as proof that a user is entitled to perform an action.

Step 5: Implement the Harness With Conversation History

The next step is to create a reusable execution layer that maintains the conversation and calls the chat client.

The following implementation keeps conversation history in memory for the lifetime of a harness instance. It also accepts a cancellation token so callers can stop an ongoing model request.

using Microsoft.Extensions.AI;

public sealed class AgentHarness
{
    private readonly IChatClient _client;
    private readonly ChatOptions _options;

    private readonly List<ChatMessage> _history =
    [
        new ChatMessage(
            ChatRole.System,
            """
            You are an operations assistant.
            Use the available service-status tool when
            answering questions about service health.
            Distinguish retrieved status from assumptions.
            """)
    ];

    public AgentHarness(
        IChatClient client,
        ChatOptions options)
    {
        _client = client;
        _options = options;
    }

    public async Task<string> RunAsync(
        string userInput,
        CancellationToken cancellationToken = default)
    {
        if (string.IsNullOrWhiteSpace(userInput))
        {
            throw new ArgumentException(
                "Input cannot be empty.",
                nameof(userInput));
        }

        _history.Add(
            new ChatMessage(ChatRole.User, userInput));

        var response = await _client.GetResponseAsync(
            _history,
            _options,
            cancellationToken);

        _history.AddMessages(response);

        return response.Text ?? string.Empty;
    }
}

This harness provides three basic capabilities:

  1. It preserves messages across calls.

  2. It supplies the configured tools with each request.

  3. It propagates cancellation to the chat-client operation.

The call to AddMessages(response) appends the response messages to the history, preserving the assistant's response content and relevant function-call information provided by the response.

The function-invocation pipeline handles the model's tool-call cycle, while the harness manages the surrounding conversation.

This is a useful starting point, but it is not yet a complete production implementation. In particular, the history is not persisted, concurrent calls are not synchronized, and there is no explicit request-level budget or retention policy.

Step 6: Run the Agent

The harness can now be instantiated and used from the application's entry point.

using Microsoft.Extensions.AI;
using OllamaSharp;

IChatClient client = new ChatClientBuilder(
    new OllamaApiClient(
        new Uri("http://localhost:11434/"),
        "llama3.1"))
    .UseFunctionInvocation()
    .Build();

var tools = new ChatOptions
{
    Tools =
    [
        AIFunctionFactory.Create(GetServiceStatus)
    ]
};

var agent = new AgentHarness(client, tools);

var answer = await agent.RunAsync(
    "Is the payments service healthy?");

Console.WriteLine(answer);

The model can request GetServiceStatus with the payments argument, receive the tool result, and use that result to formulate a response.

The expected behavior depends on the selected model's tool-calling support and its interpretation of the instructions. The application should not assume that every model will always call the tool correctly.

For important operational decisions, validate the returned information independently and avoid treating a natural-language response as a substitute for the underlying service result.

The sample separates tool registration from conversation execution, making it easier to test the harness with a fake IChatClient or a different provider implementation.

Add Execution Limits and Failure Handling

Automatic function invocation introduces a loop: the model can request a tool, receive its result, and request another tool before producing a final answer.

That loop must be bounded. A model can repeatedly request tools, fail to reach a useful conclusion, or consume more time than the calling application allows.

FunctionInvokingChatClient exposes MaximumIterationsPerRequest, which limits the number of function-invocation iterations for a request.

You can configure the invocation component explicitly when you need tighter control:

using Microsoft.Extensions.AI;
using OllamaSharp;

var invocationClient = new FunctionInvokingChatClient(
    new OllamaApiClient(
        new Uri("http://localhost:11434/"),
        "llama3.1"))
{
    MaximumIterationsPerRequest = 5
};

IChatClient client = invocationClient;

This configuration limits consecutive function-invocation iterations. It is not a complete resource budget: it does not by itself establish an end-to-end wall-clock deadline, a maximum token cost, or a global limit across multiple user requests.

For production use, combine the iteration limit with:

Avoid adding broad automatic retries around the entire agent loop without considering side effects. If a tool successfully changes external state but the response is lost, repeating the whole operation can duplicate that change.

Read-only tools are easier to retry safely. Mutating tools should use idempotency keys, durable operation records, or an approval workflow where appropriate.

Preserve Conversation History Without Growing It Forever

The example stores all messages in a list. That is suitable for a short demonstration, but long-running conversations can consume increasing amounts of memory and exceed the model's context window.

A production harness should define a history policy.

Possible strategies include:

Do not blindly remove earlier messages if they contain tool-call results that are needed to maintain a valid conversation sequence. The history-reduction process must preserve the structural requirements of the selected provider and the information required for the next model call.

For multi-user applications, do not share a single mutable history list between requests. Use a separate session or conversation object for each independent interaction, and enforce authorization when loading persisted history.

A robust design separates the stateless chat-client pipeline from the session-specific state maintained by the harness.

Add Dependency Injection and Observability

As the application grows, the harness should be registered through .NET dependency injection instead of being constructed directly in every endpoint.

The application can register its provider-specific client, configure the function-invocation pipeline, and inject the resulting client into a scoped or otherwise appropriately managed harness.

Choose the lifetime carefully. A chat client may be reusable across requests, while conversation history should normally be scoped to an individual conversation. Do not make session-specific mutable state a singleton.

Microsoft.Extensions.AI also supports composable functionality such as OpenTelemetry instrumentation and caching. These capabilities can help teams observe model calls and identify latency, error, and usage patterns.

For example, telemetry can record:

Avoid enabling sensitive-content capture indiscriminately. Prompts, tool results, and conversation history may contain private or operationally sensitive information.

Common Implementation Mistakes

Treating the chat client as the entire agent: A chat client provides model interaction, while the harness manages the execution policy and session lifecycle.

Exposing too many tools: Register only the functions needed for the current task. Keep authorization and input validation in application code.

Allowing unbounded tool loops: Configure iteration limits and request deadlines so the agent cannot consume unlimited resources.

Sharing mutable conversation history: Separate state between conversations and protect concurrent access to each session.

Retrying side effects without idempotency: External operations can succeed even when the client does not receive the result. Use safe retry semantics.

Assuming the model will always use tools correctly: Tool selection depends on model behavior. Validate important results and handle missing or malformed tool requests.

Keeping every message forever: Implement a history-retention or compaction strategy that respects context limits and conversation semantics.

Summary

Microsoft.Extensions.AI provides the abstractions needed to build a reusable .NET agent harness without coupling application logic to a specific AI provider.

By combining IChatClient, ChatClientBuilder, AIFunctionFactory, and function-invocation support, developers can build an execution layer that preserves conversation history and lets a model invoke controlled C# functions.

A production harness needs additional safeguards: bounded tool execution, cancellation, session isolation, history management, idempotent side effects, and observability. Those concerns should remain explicit application responsibilities rather than being delegated to the model.

Start with a small harness and a read-only tool, then add persistence and operational controls as the workload requires them. This approach keeps the implementation understandable while providing a foundation that can evolve into a more capable agent runtime.