Large Language Models (LLMs) can be used directly through APIs, but generic models often struggle with highly specialized terminology, response formats, internal business processes, and domain-specific instructions.

Fine-tuning an open-source LLM provides another approach. Instead of changing the model's knowledge base entirely, fine-tuning adjusts the model's parameters so it learns a specific behavior, style, task, or domain pattern from a curated training dataset.

For .NET developers, a practical architecture is to use the Python ecosystem for model training and expose the resulting model through an inference API that can be consumed by an ASP.NET Core application.

This article demonstrates that architecture using:

The goal is not to train an LLM from scratch. Instead, we will fine-tune an existing open-source model using parameter-efficient fine-tuning and integrate it with an ASP.NET Core application.

1. What Is LLM Fine-Tuning?

Fine-tuning takes a pretrained language model and trains it further using a task-specific dataset.

A pretrained model may already understand general language patterns:

Customer: How do I reset my password?

Model:
You can reset your password using the account settings...

However, an enterprise application may require a specific response format:

Intent: PASSWORD_RESET
Priority: LOW
Action: SEND_RESET_LINK

A fine-tuned model can learn this type of behavior from examples.

Conceptually:

Pretrained LLM
      |
      v
Domain-Specific Dataset
      |
      v
Fine-Tuning
      |
      v
Specialized Model
      |
      v
Inference API
      |
      v
ASP.NET Core Application

Fine-tuning is particularly useful when the application needs consistent behavior rather than simply retrieving additional information.

2. Fine-Tuning vs. RAG

Fine-tuning and Retrieval-Augmented Generation solve different problems.

Requirement Fine-Tuning RAG
Learn response style Strong fit Limited
Learn structured output Strong fit Possible
Add frequently changing information Poor fit Strong fit
Company documentation Sometimes Strong fit
Domain-specific behavior Strong fit Partial
Private knowledge retrieval Not ideal Strong fit
Reduce repeated prompting Strong fit Partial
Real-time information No Yes

For example, suppose an insurance application needs to answer questions using documents that change every week.

RAG is generally more appropriate because the documents can remain outside the model and be retrieved at runtime.

If the application needs the model to consistently classify insurance claims into predefined categories, fine-tuning may be appropriate.

In many enterprise systems, the architecture is actually:

User Request
     |
     v
ASP.NET Core
     |
     +----> RAG Retrieval
     |
     +----> Fine-Tuned LLM
     |
     v
Final Response

The two techniques are not mutually exclusive.

3. Choosing an Open-Source Model

Before fine-tuning, select a model that matches the application's requirements.

Important factors include:

For experimentation, smaller instruction-tuned models are usually easier to work with than very large models.

The model should also support the training and inference libraries you intend to use.

For example:

Model
  |
  +-- Transformers compatibility
  +-- PEFT compatibility
  +-- Tokenizer availability
  +-- Quantization support
  +-- License suitable for your application

Do not select a model solely because it has a large parameter count.

A smaller model that is properly fine-tuned can be more practical for a specific production task.

4. Preparing the Training Dataset

The dataset is often more important than the training configuration.

For an instruction-following model, a dataset can contain examples such as:

{
  "instruction": "Classify the following support request.",
  "input": "I cannot access my account because I forgot my password.",
  "output": "PASSWORD_RESET"
}

Another example:

{
  "instruction": "Classify the following support request.",
  "input": "The application crashes whenever I upload a PDF.",
  "output": "APPLICATION_ERROR"
}

A larger dataset might contain thousands of carefully reviewed examples.

The dataset should be:

Poor-quality examples can teach the model incorrect behavior.

5. Instruction Dataset Format

A common approach is to convert each example into a conversational or instruction format.

For example:

### Instruction:
Classify the support request.

### Input:
The application crashes whenever I upload a PDF.

### Response:
APPLICATION_ERROR

During preprocessing, these examples are tokenized and converted into model-compatible tensors.

A JSONL file is convenient because each line represents one training example:

{"instruction":"Classify the request.","input":"I forgot my password.","output":"PASSWORD_RESET"}
{"instruction":"Classify the request.","input":"The application crashes.","output":"APPLICATION_ERROR"}
{"instruction":"Classify the request.","input":"I need to change my billing address.","output":"ACCOUNT_UPDATE"}

6. Install the Python Training Environment

The training portion can be isolated from the .NET application.

Create a Python environment:

python -m venv .venv

Activate it on Windows:

.venv\Scripts\activate

On Linux:

source .venv/bin/activate

Install the required libraries:

pip install torch
pip install transformers
pip install datasets
pip install peft
pip install accelerate

Depending on the model and hardware configuration, additional packages may be required for quantization.

7. Load the Dataset

Using Hugging Face Datasets:

from datasets import load_dataset

dataset = load_dataset(
    "json",
    data_files="training-data.jsonl"
)

print(dataset)

A dataset might be divided into training and validation sets:

dataset = dataset["train"].train_test_split(
    test_size=0.1,
    seed=42
)

train_dataset = dataset["train"]
validation_dataset = dataset["test"]

Keeping validation data separate is important.

The model should not be evaluated only against examples it has already seen during training.

8. Load the Tokenizer

The tokenizer converts text into token IDs understood by the model.

from transformers import AutoTokenizer

model_name = "YOUR_MODEL_NAME"

tokenizer = AutoTokenizer.from_pretrained(model_name)

if tokenizer.pad_token is None:
    tokenizer.pad_token = tokenizer.eos_token

The tokenizer must correspond to the selected model.

Using an incompatible tokenizer can produce incorrect input representations.

9. Format the Dataset

Create a function that converts each record into the format used for training.

def format_example(example):
    text = f"""
### Instruction:
{example["instruction"]}

### Input:
{example["input"]}

### Response:
{example["output"]}
"""
    return {"text": text.strip()}

Apply it to the dataset:

train_dataset = train_dataset.map(format_example)
validation_dataset = validation_dataset.map(format_example)

Tokenization can then be performed:

def tokenize(example):
    return tokenizer(
        example["text"],
        truncation=True,
        max_length=1024
    )

tokenized_train = train_dataset.map(
    tokenize,
    batched=True
)

tokenized_validation = validation_dataset.map(
    tokenize,
    batched=True
)

The max_length value should be selected according to the model and dataset.

10. Why LoRA Is Useful

Full fine-tuning modifies a large portion of the model's parameters.

That can require significant GPU memory.

LoRA, or Low-Rank Adaptation, takes a different approach.

Instead of updating all model parameters, LoRA introduces smaller trainable matrices into selected model layers.

Conceptually:

Original Model Parameters
          |
          | frozen
          v
   ----------------
   |              |
   |   LLM        |
   |              |
   ----------------
          |
      LoRA Layers
          |
          v
   Trainable Parameters

This significantly reduces the number of parameters that need to be trained.

The resulting adapter can then be stored separately from the original model.

11. Configure LoRA

Using PEFT:

from peft import LoraConfig, TaskType

lora_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM,
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    target_modules=[
        "q_proj",
        "v_proj"
    ]
)

Important parameters include:

r

Controls the rank of the LoRA matrices.

Higher values generally provide more trainable capacity but increase memory and computation requirements.

lora_alpha

Controls the scaling of the LoRA contribution.

lora_dropout

Adds regularization during training.

target_modules

Defines which model layers receive LoRA adapters.

The correct target module names depend on the selected model architecture.

12. Load the Model

from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto"
)

For larger models, quantization can be combined with parameter-efficient fine-tuning to reduce GPU memory requirements.

For example, a 4-bit training configuration can be used when supported by the model and environment.

The exact configuration should be tested against the available GPU because hardware compatibility differs between environments.

13. Configure Training

Using the Transformers training stack:

from transformers import TrainingArguments

training_args = TrainingArguments(
    output_dir="./fine-tuned-model",
    num_train_epochs=3,
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,
    learning_rate=2e-4,
    logging_steps=10,
    save_steps=100,
    evaluation_strategy="steps",
    eval_steps=100,
    fp16=True,
    report_to="none"
)

These values are examples, not universal recommendations.

Training parameters should be tuned based on:

14. Train the Model

The model can then be wrapped with the LoRA configuration and trained.

from peft import get_peft_model

model = get_peft_model(
    model,
    lora_config
)

model.print_trainable_parameters()

A training loop can then be created using the Hugging Face training infrastructure.

For example:

from transformers import Trainer

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_train,
    eval_dataset=tokenized_validation
)

trainer.train()

After training:

trainer.save_model("./fine-tuned-model")
tokenizer.save_pretrained("./fine-tuned-model")

The resulting directory contains the trained adapter and supporting model configuration.

15. Evaluate the Fine-Tuned Model

Training loss alone is not sufficient to determine whether the model is useful.

Create a separate evaluation dataset.

For classification:

Input:
My password reset link has expired.

Expected:
PASSWORD_RESET

For structured generation:

{
  "category": "PASSWORD_RESET",
  "priority": "LOW"
}

Evaluate metrics appropriate to the task.

For classification:

For text generation:

For enterprise systems, task-specific evaluation is usually more meaningful than a generic language benchmark.

16. Build an Inference API

The easiest architecture for ASP.NET Core integration is to expose the model through an HTTP API.

For example:

ASP.NET Core
     |
     | HTTP POST
     v
Python Inference API
     |
     v
Fine-Tuned LLM

The inference service can be implemented using FastAPI.

Install it:

pip install fastapi uvicorn

Create:

from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

class PredictionRequest(BaseModel):
    prompt: str

The inference implementation loads the model and generates a response.

A simplified endpoint might look like:

@app.post("/generate")
def generate(request: PredictionRequest):

    inputs = tokenizer(
        request.prompt,
        return_tensors="pt"
    )

    outputs = model.generate(
        **inputs,
        max_new_tokens=256
    )

    result = tokenizer.decode(
        outputs[0],
        skip_special_tokens=True
    )

    return {
        "response": result
    }

For production workloads, inference code should also handle:

17. Create the ASP.NET Core Application

Create an ASP.NET Core Web API:

dotnet new webapi -n LlmIntegrationApi
cd LlmIntegrationApi

The application will act as the client of the inference service.

The architecture becomes:

Client
   |
   v
ASP.NET Core Web API
   |
   | HTTP
   v
LLM Inference Service
   |
   v
Fine-Tuned Open-Source Model

This separation allows the model runtime to remain independent of the .NET application.

18. Create Request and Response Models

Create a request DTO:

public sealed class LlmRequest
{
    public string Prompt { get; set; } = string.Empty;
}

Create a response DTO:

public sealed class LlmResponse
{
    public string Response { get; set; } = string.Empty;
}

19. Create an LLM Client Service

Instead of placing HTTP logic directly inside a controller, create a dedicated service.

public interface ILlmClient
{
    Task<string> GenerateAsync(
        string prompt,
        CancellationToken cancellationToken = default);
}

Implementation:

using System.Net.Http.Json;

public sealed class LlmClient : ILlmClient
{
    private readonly HttpClient _httpClient;

    public LlmClient(HttpClient httpClient)
    {
        _httpClient = httpClient;
    }

    public async Task<string> GenerateAsync(
        string prompt,
        CancellationToken cancellationToken = default)
    {
        var request = new LlmRequest
        {
            Prompt = prompt
        };

        var response = await _httpClient.PostAsJsonAsync(
            "/generate",
            request,
            cancellationToken);

        response.EnsureSuccessStatusCode();

        var result =
            await response.Content.ReadFromJsonAsync<LlmResponse>(
                cancellationToken: cancellationToken);

        return result?.Response ?? string.Empty;
    }
}

This keeps the application architecture clean and makes the model provider replaceable.

20. Register the HTTP Client

In Program.cs:

builder.Services.AddHttpClient<ILlmClient, LlmClient>(
    client =>
    {
        client.BaseAddress =
            new Uri("http://localhost:8000");
    });

In production, the endpoint should come from configuration rather than being hard-coded.

For example:

{
  "LlmService": {
    "BaseUrl": "http://llm-service:8000"
  }
}

Then:

builder.Services.AddHttpClient<ILlmClient, LlmClient>(
    (serviceProvider, client) =>
    {
        var configuration =
            serviceProvider.GetRequiredService<IConfiguration>();

        var baseUrl =
            configuration["LlmService:BaseUrl"];

        client.BaseAddress =
            new Uri(baseUrl!);
    });

21. Create the ASP.NET Core Controller

using Microsoft.AspNetCore.Mvc;

[ApiController]
[Route("api/[controller]")]
public class AiController : ControllerBase
{
    private readonly ILlmClient _llmClient;

    public AiController(ILlmClient llmClient)
    {
        _llmClient = llmClient;
    }

    [HttpPost("generate")]
    public async Task<IActionResult> Generate(
        [FromBody] LlmRequest request,
        CancellationToken cancellationToken)
    {
        if (string.IsNullOrWhiteSpace(request.Prompt))
        {
            return BadRequest("Prompt is required.");
        }

        var response =
            await _llmClient.GenerateAsync(
                request.Prompt,
                cancellationToken);

        return Ok(new
        {
            response
        });
    }
}

The client can now send:

POST /api/ai/generate
Content-Type: application/json

with:

{
  "prompt": "Classify this support request: I forgot my password."
}

The ASP.NET Core API forwards the request to the inference service and returns the generated response.

22. Add Timeout and Resilience

LLM inference can take considerably longer than a conventional REST request.

Configure an appropriate HTTP timeout:

builder.Services.AddHttpClient<ILlmClient, LlmClient>(
    client =>
    {
        client.BaseAddress =
            new Uri("http://localhost:8000");

        client.Timeout =
            TimeSpan.FromSeconds(60);
    });

Production systems should also consider:

Blindly retrying every failed generation request can increase load on an already overloaded inference server, so retry policies should distinguish transient failures from model or validation errors.

23. Secure the Model Endpoint

The model inference service should not necessarily be exposed publicly.

A common architecture is:

Internet
   |
   v
ASP.NET Core
   |
Private Network
   |
   v
LLM Service

The ASP.NET Core application becomes the controlled entry point.

Security controls can include:

Never embed API keys or credentials directly inside source code.

24. Containerize the Services

A production architecture can use separate containers:

+--------------------------+
| ASP.NET Core Container   |
|                          |
| REST API                 |
+------------+-------------+
             |
             | HTTP
             v
+--------------------------+
| LLM Container            |
|                          |
| Python                   |
| Transformers             |
| PEFT                     |
| Fine-Tuned Model         |
+--------------------------+

A simplified ASP.NET Core Dockerfile:

FROM mcr.microsoft.com/dotnet/aspnet:8.0 AS base

WORKDIR /app
EXPOSE 8080

FROM mcr.microsoft.com/dotnet/sdk:8.0 AS build

WORKDIR /src

COPY . .

RUN dotnet restore

RUN dotnet publish \
    -c Release \
    -o /app/publish

FROM base AS final

WORKDIR /app

COPY --from=build /app/publish .

ENTRYPOINT ["dotnet", "LlmIntegrationApi.dll"]

The inference container would use a Python-compatible base image and include the model runtime and dependencies.

For GPU inference, the container and host environment must also be configured for the appropriate GPU runtime.

25. Managing Model Versions

Fine-tuned models should be treated like software releases.

Instead of:

model/

use versioned artifacts:

models/
├── support-classifier-v1/
├── support-classifier-v2/
└── support-classifier-v3/

Each version should have associated metadata:

{
  "model": "support-classifier",
  "version": "3",
  "baseModel": "open-source-model",
  "datasetVersion": "2026-09",
  "trainingEpochs": 3,
  "evaluationScore": 0.94
}

This makes rollback and reproducibility significantly easier.

26. Monitor Production Inference

Deploying the model is not the end of the process.

Monitor:

Application Metrics

LLM Metrics

Quality Metrics

A model can have excellent validation metrics but still perform poorly on real-world requests.

Production feedback should therefore become part of the model improvement process.

27. Fine-Tuning Does Not Automatically Eliminate Hallucinations

A common misconception is that fine-tuning makes an LLM factually reliable.

It does not.

Fine-tuning primarily teaches patterns and behavior represented in the training data.

For factual enterprise information, a better architecture may be:

User Query
     |
     v
ASP.NET Core
     |
     v
Retriever
     |
     v
Enterprise Knowledge
     |
     v
Fine-Tuned LLM
     |
     v
Validated Response

The fine-tuned model can provide consistent behavior while RAG supplies current information.

28. When Should You Fine-Tune an LLM?

Fine-tuning is worth considering when you repeatedly need the model to perform a particular task.

Good candidates include:

Fine-tuning is less suitable when the primary requirement is simply giving the model access to frequently changing documents.

In that scenario, RAG is often a better first approach.

29. Common Fine-Tuning Mistakes

Training on Too Little Data

A handful of examples rarely teaches robust behavior.

Quality and diversity matter.

Using Low-Quality Examples

The model learns patterns present in the dataset, including bad patterns.

Data Leakage

Do not allow evaluation examples to accidentally enter the training dataset.

Overfitting

Excessive training can make the model perform well on training examples while generalizing poorly.

Ignoring the Base Model License

Open-source does not automatically mean unrestricted commercial use.

Review the model's license before deployment.

Ignoring Inference Requirements

A model that can be trained successfully may still be expensive or slow to serve.

Evaluating Only With Loss

Training loss does not necessarily represent business-task performance.

30. Recommended Architecture for Enterprise .NET Applications

A practical architecture can look like this:

                    ┌──────────────────────┐
                    │      Client App      │
                    └──────────┬───────────┘
                               │
                               ▼
                    ┌──────────────────────┐
                    │    ASP.NET Core API  │
                    │                      │
                    │ Authentication       │
                    │ Authorization        │
                    │ Validation            │
                    │ Rate Limiting        │
                    └──────────┬───────────┘
                               │
                     ┌─────────┴─────────┐
                     │                   │
                     ▼                   ▼
             ┌───────────────┐   ┌───────────────┐
             │ RAG / Search  │   │ LLM Service   │
             │               │   │               │
             │ Vector DB     │   │ Fine-Tuned    │
             │ Documents     │   │ Open-Source   │
             └───────┬───────┘   │ Model         │
                     │           └───────┬───────┘
                     └─────────┬─────────┘
                               ▼
                    ┌──────────────────────┐
                    │ Response Validation  │
                    └──────────┬───────────┘
                               │
                               ▼
                         Client Response

This architecture separates application logic from model infrastructure.

The ASP.NET Core layer handles business and security concerns, while the model service handles inference.

31. Key Takeaways

Fine-tuning an open-source LLM does not require building a language model from scratch.

A practical implementation can use:

Open-Source LLM
       +
Domain-Specific Dataset
       +
LoRA / PEFT
       +
Hugging Face
       +
Python Inference API
       +
ASP.NET Core

The most important engineering decisions are not simply the number of training epochs or GPU size.

They include:

  1. Selecting an appropriate base model.
  2. Creating high-quality training data.
  3. Choosing between full fine-tuning and parameter-efficient methods.
  4. Evaluating the model against unseen examples.
  5. Separating model inference from application logic.
  6. Securing the inference endpoint.
  7. Monitoring latency and model quality.
  8. Versioning datasets and model artifacts.
  9. Combining fine-tuning with RAG when both behavior and external knowledge are required.

For .NET teams, this architecture makes it possible to keep the core business application in ASP.NET Core and C# while using the mature Python ecosystem for model training and inference.

The result is a flexible architecture where the fine-tuned model can evolve independently from the application layer while remaining accessible through a conventional HTTP API.