Large Language Models (LLMs) can be used directly through APIs, but generic models often struggle with highly specialized terminology, response formats, internal business processes, and domain-specific instructions.
Fine-tuning an open-source LLM provides another approach. Instead of changing the model's knowledge base entirely, fine-tuning adjusts the model's parameters so it learns a specific behavior, style, task, or domain pattern from a curated training dataset.
For .NET developers, a practical architecture is to use the Python ecosystem for model training and expose the resulting model through an inference API that can be consumed by an ASP.NET Core application.
This article demonstrates that architecture using:
- Python
- PyTorch
- Hugging Face Transformers
- PEFT
- LoRA
- An open-source causal language model
- ASP.NET Core Web API
- HTTP-based model inference
- Docker-ready deployment concepts
The goal is not to train an LLM from scratch. Instead, we will fine-tune an existing open-source model using parameter-efficient fine-tuning and integrate it with an ASP.NET Core application.
1. What Is LLM Fine-Tuning?
Fine-tuning takes a pretrained language model and trains it further using a task-specific dataset.
A pretrained model may already understand general language patterns:
Customer: How do I reset my password?
Model:
You can reset your password using the account settings...
However, an enterprise application may require a specific response format:
Intent: PASSWORD_RESET
Priority: LOW
Action: SEND_RESET_LINK
A fine-tuned model can learn this type of behavior from examples.
Conceptually:
Pretrained LLM
|
v
Domain-Specific Dataset
|
v
Fine-Tuning
|
v
Specialized Model
|
v
Inference API
|
v
ASP.NET Core Application
Fine-tuning is particularly useful when the application needs consistent behavior rather than simply retrieving additional information.
2. Fine-Tuning vs. RAG
Fine-tuning and Retrieval-Augmented Generation solve different problems.
| Requirement | Fine-Tuning | RAG |
|---|---|---|
| Learn response style | Strong fit | Limited |
| Learn structured output | Strong fit | Possible |
| Add frequently changing information | Poor fit | Strong fit |
| Company documentation | Sometimes | Strong fit |
| Domain-specific behavior | Strong fit | Partial |
| Private knowledge retrieval | Not ideal | Strong fit |
| Reduce repeated prompting | Strong fit | Partial |
| Real-time information | No | Yes |
For example, suppose an insurance application needs to answer questions using documents that change every week.
RAG is generally more appropriate because the documents can remain outside the model and be retrieved at runtime.
If the application needs the model to consistently classify insurance claims into predefined categories, fine-tuning may be appropriate.
In many enterprise systems, the architecture is actually:
User Request
|
v
ASP.NET Core
|
+----> RAG Retrieval
|
+----> Fine-Tuned LLM
|
v
Final Response
The two techniques are not mutually exclusive.
3. Choosing an Open-Source Model
Before fine-tuning, select a model that matches the application's requirements.
Important factors include:
- Model architecture
- Parameter count
- Context length
- License
- Hardware requirements
- Language support
- Existing benchmark performance
- Quantization support
- Community ecosystem
- Inference latency
For experimentation, smaller instruction-tuned models are usually easier to work with than very large models.
The model should also support the training and inference libraries you intend to use.
For example:
Model
|
+-- Transformers compatibility
+-- PEFT compatibility
+-- Tokenizer availability
+-- Quantization support
+-- License suitable for your application
Do not select a model solely because it has a large parameter count.
A smaller model that is properly fine-tuned can be more practical for a specific production task.
4. Preparing the Training Dataset
The dataset is often more important than the training configuration.
For an instruction-following model, a dataset can contain examples such as:
{
"instruction": "Classify the following support request.",
"input": "I cannot access my account because I forgot my password.",
"output": "PASSWORD_RESET"
}
Another example:
{
"instruction": "Classify the following support request.",
"input": "The application crashes whenever I upload a PDF.",
"output": "APPLICATION_ERROR"
}
A larger dataset might contain thousands of carefully reviewed examples.
The dataset should be:
- Consistent
- Representative
- Free from unnecessary duplication
- Properly formatted
- Free from sensitive information unless appropriately handled
- Balanced across important categories
Poor-quality examples can teach the model incorrect behavior.
5. Instruction Dataset Format
A common approach is to convert each example into a conversational or instruction format.
For example:
### Instruction:
Classify the support request.
### Input:
The application crashes whenever I upload a PDF.
### Response:
APPLICATION_ERROR
During preprocessing, these examples are tokenized and converted into model-compatible tensors.
A JSONL file is convenient because each line represents one training example:
{"instruction":"Classify the request.","input":"I forgot my password.","output":"PASSWORD_RESET"}
{"instruction":"Classify the request.","input":"The application crashes.","output":"APPLICATION_ERROR"}
{"instruction":"Classify the request.","input":"I need to change my billing address.","output":"ACCOUNT_UPDATE"}
6. Install the Python Training Environment
The training portion can be isolated from the .NET application.
Create a Python environment:
python -m venv .venv
Activate it on Windows:
.venv\Scripts\activate
On Linux:
source .venv/bin/activate
Install the required libraries:
pip install torch
pip install transformers
pip install datasets
pip install peft
pip install accelerate
Depending on the model and hardware configuration, additional packages may be required for quantization.
7. Load the Dataset
Using Hugging Face Datasets:
from datasets import load_dataset
dataset = load_dataset(
"json",
data_files="training-data.jsonl"
)
print(dataset)
A dataset might be divided into training and validation sets:
dataset = dataset["train"].train_test_split(
test_size=0.1,
seed=42
)
train_dataset = dataset["train"]
validation_dataset = dataset["test"]
Keeping validation data separate is important.
The model should not be evaluated only against examples it has already seen during training.
8. Load the Tokenizer
The tokenizer converts text into token IDs understood by the model.
from transformers import AutoTokenizer
model_name = "YOUR_MODEL_NAME"
tokenizer = AutoTokenizer.from_pretrained(model_name)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
The tokenizer must correspond to the selected model.
Using an incompatible tokenizer can produce incorrect input representations.
9. Format the Dataset
Create a function that converts each record into the format used for training.
def format_example(example):
text = f"""
### Instruction:
{example["instruction"]}
### Input:
{example["input"]}
### Response:
{example["output"]}
"""
return {"text": text.strip()}
Apply it to the dataset:
train_dataset = train_dataset.map(format_example)
validation_dataset = validation_dataset.map(format_example)
Tokenization can then be performed:
def tokenize(example):
return tokenizer(
example["text"],
truncation=True,
max_length=1024
)
tokenized_train = train_dataset.map(
tokenize,
batched=True
)
tokenized_validation = validation_dataset.map(
tokenize,
batched=True
)
The max_length value should be selected according to the model and dataset.
10. Why LoRA Is Useful
Full fine-tuning modifies a large portion of the model's parameters.
That can require significant GPU memory.
LoRA, or Low-Rank Adaptation, takes a different approach.
Instead of updating all model parameters, LoRA introduces smaller trainable matrices into selected model layers.
Conceptually:
Original Model Parameters
|
| frozen
v
----------------
| |
| LLM |
| |
----------------
|
LoRA Layers
|
v
Trainable Parameters
This significantly reduces the number of parameters that need to be trained.
The resulting adapter can then be stored separately from the original model.
11. Configure LoRA
Using PEFT:
from peft import LoraConfig, TaskType
lora_config = LoraConfig(
task_type=TaskType.CAUSAL_LM,
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=[
"q_proj",
"v_proj"
]
)
Important parameters include:
r
Controls the rank of the LoRA matrices.
Higher values generally provide more trainable capacity but increase memory and computation requirements.
lora_alpha
Controls the scaling of the LoRA contribution.
lora_dropout
Adds regularization during training.
target_modules
Defines which model layers receive LoRA adapters.
The correct target module names depend on the selected model architecture.
12. Load the Model
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto"
)
For larger models, quantization can be combined with parameter-efficient fine-tuning to reduce GPU memory requirements.
For example, a 4-bit training configuration can be used when supported by the model and environment.
The exact configuration should be tested against the available GPU because hardware compatibility differs between environments.
13. Configure Training
Using the Transformers training stack:
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="./fine-tuned-model",
num_train_epochs=3,
per_device_train_batch_size=2,
gradient_accumulation_steps=8,
learning_rate=2e-4,
logging_steps=10,
save_steps=100,
evaluation_strategy="steps",
eval_steps=100,
fp16=True,
report_to="none"
)
These values are examples, not universal recommendations.
Training parameters should be tuned based on:
- Dataset size
- GPU memory
- Model architecture
- Sequence length
- Desired output quality
- Training loss
- Validation performance
14. Train the Model
The model can then be wrapped with the LoRA configuration and trained.
from peft import get_peft_model
model = get_peft_model(
model,
lora_config
)
model.print_trainable_parameters()
A training loop can then be created using the Hugging Face training infrastructure.
For example:
from transformers import Trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized_train,
eval_dataset=tokenized_validation
)
trainer.train()
After training:
trainer.save_model("./fine-tuned-model")
tokenizer.save_pretrained("./fine-tuned-model")
The resulting directory contains the trained adapter and supporting model configuration.
15. Evaluate the Fine-Tuned Model
Training loss alone is not sufficient to determine whether the model is useful.
Create a separate evaluation dataset.
For classification:
Input:
My password reset link has expired.
Expected:
PASSWORD_RESET
For structured generation:
{
"category": "PASSWORD_RESET",
"priority": "LOW"
}
Evaluate metrics appropriate to the task.
For classification:
- Accuracy
- Precision
- Recall
- F1 score
For text generation:
- Exact match
- ROUGE
- BLEU where appropriate
- Human evaluation
- Structured-output validity
- Task-specific metrics
For enterprise systems, task-specific evaluation is usually more meaningful than a generic language benchmark.
16. Build an Inference API
The easiest architecture for ASP.NET Core integration is to expose the model through an HTTP API.
For example:
ASP.NET Core
|
| HTTP POST
v
Python Inference API
|
v
Fine-Tuned LLM
The inference service can be implemented using FastAPI.
Install it:
pip install fastapi uvicorn
Create:
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI()
class PredictionRequest(BaseModel):
prompt: str
The inference implementation loads the model and generates a response.
A simplified endpoint might look like:
@app.post("/generate")
def generate(request: PredictionRequest):
inputs = tokenizer(
request.prompt,
return_tensors="pt"
)
outputs = model.generate(
**inputs,
max_new_tokens=256
)
result = tokenizer.decode(
outputs[0],
skip_special_tokens=True
)
return {
"response": result
}
For production workloads, inference code should also handle:
- Authentication
- Request validation
- Timeouts
- Concurrency
- GPU utilization
- Token limits
- Logging
- Monitoring
- Error handling
17. Create the ASP.NET Core Application
Create an ASP.NET Core Web API:
dotnet new webapi -n LlmIntegrationApi
cd LlmIntegrationApi
The application will act as the client of the inference service.
The architecture becomes:
Client
|
v
ASP.NET Core Web API
|
| HTTP
v
LLM Inference Service
|
v
Fine-Tuned Open-Source Model
This separation allows the model runtime to remain independent of the .NET application.
18. Create Request and Response Models
Create a request DTO:
public sealed class LlmRequest
{
public string Prompt { get; set; } = string.Empty;
}
Create a response DTO:
public sealed class LlmResponse
{
public string Response { get; set; } = string.Empty;
}
19. Create an LLM Client Service
Instead of placing HTTP logic directly inside a controller, create a dedicated service.
public interface ILlmClient
{
Task<string> GenerateAsync(
string prompt,
CancellationToken cancellationToken = default);
}
Implementation:
using System.Net.Http.Json;
public sealed class LlmClient : ILlmClient
{
private readonly HttpClient _httpClient;
public LlmClient(HttpClient httpClient)
{
_httpClient = httpClient;
}
public async Task<string> GenerateAsync(
string prompt,
CancellationToken cancellationToken = default)
{
var request = new LlmRequest
{
Prompt = prompt
};
var response = await _httpClient.PostAsJsonAsync(
"/generate",
request,
cancellationToken);
response.EnsureSuccessStatusCode();
var result =
await response.Content.ReadFromJsonAsync<LlmResponse>(
cancellationToken: cancellationToken);
return result?.Response ?? string.Empty;
}
}
This keeps the application architecture clean and makes the model provider replaceable.
20. Register the HTTP Client
In Program.cs:
builder.Services.AddHttpClient<ILlmClient, LlmClient>(
client =>
{
client.BaseAddress =
new Uri("http://localhost:8000");
});
In production, the endpoint should come from configuration rather than being hard-coded.
For example:
{
"LlmService": {
"BaseUrl": "http://llm-service:8000"
}
}
Then:
builder.Services.AddHttpClient<ILlmClient, LlmClient>(
(serviceProvider, client) =>
{
var configuration =
serviceProvider.GetRequiredService<IConfiguration>();
var baseUrl =
configuration["LlmService:BaseUrl"];
client.BaseAddress =
new Uri(baseUrl!);
});
21. Create the ASP.NET Core Controller
using Microsoft.AspNetCore.Mvc;
[ApiController]
[Route("api/[controller]")]
public class AiController : ControllerBase
{
private readonly ILlmClient _llmClient;
public AiController(ILlmClient llmClient)
{
_llmClient = llmClient;
}
[HttpPost("generate")]
public async Task<IActionResult> Generate(
[FromBody] LlmRequest request,
CancellationToken cancellationToken)
{
if (string.IsNullOrWhiteSpace(request.Prompt))
{
return BadRequest("Prompt is required.");
}
var response =
await _llmClient.GenerateAsync(
request.Prompt,
cancellationToken);
return Ok(new
{
response
});
}
}
The client can now send:
POST /api/ai/generate
Content-Type: application/json
with:
{
"prompt": "Classify this support request: I forgot my password."
}
The ASP.NET Core API forwards the request to the inference service and returns the generated response.
22. Add Timeout and Resilience
LLM inference can take considerably longer than a conventional REST request.
Configure an appropriate HTTP timeout:
builder.Services.AddHttpClient<ILlmClient, LlmClient>(
client =>
{
client.BaseAddress =
new Uri("http://localhost:8000");
client.Timeout =
TimeSpan.FromSeconds(60);
});
Production systems should also consider:
- Retry policies
- Circuit breakers
- Cancellation tokens
- Request queues
- Rate limiting
- Health checks
Blindly retrying every failed generation request can increase load on an already overloaded inference server, so retry policies should distinguish transient failures from model or validation errors.
23. Secure the Model Endpoint
The model inference service should not necessarily be exposed publicly.
A common architecture is:
Internet
|
v
ASP.NET Core
|
Private Network
|
v
LLM Service
The ASP.NET Core application becomes the controlled entry point.
Security controls can include:
- Authentication
- Authorization
- API keys
- Network isolation
- TLS
- Request validation
- Rate limiting
- Secrets management
- Audit logging
Never embed API keys or credentials directly inside source code.
24. Containerize the Services
A production architecture can use separate containers:
+--------------------------+
| ASP.NET Core Container |
| |
| REST API |
+------------+-------------+
|
| HTTP
v
+--------------------------+
| LLM Container |
| |
| Python |
| Transformers |
| PEFT |
| Fine-Tuned Model |
+--------------------------+
A simplified ASP.NET Core Dockerfile:
FROM mcr.microsoft.com/dotnet/aspnet:8.0 AS base
WORKDIR /app
EXPOSE 8080
FROM mcr.microsoft.com/dotnet/sdk:8.0 AS build
WORKDIR /src
COPY . .
RUN dotnet restore
RUN dotnet publish \
-c Release \
-o /app/publish
FROM base AS final
WORKDIR /app
COPY --from=build /app/publish .
ENTRYPOINT ["dotnet", "LlmIntegrationApi.dll"]
The inference container would use a Python-compatible base image and include the model runtime and dependencies.
For GPU inference, the container and host environment must also be configured for the appropriate GPU runtime.
25. Managing Model Versions
Fine-tuned models should be treated like software releases.
Instead of:
model/
use versioned artifacts:
models/
├── support-classifier-v1/
├── support-classifier-v2/
└── support-classifier-v3/
Each version should have associated metadata:
{
"model": "support-classifier",
"version": "3",
"baseModel": "open-source-model",
"datasetVersion": "2026-09",
"trainingEpochs": 3,
"evaluationScore": 0.94
}
This makes rollback and reproducibility significantly easier.
26. Monitor Production Inference
Deploying the model is not the end of the process.
Monitor:
Application Metrics
- Request count
- Error rate
- Response time
- HTTP status codes
LLM Metrics
- Input tokens
- Output tokens
- Generation latency
- GPU utilization
- Memory usage
- Context length
Quality Metrics
- Classification accuracy
- Invalid responses
- User feedback
- Hallucination rate
- Task completion rate
A model can have excellent validation metrics but still perform poorly on real-world requests.
Production feedback should therefore become part of the model improvement process.
27. Fine-Tuning Does Not Automatically Eliminate Hallucinations
A common misconception is that fine-tuning makes an LLM factually reliable.
It does not.
Fine-tuning primarily teaches patterns and behavior represented in the training data.
For factual enterprise information, a better architecture may be:
User Query
|
v
ASP.NET Core
|
v
Retriever
|
v
Enterprise Knowledge
|
v
Fine-Tuned LLM
|
v
Validated Response
The fine-tuned model can provide consistent behavior while RAG supplies current information.
28. When Should You Fine-Tune an LLM?
Fine-tuning is worth considering when you repeatedly need the model to perform a particular task.
Good candidates include:
- Classification
- Intent detection
- Structured output
- Domain-specific language
- Consistent response style
- Specialized extraction
- Code generation patterns
- Industry-specific workflows
Fine-tuning is less suitable when the primary requirement is simply giving the model access to frequently changing documents.
In that scenario, RAG is often a better first approach.
29. Common Fine-Tuning Mistakes
Training on Too Little Data
A handful of examples rarely teaches robust behavior.
Quality and diversity matter.
Using Low-Quality Examples
The model learns patterns present in the dataset, including bad patterns.
Data Leakage
Do not allow evaluation examples to accidentally enter the training dataset.
Overfitting
Excessive training can make the model perform well on training examples while generalizing poorly.
Ignoring the Base Model License
Open-source does not automatically mean unrestricted commercial use.
Review the model's license before deployment.
Ignoring Inference Requirements
A model that can be trained successfully may still be expensive or slow to serve.
Evaluating Only With Loss
Training loss does not necessarily represent business-task performance.
30. Recommended Architecture for Enterprise .NET Applications
A practical architecture can look like this:
┌──────────────────────┐
│ Client App │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ ASP.NET Core API │
│ │
│ Authentication │
│ Authorization │
│ Validation │
│ Rate Limiting │
└──────────┬───────────┘
│
┌─────────┴─────────┐
│ │
▼ ▼
┌───────────────┐ ┌───────────────┐
│ RAG / Search │ │ LLM Service │
│ │ │ │
│ Vector DB │ │ Fine-Tuned │
│ Documents │ │ Open-Source │
└───────┬───────┘ │ Model │
│ └───────┬───────┘
└─────────┬─────────┘
▼
┌──────────────────────┐
│ Response Validation │
└──────────┬───────────┘
│
▼
Client Response
This architecture separates application logic from model infrastructure.
The ASP.NET Core layer handles business and security concerns, while the model service handles inference.
31. Key Takeaways
Fine-tuning an open-source LLM does not require building a language model from scratch.
A practical implementation can use:
Open-Source LLM
+
Domain-Specific Dataset
+
LoRA / PEFT
+
Hugging Face
+
Python Inference API
+
ASP.NET Core
The most important engineering decisions are not simply the number of training epochs or GPU size.
They include:
- Selecting an appropriate base model.
- Creating high-quality training data.
- Choosing between full fine-tuning and parameter-efficient methods.
- Evaluating the model against unseen examples.
- Separating model inference from application logic.
- Securing the inference endpoint.
- Monitoring latency and model quality.
- Versioning datasets and model artifacts.
- Combining fine-tuning with RAG when both behavior and external knowledge are required.
For .NET teams, this architecture makes it possible to keep the core business application in ASP.NET Core and C# while using the mature Python ecosystem for model training and inference.
The result is a flexible architecture where the fine-tuned model can evolve independently from the application layer while remaining accessible through a conventional HTTP API.

Join the conversation! Your thoughts help the community grow.