AI features can be easy to add and surprisingly difficult to debug.
A normal API call is fairly simple to follow. A request enters the application, a function runs, and a response comes back. With an AI feature, one user request can involve several model calls, retrieval steps, prompts, tool calls, retries, and evaluation logic.
Then someone asks:
Why did this request take 8 seconds and cost more than expected?
Looking only at application logs usually does not give a good answer.
This is where Sentry can help. With the Python SDK, AI-related work can be attached to the application's tracing and error reporting instead of being treated as a completely separate system.
The useful part is not simply recording that an LLM was called. It is being able to connect that call to the request that caused it.
A Small Example
Imagine a Python API that creates product descriptions with an LLM.
The request might look like this:
@app.post("/generate")
def generate_product_description(product):
prompt = build_prompt(product)
response = client.responses.create(
model="gpt-model",
input=prompt
)
return {"description": response.output_text}
At first, this looks fine.
But in a real application, you may eventually need to answer questions such as:
Which requests are making the most model calls?
Which model calls are slow?
How many tokens are being used?
Are retries increasing the cost?
Which prompts produce poor results?
Did a failed request come from the model, the database, or our own code?
Is one endpoint responsible for most of the AI usage?
Those questions require more than a print() statement around the API call.
What Sentry Adds
Sentry is already commonly used for application errors and performance monitoring.
For an AI-enabled Python application, the same monitoring approach can be extended to the part of the application that talks to a model.
A simplified request might look like:
HTTP Request
|
v
Python API
|
+---- Database
|
+---- Retrieval
|
+---- LLM Call
|
+---- Tokens
+---- Latency
+---- Result
Instead of looking at each operation separately, tracing lets you inspect them as part of the same request.
That makes troubleshooting much easier.
Start With the Python SDK
The first step is adding the Sentry Python SDK to the application.
For example:
pip install sentry-sdk
Then initialize it early in the application:
import sentry_sdk
sentry_sdk.init(
dsn="YOUR_SENTRY_DSN",
traces_sample_rate=0.1
)
The exact sampling configuration should depend on your application's traffic and monitoring requirements.
For a development environment, you may choose a higher sampling rate while testing. A busy production service may need a lower rate to control the amount of tracing data collected.
The important thing is to initialize Sentry before the work you want to observe begins.
Trace the Request Before Looking at Tokens
Token usage is useful, but it does not tell the whole story.
Suppose a request uses 4,000 input tokens.
Is that good or bad?
It depends.
Maybe the request also spent three seconds retrieving documents and another four seconds waiting for the model.
The better view is:
Request
|
+-- Database: 120 ms
|
+-- Retrieval: 1.8 sec
|
+-- Model call: 4.2 sec
|
+-- Response processing: 150 ms
Now there is something useful to investigate.
If model calls are responsible for most of the time, look at the model integration.
If retrieval is slow, changing the model will not solve the main problem.
Add AI Work to a Trace
For custom AI code, Sentry's tracing APIs can be used to create spans around important operations.
A simplified example:
import sentry_sdk
def generate_description(product):
with sentry_sdk.start_span(
op="ai.generate",
name="Generate product description"
) as span:
span.set_data("model", "gpt-model")
span.set_data("product_id", product["id"])
response = client.responses.create(
model="gpt-model",
input=build_prompt(product)
)
return response.output_text
Now the model operation has a place in the application's trace.
The exact span names and instrumentation approach can vary depending on the libraries being used, but the idea is simple: surround meaningful AI operations with trace information.
Do not create a span for every tiny line of Python.
A span should represent something you may actually want to investigate.
Recording Token Usage
Token counts become particularly useful when they are attached to the model operation.
For example, if the provider response gives you usage information, you can record it:
usage = response.usage
span.set_data(
"input_tokens",
usage.input_tokens
)
span.set_data(
"output_tokens",
usage.output_tokens
)
You might also record the total:
span.set_data(
"total_tokens",
usage.input_tokens + usage.output_tokens
)
The field names depend on the model provider and response format.
The important thing is to capture the values consistently.
Once the data is attached to the trace, you can start looking for patterns.
For example:
Request A -> 800 tokens
Request B -> 1,100 tokens
Request C -> 7,900 tokens
Request D -> 950 tokens
Request C deserves a closer look.
Maybe the prompt is too large.
Maybe retrieval is returning too many documents.
Maybe the application is sending conversation history that no longer needs to be included.
Do Not Store the Whole Prompt by Default
This is one area where AI observability needs some care.
It can be tempting to store everything:
span.set_data("prompt", prompt)
span.set_data("response", response.output_text)
That may make debugging easier, but prompts and responses can contain sensitive information.
They may include:
customer information
internal documents
access details
source code
business data
personal information
Before capturing prompt or response content, decide whether the data is actually safe to store.
A better approach is often to record metadata:
span.set_data("model", model_name)
span.set_data("input_tokens", input_tokens)
span.set_data("output_tokens", output_tokens)
span.set_data("prompt_version", "product-description-v3")
Now you can compare requests without putting the entire conversation into your monitoring system.
Prompt Versions Are Surprisingly Useful
If your team changes prompts frequently, record a prompt version.
For example:
span.set_data("prompt_version", "support-agent-v5")
Suppose the application starts producing more expensive requests after a deployment.
Without a prompt version, you may have to guess what changed.
With it, you can compare:
v4 -> average 1,200 input tokens
v5 -> average 2,900 input tokens
That gives the development team a much clearer starting point.
Prompt changes should be treated like code changes when they affect application behavior.
AI Evaluation Can Be Traced Too
Monitoring model calls is only half the problem.
Many AI applications also evaluate the generated result.
For example:
answer = generate_answer(question)
score = evaluate_answer(
question,
answer
)
The evaluation itself may involve another model call.
Now one user request contains two different AI operations:
User Request
|
+---- Generate Answer
|
+---- Evaluate Answer
It is useful to keep those operations separate.
For example:
with sentry_sdk.start_span(
op="ai.generate",
name="Generate answer"
) as generation_span:
answer = generate_answer(question)
with sentry_sdk.start_span(
op="ai.evaluate",
name="Evaluate answer"
) as evaluation_span:
score = evaluate_answer(question, answer)
Now a slow evaluation does not look like a slow generation request.
That distinction matters when the evaluation process itself uses a model.
Measuring Evaluation Results
Suppose your application evaluates answers using a score from 0 to 1.
You can attach the result:
evaluation_span.set_data(
"evaluation_score",
score
)
You could also record the evaluation type:
evaluation_span.set_data(
"evaluation_type",
"answer_relevance"
)
Over time, this gives you a way to compare application behavior with model usage.
For example, you might discover that a particular prompt version uses more tokens without improving evaluation results.
That is a much more useful finding than simply knowing that the application made more API calls.
Tracing Retrieval-Augmented Applications
RAG applications add another layer.
A typical request may look like:
Question
|
v
Create Embedding
|
v
Vector Search
|
v
Select Documents
|
v
Build Prompt
|
v
LLM
|
v
Answer
If the final answer is poor, the model may not be the problem.
The vector search could have returned irrelevant documents.
For that reason, tracing should cover important retrieval operations as well.
For example:
with sentry_sdk.start_span(
op="ai.retrieval",
name="Search product documentation"
) as span:
documents = search_documents(question)
span.set_data(
"documents_found",
len(documents)
)
You do not necessarily need to store the document contents.
Knowing how many documents were retrieved, how long the search took, and which retrieval strategy was used can already be valuable.
Finding Expensive Requests
Once token information is attached to traces, you can investigate outliers.
Imagine these requests:
Request | Input Tokens | Output Tokens | Time |
|---|---|---|---|
A | 720 | 180 | 1.2s |
B | 850 | 210 | 1.4s |
C | 6,900 | 240 | 3.9s |
D | 780 | 190 | 1.3s |
Request C stands out immediately.
The output is similar to the other requests, but the input is much larger.
That often points toward things such as:
oversized conversation history
too many retrieved documents
duplicated context
unnecessarily large system instructions
a prompt construction bug
Observability turns that into an investigation rather than a guess.
Finding Slow AI Requests
Latency should also be looked at separately from token usage.
A request with 2,000 tokens might still be slow because of:
provider latency
network problems
retries
multiple sequential model calls
slow retrieval
application-side processing
For example:
Total request: 7.5 sec
Retrieval: 0.8 sec
Model call 1: 2.1 sec
Model call 2: 3.7 sec
Processing: 0.9 sec
The obvious optimization may be to investigate why two model calls are being made sequentially.
Without tracing, the application may simply appear “slow.”
Add Error Context
Errors become much easier to diagnose when they are connected to the AI operation that caused them.
For example:
try:
response = client.responses.create(
model=model_name,
input=prompt
)
except Exception as exc:
sentry_sdk.capture_exception(exc)
raise
The surrounding trace can provide additional context.
Instead of seeing:
TimeoutError
you may be able to determine:
Endpoint: /generate
Operation: ai.generate
Model: gpt-model
Prompt version: product-description-v3
Input tokens: 4,200
That is considerably more useful when investigating an incident.
Sampling Matters
You do not always need to trace every request.
If an application receives thousands of requests per minute, collecting detailed traces for everything can create unnecessary monitoring volume.
Sampling lets you control how much tracing data is collected.
For example:
sentry_sdk.init(
dsn="YOUR_SENTRY_DSN",
traces_sample_rate=0.1
)
This is only an example value.
The right setting depends on traffic, debugging needs, and your monitoring configuration.
During a short debugging exercise, you may temporarily collect more data. For normal operation, you may choose a lower rate.
Keep Observability Data Useful
It is easy to collect too much information.
For an AI application, I would start with a small set of fields:
model
prompt_version
input_tokens
output_tokens
latency
evaluation_score
retrieval_count
operation_name
Then add more when there is a real debugging need.
This keeps traces easier to understand and reduces the chance of accidentally storing sensitive content.
A Simple Pattern for an AI Service
A small service can bring these ideas together:
import sentry_sdk
def generate_answer(question: str):
with sentry_sdk.start_span(
op="ai.generate",
name="Generate answer"
) as span:
span.set_data(
"prompt_version",
"support-v3"
)
response = client.responses.create(
model="gpt-model",
input=question
)
if response.usage:
span.set_data(
"input_tokens",
response.usage.input_tokens
)
span.set_data(
"output_tokens",
response.usage.output_tokens
)
return response.output_text
It is intentionally small.
You can build more detailed instrumentation later, but starting with a few useful fields is better than creating a complicated monitoring layer that nobody uses.
What I Would Monitor First
For a new AI feature, I would start with four things:
Latency
How long does each model operation take?
Token usage
How much input and output are we sending?
Failures
Which requests are failing and where?
Evaluation
Are changes to prompts or models actually improving the result?
Once those four areas are visible, other questions become much easier to answer.
Conclusion
Adding AI to a Python application introduces another layer that needs monitoring. A normal application trace can tell you that an endpoint was slow or failed, but it may not tell you what happened inside the AI workflow.
Sentry's Python SDK can be used to connect those AI operations to the rest of the request.
Start small. Trace the model calls, record token counts, keep prompt versions, add evaluation results where useful, and be careful about putting prompt or response content into monitoring data.
For RAG applications, include retrieval in the trace as well. For AI evaluation, keep generation and evaluation as separate operations.
The goal is not to collect every possible piece of information.
The goal is to be able to look at a failed, slow, or expensive request and answer a simple question:
What actually happened?

Join the conversation! Your thoughts help the community grow.