AI features can be easy to add and surprisingly difficult to debug.

A normal API call is fairly simple to follow. A request enters the application, a function runs, and a response comes back. With an AI feature, one user request can involve several model calls, retrieval steps, prompts, tool calls, retries, and evaluation logic.

Then someone asks:

Why did this request take 8 seconds and cost more than expected?

Looking only at application logs usually does not give a good answer.

This is where Sentry can help. With the Python SDK, AI-related work can be attached to the application's tracing and error reporting instead of being treated as a completely separate system.

The useful part is not simply recording that an LLM was called. It is being able to connect that call to the request that caused it.

A Small Example

Imagine a Python API that creates product descriptions with an LLM.

The request might look like this:

@app.post("/generate")
def generate_product_description(product):
    prompt = build_prompt(product)

    response = client.responses.create(
        model="gpt-model",
        input=prompt
    )

    return {"description": response.output_text}

At first, this looks fine.

But in a real application, you may eventually need to answer questions such as:

  • Which requests are making the most model calls?

  • Which model calls are slow?

  • How many tokens are being used?

  • Are retries increasing the cost?

  • Which prompts produce poor results?

  • Did a failed request come from the model, the database, or our own code?

  • Is one endpoint responsible for most of the AI usage?

Those questions require more than a print() statement around the API call.

What Sentry Adds

Sentry is already commonly used for application errors and performance monitoring.

For an AI-enabled Python application, the same monitoring approach can be extended to the part of the application that talks to a model.

A simplified request might look like:

HTTP Request
     |
     v
Python API
     |
     +---- Database
     |
     +---- Retrieval
     |
     +---- LLM Call
              |
              +---- Tokens
              +---- Latency
              +---- Result

Instead of looking at each operation separately, tracing lets you inspect them as part of the same request.

That makes troubleshooting much easier.

Start With the Python SDK

The first step is adding the Sentry Python SDK to the application.

For example:

pip install sentry-sdk

Then initialize it early in the application:

import sentry_sdk

sentry_sdk.init(
    dsn="YOUR_SENTRY_DSN",
    traces_sample_rate=0.1
)

The exact sampling configuration should depend on your application's traffic and monitoring requirements.

For a development environment, you may choose a higher sampling rate while testing. A busy production service may need a lower rate to control the amount of tracing data collected.

The important thing is to initialize Sentry before the work you want to observe begins.

Trace the Request Before Looking at Tokens

Token usage is useful, but it does not tell the whole story.

Suppose a request uses 4,000 input tokens.

Is that good or bad?

It depends.

Maybe the request also spent three seconds retrieving documents and another four seconds waiting for the model.

The better view is:

Request
  |
  +-- Database: 120 ms
  |
  +-- Retrieval: 1.8 sec
  |
  +-- Model call: 4.2 sec
  |
  +-- Response processing: 150 ms

Now there is something useful to investigate.

If model calls are responsible for most of the time, look at the model integration.

If retrieval is slow, changing the model will not solve the main problem.

Add AI Work to a Trace

For custom AI code, Sentry's tracing APIs can be used to create spans around important operations.

A simplified example:

import sentry_sdk

def generate_description(product):
    with sentry_sdk.start_span(
        op="ai.generate",
        name="Generate product description"
    ) as span:

        span.set_data("model", "gpt-model")
        span.set_data("product_id", product["id"])

        response = client.responses.create(
            model="gpt-model",
            input=build_prompt(product)
        )

        return response.output_text

Now the model operation has a place in the application's trace.

The exact span names and instrumentation approach can vary depending on the libraries being used, but the idea is simple: surround meaningful AI operations with trace information.

Do not create a span for every tiny line of Python.

A span should represent something you may actually want to investigate.

Recording Token Usage

Token counts become particularly useful when they are attached to the model operation.

For example, if the provider response gives you usage information, you can record it:

usage = response.usage

span.set_data(
    "input_tokens",
    usage.input_tokens
)

span.set_data(
    "output_tokens",
    usage.output_tokens
)

You might also record the total:

span.set_data(
    "total_tokens",
    usage.input_tokens + usage.output_tokens
)

The field names depend on the model provider and response format.

The important thing is to capture the values consistently.

Once the data is attached to the trace, you can start looking for patterns.

For example:

Request A -> 800 tokens
Request B -> 1,100 tokens
Request C -> 7,900 tokens
Request D -> 950 tokens

Request C deserves a closer look.

Maybe the prompt is too large.

Maybe retrieval is returning too many documents.

Maybe the application is sending conversation history that no longer needs to be included.

Do Not Store the Whole Prompt by Default

This is one area where AI observability needs some care.

It can be tempting to store everything:

span.set_data("prompt", prompt)
span.set_data("response", response.output_text)

That may make debugging easier, but prompts and responses can contain sensitive information.

They may include:

  • customer information

  • internal documents

  • access details

  • source code

  • business data

  • personal information

Before capturing prompt or response content, decide whether the data is actually safe to store.

A better approach is often to record metadata:

span.set_data("model", model_name)
span.set_data("input_tokens", input_tokens)
span.set_data("output_tokens", output_tokens)
span.set_data("prompt_version", "product-description-v3")

Now you can compare requests without putting the entire conversation into your monitoring system.

Prompt Versions Are Surprisingly Useful

If your team changes prompts frequently, record a prompt version.

For example:

span.set_data("prompt_version", "support-agent-v5")

Suppose the application starts producing more expensive requests after a deployment.

Without a prompt version, you may have to guess what changed.

With it, you can compare:

v4 -> average 1,200 input tokens
v5 -> average 2,900 input tokens

That gives the development team a much clearer starting point.

Prompt changes should be treated like code changes when they affect application behavior.

AI Evaluation Can Be Traced Too

Monitoring model calls is only half the problem.

Many AI applications also evaluate the generated result.

For example:

answer = generate_answer(question)

score = evaluate_answer(
    question,
    answer
)

The evaluation itself may involve another model call.

Now one user request contains two different AI operations:

User Request
    |
    +---- Generate Answer
    |
    +---- Evaluate Answer

It is useful to keep those operations separate.

For example:

with sentry_sdk.start_span(
    op="ai.generate",
    name="Generate answer"
) as generation_span:

    answer = generate_answer(question)


with sentry_sdk.start_span(
    op="ai.evaluate",
    name="Evaluate answer"
) as evaluation_span:

    score = evaluate_answer(question, answer)

Now a slow evaluation does not look like a slow generation request.

That distinction matters when the evaluation process itself uses a model.

Measuring Evaluation Results

Suppose your application evaluates answers using a score from 0 to 1.

You can attach the result:

evaluation_span.set_data(
    "evaluation_score",
    score
)

You could also record the evaluation type:

evaluation_span.set_data(
    "evaluation_type",
    "answer_relevance"
)

Over time, this gives you a way to compare application behavior with model usage.

For example, you might discover that a particular prompt version uses more tokens without improving evaluation results.

That is a much more useful finding than simply knowing that the application made more API calls.

Tracing Retrieval-Augmented Applications

RAG applications add another layer.

A typical request may look like:

Question
   |
   v
Create Embedding
   |
   v
Vector Search
   |
   v
Select Documents
   |
   v
Build Prompt
   |
   v
LLM
   |
   v
Answer

If the final answer is poor, the model may not be the problem.

The vector search could have returned irrelevant documents.

For that reason, tracing should cover important retrieval operations as well.

For example:

with sentry_sdk.start_span(
    op="ai.retrieval",
    name="Search product documentation"
) as span:

    documents = search_documents(question)

    span.set_data(
        "documents_found",
        len(documents)
    )

You do not necessarily need to store the document contents.

Knowing how many documents were retrieved, how long the search took, and which retrieval strategy was used can already be valuable.

Finding Expensive Requests

Once token information is attached to traces, you can investigate outliers.

Imagine these requests:

Request

Input Tokens

Output Tokens

Time

A

720

180

1.2s

B

850

210

1.4s

C

6,900

240

3.9s

D

780

190

1.3s

Request C stands out immediately.

The output is similar to the other requests, but the input is much larger.

That often points toward things such as:

  • oversized conversation history

  • too many retrieved documents

  • duplicated context

  • unnecessarily large system instructions

  • a prompt construction bug

Observability turns that into an investigation rather than a guess.

Finding Slow AI Requests

Latency should also be looked at separately from token usage.

A request with 2,000 tokens might still be slow because of:

  • provider latency

  • network problems

  • retries

  • multiple sequential model calls

  • slow retrieval

  • application-side processing

For example:

Total request: 7.5 sec

Retrieval:     0.8 sec
Model call 1:  2.1 sec
Model call 2:  3.7 sec
Processing:    0.9 sec

The obvious optimization may be to investigate why two model calls are being made sequentially.

Without tracing, the application may simply appear “slow.”

Add Error Context

Errors become much easier to diagnose when they are connected to the AI operation that caused them.

For example:

try:
    response = client.responses.create(
        model=model_name,
        input=prompt
    )

except Exception as exc:
    sentry_sdk.capture_exception(exc)
    raise

The surrounding trace can provide additional context.

Instead of seeing:

TimeoutError

you may be able to determine:

Endpoint: /generate
Operation: ai.generate
Model: gpt-model
Prompt version: product-description-v3
Input tokens: 4,200

That is considerably more useful when investigating an incident.

Sampling Matters

You do not always need to trace every request.

If an application receives thousands of requests per minute, collecting detailed traces for everything can create unnecessary monitoring volume.

Sampling lets you control how much tracing data is collected.

For example:

sentry_sdk.init(
    dsn="YOUR_SENTRY_DSN",
    traces_sample_rate=0.1
)

This is only an example value.

The right setting depends on traffic, debugging needs, and your monitoring configuration.

During a short debugging exercise, you may temporarily collect more data. For normal operation, you may choose a lower rate.

Keep Observability Data Useful

It is easy to collect too much information.

For an AI application, I would start with a small set of fields:

model
prompt_version
input_tokens
output_tokens
latency
evaluation_score
retrieval_count
operation_name

Then add more when there is a real debugging need.

This keeps traces easier to understand and reduces the chance of accidentally storing sensitive content.

A Simple Pattern for an AI Service

A small service can bring these ideas together:

import sentry_sdk


def generate_answer(question: str):

    with sentry_sdk.start_span(
        op="ai.generate",
        name="Generate answer"
    ) as span:

        span.set_data(
            "prompt_version",
            "support-v3"
        )

        response = client.responses.create(
            model="gpt-model",
            input=question
        )

        if response.usage:
            span.set_data(
                "input_tokens",
                response.usage.input_tokens
            )

            span.set_data(
                "output_tokens",
                response.usage.output_tokens
            )

        return response.output_text

It is intentionally small.

You can build more detailed instrumentation later, but starting with a few useful fields is better than creating a complicated monitoring layer that nobody uses.

What I Would Monitor First

For a new AI feature, I would start with four things:

Latency

How long does each model operation take?

Token usage

How much input and output are we sending?

Failures

Which requests are failing and where?

Evaluation

Are changes to prompts or models actually improving the result?

Once those four areas are visible, other questions become much easier to answer.

Conclusion

Adding AI to a Python application introduces another layer that needs monitoring. A normal application trace can tell you that an endpoint was slow or failed, but it may not tell you what happened inside the AI workflow.

Sentry's Python SDK can be used to connect those AI operations to the rest of the request.

Start small. Trace the model calls, record token counts, keep prompt versions, add evaluation results where useful, and be careful about putting prompt or response content into monitoring data.

For RAG applications, include retrieval in the trace as well. For AI evaluation, keep generation and evaluation as separate operations.

The goal is not to collect every possible piece of information.

The goal is to be able to look at a failed, slow, or expensive request and answer a simple question:

What actually happened?