Introduction

Modern applications often need to do more than store data. They also need to trigger work when something changes.

For example, an e-commerce application may create an order and then need to:

  • Send a confirmation email

  • Reserve inventory

  • Start payment processing

  • Update analytics

  • Notify another service

  • Generate an invoice

A common architecture is to store the order in a database and send a message to a separate messaging service.

That works well, but it introduces another system that developers have to manage.

Google Cloud Spanner now provides Spanner Queues, bringing transactional messaging capabilities closer to the database. The idea is to allow applications to combine database changes and messages in a transactional workflow.

This is particularly interesting for distributed applications because it can reduce the gap between "the database says this happened" and "another service was told that it happened."

A simplified architecture looks like this:

Application
     |
     v
Cloud Spanner
     |
     +---- Database Transaction
     |
     +---- Queue Message
     |
     v
Downstream Consumer

The important concept is that the database operation and the message can participate in a coordinated transaction.

What Is Google Cloud Spanner?

Google Cloud Spanner is a distributed relational database designed for applications that need relational database capabilities while scaling across infrastructure.

A traditional database architecture may look like:

Application
     |
     v
Database
     |
     v
Data

A distributed database architecture is more complex:

                 Application
                     |
                     v
              Cloud Spanner
             /      |       \
            /       |        \
           v        v         v
        Node      Node       Node
          |         |          |
          +---------+----------+
                    |
                    v
              Distributed Data

Spanner is designed around distributed operation, which makes it suitable for applications that need strong consistency and large-scale relational workloads.

What Is a Queue?

A queue is a mechanism for passing work from one component to another asynchronously.

For example:

Order Service
     |
     v
Queue
     |
     v
Email Service

The order service does not have to wait for the email service to finish.

It can place a message into the queue and continue.

The email service can process the message later.

Why Applications Use Queues

Queues help separate application components.

Suppose an order API directly calls five downstream services:

Order API
  |
  +--> Inventory
  |
  +--> Payment
  |
  +--> Email
  |
  +--> Analytics
  |
  +--> Invoice

The request becomes dependent on all five services.

A queue-based architecture can reduce that coupling:

Order API
    |
    v
Database
    |
    v
Queue
    |
    +----> Inventory Worker
    |
    +----> Email Worker
    |
    +----> Analytics Worker
    |
    +----> Invoice Worker

The application can complete the main database transaction and allow downstream workers to process asynchronous tasks.

The Transactional Messaging Problem

The difficult part is ensuring that the database change and message publication do not get out of sync.

Consider this workflow:

BEGIN TRANSACTION

Insert Order

COMMIT

Send Message

What happens if the application crashes after the commit but before sending the message?

The order exists.

The message does not.

Now the system has inconsistent state.

The opposite problem is also possible:

Send Message

Insert Order

COMMIT

If the database transaction fails, the message may still exist even though the order was never created.

This is a classic distributed consistency problem.

The Transactional Queue Approach

A transactional queue aims to bring the database operation and message publication into a coordinated transaction.

Conceptually:

BEGIN
   |
   +--> Update Database
   |
   +--> Enqueue Message
   |
COMMIT

If the transaction succeeds:

Database Change
      +
Queue Message
      |
      v
Committed

If the transaction fails:

Database Change
      +
Queue Message
      |
      v
Rolled Back

This can significantly simplify application architecture.

The Outbox Pattern

Before discussing transactional queues, it is useful to understand the outbox pattern.

A common solution to the database-message consistency problem is to write an event into an outbox table as part of the same database transaction.

For example:

BEGIN
   |
   +--> Insert Order
   |
   +--> Insert Outbox Event
   |
COMMIT

A separate worker then reads the outbox:

Outbox Table
     |
     v
Worker
     |
     v
Message Broker
     |
     v
Consumer

This works, but the application now has to manage:

  • Outbox tables

  • Polling

  • Worker processes

  • Retry logic

  • Message state

  • Cleanup

  • Broker integration

A transactional queue can reduce some of this complexity by making the queue part of the database-oriented workflow.

Spanner Queues and the Outbox Idea

Spanner Queues can be viewed as a way to connect transactional database work with asynchronous processing without forcing developers to build every part of an outbox system manually.

The basic idea is:

Application
     |
     v
Spanner Transaction
     |
     +---- Data Change
     |
     +---- Queue Message
     |
     v
Committed
     |
     v
Consumer

This can be especially useful when the event is directly related to the database transaction.

Example: Creating an Order

Consider an order service.

The application receives:

{
  "customerId": 100,
  "productId": 500,
  "quantity": 2
}

The service needs to create the order and trigger downstream processing.

A transactional workflow could look like:

Create Order Request
        |
        v
Validate Request
        |
        v
Spanner Transaction
        |
        +---- Create Order
        |
        +---- Enqueue OrderCreated
        |
        v
Commit
        |
        v
Order Worker

The worker can then process the event:

OrderCreated
    |
    +----> Reserve Inventory
    |
    +----> Send Email
    |
    +----> Start Fulfillment

The order and event remain tied together at the transaction boundary.

Why This Matters for Microservices

Microservices often create distributed workflows.

For example:

Order Service
      |
      v
Inventory Service
      |
      v
Payment Service
      |
      v
Shipping Service

Synchronous communication creates tight dependencies.

A queue can introduce asynchronous boundaries:

Order Service
      |
      v
Queue
      |
      +----> Inventory
      |
      +----> Payment
      |
      +----> Shipping

This can make services more independent.

The database transaction ensures that the message representing the state change is not accidentally separated from the database update.

Asynchronous Processing

A queue allows the producer and consumer to operate at different speeds.

For example:

Producer
  |
  | 100 messages/sec
  v
Queue
  |
  | 50 messages/sec
  v
Consumer

The queue can temporarily hold work while consumers process it.

This can help applications handle traffic bursts without forcing every request to wait for downstream processing.

Queue Consumers

A consumer reads messages and performs the associated work.

For example:

Queue
  |
  v
Consumer
  |
  +--> Read Message
  |
  +--> Validate
  |
  +--> Process
  |
  +--> Acknowledge

A well-designed consumer should be able to handle retries.

A message may be delivered again after a temporary failure.

Therefore, consumers should generally be designed to be idempotent.

What Is Idempotency?

An operation is idempotent when performing it multiple times does not produce an incorrect additional effect.

For example:

Process Order 100
Process Order 100

The system should not accidentally create two invoices simply because the same message was delivered twice.

A consumer might use an event identifier:

EventId = abc123

Before applying a non-idempotent operation, it can determine whether that event has already been processed.

Designing Idempotent Consumers

A consumer can maintain processing state:

Event
 |
 v
Check EventId
 |
 +---- Already Processed ---> Skip
 |
 +---- New Event -----------> Process
                              |
                              v
                         Mark Processed

The exact implementation depends on the application.

The important principle is that asynchronous systems should assume that retries can happen.

Transactional Messaging Does Not Remove Retries

A transactional queue can improve consistency between database changes and message publication.

It does not eliminate every distributed-system failure.

A consumer can still fail after receiving a message.

For example:

Message Delivered
      |
      v
Consumer Starts
      |
      v
External API Call
      |
      X
Consumer Crashes

The message may need to be delivered again.

This is why retry handling and idempotency remain important.

Error Handling

A production consumer should distinguish between temporary and permanent failures.

For example:

Message
  |
  v
Processing
  |
  +---- Temporary Error
  |         |
  |         v
  |       Retry
  |
  +---- Permanent Error
            |
            v
       Failure Handling

A temporary network failure may be retried.

An invalid message schema may need to be rejected and investigated instead.

Poison Messages

A poison message is a message that repeatedly fails processing.

For example:

Message
   |
   v
Attempt 1 -> Failure
   |
   v
Attempt 2 -> Failure
   |
   v
Attempt 3 -> Failure
   |
   v
Attempt 4 -> Failure

Allowing unlimited retries can waste resources.

Production systems should have a strategy for messages that cannot be processed successfully.

Ordering Considerations

Some applications care about message order.

For example:

OrderCreated
     |
     v
OrderPaid
     |
     v
OrderShipped

Processing OrderShipped before OrderPaid may create an invalid workflow.

Developers should therefore understand the ordering guarantees of the queue architecture and design consumers accordingly.

Where strict ordering is not guaranteed or not required, the application should use event state and idempotency rather than relying on message arrival order.

Schema Design for Events

Messages should contain enough information for consumers to process them.

For example:

{
  "eventId": "abc123",
  "eventType": "OrderCreated",
  "orderId": 5001,
  "customerId": 100,
  "createdAt": "2026-10-06T10:30:00Z"
}

A good event contract should be explicit.

Avoid sending an entire database row simply because it is convenient.

Include the information that consumers actually need.

Event Versioning

Event formats can change over time.

For example:

OrderCreated v1

may eventually become:

OrderCreated v2

Consumers should be designed with compatibility in mind.

A change to an event schema can affect every downstream service consuming that event.

Queue-Based Architecture

A larger application might look like:

                    Cloud Spanner
                         |
             +-----------+-----------+
             |                       |
             v                       v
         Application             Spanner Queue
                                     |
                   +-----------------+-----------------+
                   |                 |                 |
                   v                 v                 v
              Inventory          Payments          Notifications
                Worker             Worker              Worker
                   |                 |                 |
                   v                 v                 v
              Inventory DB      Payment API        Email Service

The database remains the system where the primary transaction occurs.

The queue provides an asynchronous path for downstream work.

Common Mistakes

Treating the Queue as a Database Replacement

A queue represents work or events. It should not automatically become the primary source of business data.

Persistent business state should remain in the appropriate data store.

Assuming Messages Are Processed Exactly Once

Distributed systems commonly need to handle retries.

Consumers should be designed to tolerate duplicate processing.

Putting Too Much Data Into Messages

Large messages can make processing and versioning more difficult.

Send the information consumers need rather than duplicating entire database records unnecessarily.

Ignoring Consumer Failures

A producer can successfully publish a message while the consumer fails.

Design retry and failure handling from the beginning.

Creating Tight Coupling Between Events and Database Schemas

Consumers should depend on stable event contracts rather than internal database implementation details.

Assuming Transactions Solve Everything

Transactional publication can solve one consistency problem, but it does not eliminate downstream failures, external API failures, duplicate processing, or business-level errors.

Best Practices

Keep Transactions Small

A transaction should contain the database operations and queue operations that must succeed together.

Avoid putting slow external network calls inside the database transaction.

Make Consumers Idempotent

Assume a message can be delivered more than once.

Use event identifiers or another reliable mechanism to prevent duplicate business effects.

Define Clear Event Contracts

Document event names, required fields, data types, and versioning rules.

Separate Data Changes From Asynchronous Work

The transaction should record the business state and the message needed to trigger downstream processing.

The consumer should perform the longer-running work after the transaction completes.

Design for Failure

Expect:

Timeouts
Retries
Duplicate Messages
Consumer Failures
Network Failures
Invalid Data

A reliable queue architecture is designed around these conditions rather than treating them as rare exceptions.

Monitor Queue Processing

Track:

  • Message throughput

  • Processing latency

  • Failure count

  • Retry count

  • Consumer health

  • Backlog size

A queue that silently accumulates messages is a production problem even if the database itself appears healthy.

Advantages

Stronger Consistency Between Data and Messages

The most important advantage is reducing the chance that a database change succeeds while its corresponding message is lost. By bringing the message operation into the transactional workflow, the application can maintain a much stronger relationship between persistent state and asynchronous work.

Less Custom Outbox Infrastructure

Teams using an outbox pattern often need tables, polling workers, cleanup logic, retry handling, and message publication code. A database-integrated queue can reduce some of this application infrastructure and make the architecture easier to reason about.

Better Support for Event-Driven Applications

Applications can record a state change and produce the corresponding asynchronous event as part of the same workflow. This makes transactional business operations easier to connect with downstream services.

Useful for Distributed Systems

Spanner is designed for distributed workloads, and transactional queue capabilities can help developers build asynchronous processing into applications without treating messaging as a completely unrelated subsystem.

Cleaner Microservice Boundaries

A service can commit its own state and publish an event for other services to consume. This reduces the need for synchronous calls between every component and allows consumers to process work independently.

Disadvantages

Consumers Still Need Robust Error Handling

Transactional message publication does not guarantee that the consumer will successfully process the message. Consumers still need retry handling, idempotency, failure handling, and monitoring.

Adds Asynchronous Complexity

Queues introduce eventual processing behavior. Developers now need to reason about message delivery, retries, processing state, failures, and monitoring in addition to normal database transactions.

Event Contracts Need Maintenance

Once multiple services consume an event, changing its structure can become difficult. Event schemas need versioning and compatibility rules to prevent one service from breaking another.

Debugging Can Become More Difficult

A synchronous request is usually easier to trace from start to finish. With queues, the original transaction and downstream processing happen at different times, so developers need distributed tracing and useful identifiers to connect the operations.

Not Every Workflow Needs a Queue

A queue adds value when work can be asynchronous or when services need decoupling. For a simple operation where the caller needs an immediate result, introducing asynchronous messaging may add unnecessary complexity.

Troubleshooting

Database Transaction Succeeds but Consumer Does Not Process the Message

First verify that the message was actually created and is visible to the consumer.

Then check:

Queue Status
Consumer Health
Permissions
Processing Errors
Retry State

Do not immediately change the database transaction if the problem is actually in the consumer.

Messages Are Being Processed More Than Once

Check whether the consumer is idempotent.

Use an event identifier and maintain processing state where necessary.

Queue Backlog Keeps Increasing

A growing backlog generally means producers are generating work faster than consumers can process it.

Check:

Producer Rate
Consumer Rate
Consumer Errors
Processing Time
Consumer Capacity

Adding more consumers may help, but first identify whether the bottleneck is CPU, database access, an external API, or application logic.

Consumer Keeps Retrying the Same Message

Inspect the underlying error.

If the message itself is invalid, retries will not solve the problem.

The system needs an appropriate failure-handling path for messages that cannot succeed.

Downstream Service Receives Duplicate Operations

Check the consumer's idempotency logic.

A retry should not create a second business effect when the first attempt already completed successfully but the acknowledgement path failed.

A Practical Order Processing Example

Consider an order application.

The API receives:

POST /orders

The service performs:

BEGIN TRANSACTION
      |
      +---- Insert Order
      |
      +---- Insert Order Items
      |
      +---- Enqueue OrderCreated
      |
COMMIT

After the transaction succeeds:

OrderCreated
     |
     +----> Inventory Worker
     |
     +----> Notification Worker
     |
     +----> Analytics Worker

The inventory worker can reserve stock.

The notification worker can send an email.

The analytics worker can record the business event.

The original order transaction does not have to wait for all three operations.

Why This Architecture Scales Better

Suppose email processing becomes slow.

In a synchronous design:

Order API
    |
    v
Email Service
    |
    X
Slow

The order request can become slow as a result.

With asynchronous processing:

Order API
    |
    v
Spanner
    |
    v
Queue
    |
    v
Email Worker

The order transaction can complete while the email worker processes the task independently.

This is one of the main reasons queues are useful in distributed systems.

When Should You Use Spanner Queues?

Transactional queues are a good fit when:

  • A database change needs to trigger asynchronous work.

  • The message and database change should remain consistent.

  • The application has downstream consumers.

  • Work can happen after the main transaction commits.

  • Services need to be loosely coupled.

  • The application already uses Cloud Spanner.

Examples include:

  • Order processing

  • Notifications

  • Inventory updates

  • Billing workflows

  • Audit processing

  • Data synchronization

  • Background jobs

  • Event-driven microservices

A queue may not be necessary when the operation is simple and the caller needs the result immediately.

Summary

Spanner Queues bring transactional messaging capabilities into Google Cloud Spanner, addressing one of the common challenges in event-driven applications: keeping database changes and asynchronous messages consistent.

The traditional problem looks like this:

Database Commit
      |
      X
Application Crashes
      |
      X
Message Never Sent

A transactional workflow can instead connect the two operations:

Transaction
   |
   +---- Database Change
   |
   +---- Queue Message
   |
   v
Commit

This can reduce the need for custom outbox infrastructure and make it easier to build event-driven applications around Spanner.

However, developers still need to design for the realities of distributed systems. Consumers can fail, messages can be retried, downstream APIs can be unavailable, and event contracts can evolve.

The strongest architecture combines transactional message publication with idempotent consumers, clear event contracts, controlled retries, monitoring, and good observability.

For applications already using Cloud Spanner, Spanner Queues provide an interesting way to connect relational transactions with asynchronous processing without treating messaging as a completely separate concern. The result can be a cleaner architecture for microservices and event-driven workloads where database state and downstream work need to stay closely connected.