Introduction
Event-driven applications are designed around a simple idea: one system produces an event and another system reacts to it.
For example:
Order Service
|
v
Order Created Event
|
v
Amazon EventBridge
|
+----> Billing
+----> Notifications
+----> Analytics
This architecture reduces direct dependencies between services, but it introduces an operational challenge: events can fail to reach or process successfully at the destination.
A consumer may be temporarily unavailable, a downstream API may return an error, or an event handler may contain a software defect.
Recovery therefore needs to be part of the event architecture.
Amazon EventBridge provides capabilities such as event delivery retries, dead-letter queues, and event replay that can help applications recover from failures. Ordering is a separate concern and should be designed carefully because event-driven systems should not assume that every event arrives in the order it was created.
This article explains how ordering, retries, dead-letter queues, and replay fit together and how to design a reliable recovery process.
Why Event Ordering Matters
Consider an order workflow:
Event 1: OrderCreated
Event 2: PaymentCompleted
Event 3: OrderShipped
The consumer expects:
OrderCreated
↓
PaymentCompleted
↓
OrderShipped
If events are processed out of order:
OrderCreated
↓
OrderShipped
↓
PaymentCompleted
the consumer may temporarily or permanently reach an invalid state.
However, ordering is not automatically required for every event-driven system.
For example, independent events such as:
UserViewedProduct
UserOpenedPage
may not need strict ordering.
The first design question should therefore be:
Does the business process actually depend on event order?
EventBridge Delivery and Consumer Failures
An event can be successfully accepted by the event bus while downstream processing still fails.
A simplified flow is:
Producer
|
v
EventBridge Event Bus
|
v
Rule
|
v
Target
|
+---- Success
|
+---- Failure
The failure may happen because:
The target is unavailable
The target returns an error
A network dependency fails
The consumer times out
The consumer has a software bug
The event payload is invalid for the current consumer version
A reliable architecture needs a way to distinguish temporary failures from events that require investigation.
Retries Help With Temporary Failures
A transient failure does not necessarily mean that an event is permanently invalid.
For example:
Event
|
v
Consumer
|
X---- Temporary failure
|
v
Retry
|
v
Consumer
|
v
Success
EventBridge supports retry behavior for target delivery failures.
The exact retry configuration should be chosen according to the target and business requirements.
The important principle is that retries should address transient failures, not hide permanent failures indefinitely.
Dead-Letter Queues Capture Failed Events
Eventually, an event may continue failing.
A dead-letter queue (DLQ) provides a place to retain events that could not be successfully delivered after the configured retry behavior.
The architecture becomes:
EventBridge
|
v
Target
|
+---- Success
|
+---- Retry
|
+---- Success
|
+---- Failure
|
v
DLQ
This prevents permanently failing events from disappearing silently.
A DLQ also gives operations teams an opportunity to investigate and recover events later.
What Event Replay Adds
A DLQ preserves failed events, but recovery still requires a way to process them.
Event replay provides another mechanism for reprocessing archived events.
A simplified recovery flow is:
Event Archive
|
v
Replay
|
v
EventBridge
|
v
Target
This can be useful after:
A consumer bug is fixed
A downstream service recovers
A deployment introduced a temporary failure
A new consumer needs historical events
An event-processing rule needs to be corrected
Replay should be treated as a controlled operational process rather than an automatic replacement for retries.
Retry vs DLQ vs Replay
These mechanisms solve different problems.
Mechanism | Primary purpose | Typical timing |
|---|---|---|
Retry | Recover from temporary target failure | Immediately/automatically |
DLQ | Preserve events that could not be delivered | After delivery failures |
Archive | Retain events for future use | Before a recovery need |
Replay | Reprocess retained historical events | Deliberately initiated |
Application retry | Recover from downstream business/API failure | Consumer-controlled |
Understanding this distinction prevents teams from trying to use one mechanism for every failure mode.
A Production Recovery Architecture
A robust architecture might look like:
Event Producer
|
v
EventBridge Bus
|
+------------+------------+
| |
v v
Rule A Rule B
| |
v v
Consumer A Consumer B
| |
+---+---+ +---+---+
| | | |
Success Retry Success Retry
| |
v v
DLQ DLQ
Event Archive
|
v
Replay
This architecture separates normal delivery from recovery.
Do Not Assume Replay Solves Ordering
Replay can reprocess events, but replaying events does not automatically solve application-level ordering requirements.
Suppose the original events were:
1. CustomerCreated
2. AddressUpdated
3. CustomerDeleted
If the consumer processes them incorrectly, replaying the events without understanding state transitions may reproduce the same problem.
The consumer should be designed to handle the ordering requirements of the business process.
Possible strategies include:
Sequence numbers
Version numbers
Event timestamps
State validation
Idempotency keys
Application-level ordering
Partitioning strategies where supported
Use Idempotency for Safe Replay
Replay means an event may be processed more than once.
Therefore, consumers should avoid assuming:
one event = one execution
Instead:
same event
|
+---- first processing --> execute
|
+---- replay --> detect previous processing
For example:
public async Task HandleAsync(OrderCreatedEvent message)
{
if (await _processedEvents.ExistsAsync(message.EventId))
return;
await _orderService.CreateAsync(message.OrderId);
await _processedEvents.RecordAsync(message.EventId);
}
The implementation must also consider transactional consistency. Recording an event as processed separately from the business operation can itself create failure scenarios.
For important workflows, use a transactional or otherwise reliable idempotency design appropriate to the data store.
Handle Partial Processing Carefully
Consider a consumer that performs three operations:
1. Update database
2. Call payment service
3. Send notification
If the first two succeed and the third fails, replaying the complete event may repeat operations that already succeeded.
This is why idempotency and workflow state are important.
A more explicit state model might be:
OrderCreated
|
v
DatabaseUpdated
|
v
PaymentProcessed
|
v
NotificationPending
The recovery process can then resume from the appropriate state rather than blindly repeating every operation.
Designing a Replay Process
Replay should be controlled.
A useful operational process is:
Step 1: Identify the Failure
Determine why events failed.
Examples:
Consumer deployment bug
API outage
Invalid transformation
Authentication failure
Step 2: Fix the Underlying Problem
Do not replay events while the same failure still exists.
Step 3: Select the Required Event Range
Avoid replaying unrelated events.
Step 4: Validate Consumer Behavior
Test the fixed consumer with representative events.
Step 5: Replay Carefully
Start with an appropriate scope and monitor the results.
Step 6: Verify Downstream State
Check whether the intended business state was restored.
This is especially important for events that trigger external side effects.
Monitoring Failed Events
A production event system should expose useful operational metrics.
Track:
Metric | Why it matters |
|---|---|
Events published | Producer activity |
Delivery failures | Consumer health |
Retry count | Transient failure detection |
DLQ message count | Persistent failures |
Processing latency | Consumer performance |
Replay volume | Recovery activity |
Duplicate processing | Idempotency health |
Consumer error rate | Application reliability |
A sudden increase in DLQ messages should trigger investigation rather than simply repeated replay.
Common Mistakes
Assuming Event Delivery Means Successful Processing
An event being accepted by EventBridge does not mean every target completed its business operation.
Treating Retries as a Permanent Recovery Mechanism
Retries are useful for transient failures. Permanent failures need investigation.
Replaying Without Fixing the Consumer
Replay can reproduce the same failure.
Ignoring Duplicate Processing
Retries and replay can cause an event to reach a consumer multiple times.
Assuming Events Are Automatically Ordered
Do not build business logic around ordering unless the architecture explicitly guarantees the required behavior.
Replaying Large Volumes Without Monitoring
A replay can generate significant downstream traffic.
Ignoring External Side Effects
Sending emails, charging payments, or creating external records can be dangerous to repeat.
Troubleshooting Failed Events
Step 1: Check the Target
Determine whether the destination service is healthy.
Step 2: Inspect Error Patterns
Group failures by status code, exception type, or failure reason.
Step 3: Check the DLQ
Determine whether events are accumulating.
Step 4: Identify the First Failed Event
The first failure can reveal whether the issue is systemic or related to specific payloads.
Step 5: Verify Consumer Idempotency
Confirm that replaying the event will not duplicate an external side effect.
Step 6: Test the Fixed Consumer
Use representative failed events before initiating broad recovery.
Step 7: Monitor Replay
Watch target health, error rates, processing latency, and downstream effects.
Best Practices
Decide whether ordering is actually required by the business process.
Configure appropriate retry behavior for transient failures.
Use dead-letter queues for events that cannot be delivered successfully.
Archive events when historical replay is a real recovery requirement.
Make consumers idempotent.
Use event IDs or equivalent idempotency keys.
Treat replay as an operational recovery procedure.
Fix the underlying problem before replaying events.
Monitor replay volume and downstream effects.
Protect external side effects from duplicate execution.
Track failed events with enough metadata for investigation.
Test disaster-recovery and replay procedures before they are needed.
Advantages and Disadvantages
Advantages | Disadvantages |
|---|---|
Retries can recover transient failures | Retries can increase processing load |
DLQs preserve failed events | Failed messages require operational attention |
Archives support historical recovery | Retention and storage need management |
Replay can recover after consumer fixes | Replay can produce duplicate side effects |
Event-driven recovery can reduce manual data repair | Ordering and state consistency still require application design |
Production Checklist
[ ] Event ordering requirements are documented
[ ] Retry behavior is configured
[ ] Dead-letter handling is configured
[ ] Events are archived when replay is required
[ ] Consumers are idempotent
[ ] External side effects are protected against duplicates
[ ] Failed events can be investigated
[ ] Replay procedures are documented
[ ] Replay has an appropriate scope
[ ] Consumer fixes are tested before replay
[ ] Monitoring covers retries and DLQ growth
[ ] Recovery procedures are tested periodically
Summary
Reliable EventBridge architecture is not just about successfully publishing events. It also needs a plan for what happens when delivery or processing fails.
Retries are useful for temporary failures. Dead-letter queues preserve events that continue to fail. Event archives provide historical data that can be replayed after a consumer or dependency has been fixed.
Ordering is a separate application concern. If business logic depends on event sequence, the consumer should explicitly protect that requirement rather than assuming that replay will restore the correct state.
The most important production principle is to design consumers for idempotent, observable, recoverable processing. When retries, DLQs, archives, and replay are combined with those application-level safeguards, failed events become manageable operational problems instead of permanent data-loss scenarios.

Join the conversation! Your thoughts help the community grow.