← Microservices: design boundaries and operate distributed systems
08 / 8 · 60 MIN

Replay, sequences and reconciliation

Plan event recovery without hiding gaps, duplicating effects or exceeding capacity.

1. Identity is not order

Deduplication answers whether an occurrence has already applied in that scope. Sequence answers whether required dependencies are satisfied. The fixture requires consecutive numbers per aggregate: when expecting 1 and receiving 2, it returns sequence-gap without recording completion. After applying 1, it can recover 2. In the experiment, values 7 and 9 produce total 16 and two markers. This rule is specific to the exercise; other contracts may accept out-of-order events or use snapshots. Choose policy from data meaning rather than merely network delivery order.

2. Isolating a message does not resolve its dependency

A quarantine or DLQ separates work needing diagnosis, but the business may still depend on that work. If 19 depends on 18, retrying 19 does not replace the missing effect. Do not advance a counter or insert a completion marker merely to drain the queue. Define how to correct the event, interpret its version or perform authorized recovery while retaining its original reference. SQS documentation highlights DLQ consequences when ordering matters; that provider behavior should not be generalized to every broker. The recovery plan must understand the actual product used.

3. Rebuilding a projection is a new effect

An empty v2 projection is not populated merely because v1 processed the events. The lab demonstrates separate consumer scopes: the same event applies once in each projection. During migration, create an isolated destination, control historical-schema interpretation and avoid enabling external effects during rebuilding. CloudEvents specversion identifies envelope format, not business-model version. Use the type and dataschema contract when applicable. Compare relevant counts and values before cutover, and retain an explicit decision about handling incomplete history.

4. Measure age and net capacity

The experiment’s query filters published=0 before calculating count and min(created). With pending items at ticks 10 and 80 and now=100, it obtains two pending items and a maximum age of 90 ticks. An already published item from tick 1 is outside the population. Low depth does not imply low impact when an old item blocks completion. To estimate draining, account for arrivals: a model with backlog 100, arrivals 20/s and processing 25/s has net capacity 5/s and predicts 20 seconds. This is neither a performance measurement nor an SLA guarantee.

5. Prepare a controlled replay

Before replay, bound events and consumers, confirm the corrected cause, identity policy and available capacity. Keep stop conditions for increasing errors, age or dependency saturation. Do not change historical IDs to bypass deduplication without understanding effects. To investigate sensitive payloads, use a restricted reference and minimal redacted example; avoid spreading personal data into broad channels or baggage. A support team must explain which operations recovered, which still need decisions and who authorizes further attempts. Sent-message counts do not replace that reconciliation.

6. Establish recovery without overstating evidence

The lab’s ten groups check real SQLite transactions, three abrupt process exits and injected relay and consumer interruptions. The queue is local and there is no network traffic. The experiment does not establish power-cut behavior, replication, failover, authentication, real-broker ordering or external effects. For APS handover, complement it with representative target-system tests, a replay runbook, permissions, owners and functional reconciliation. Record these limitations alongside passing results. A technically useful conclusion states what was observed and which boundaries still need evidence.

SELECT count(*) AS pending, min(created) AS oldest_created
FROM outbox
WHERE published = 0;
-- Synthetic fixture: now=100, created=[10,80] -> pending=2, oldest_age=90.
IN PRACTICE

sequence=2 arrives first and remains unmarked. Then 1 and 2 apply: two markers, final sequence 2, total 16.

Common pitfalls

Skip gaps; reuse another projection’s inbox; confuse specversion with payload version; measure only depth; declare an SLA from a model.

Related topics: Transactions and persistence · Events and identity · Operational recovery

Take this idea with you

Recovery ends with reconciled effects and owned exceptions, not merely fewer pending messages.

Create account

Reference: Queue-based load leveling · Microservice architecture patterns and scoped platform examples; primary guidance consulted 2026-09-30