1. Separate acceptance, processing and acknowledgment
An instruction can have been accepted by the broker even when its producer receives no response. It can also have produced its business effect before the consumer manages to acknowledge the message. These intervals belong in the design. Under PeekLock, an uncompleted message can be delivered again; under ReceiveAndDelete, failure after delivery can leave work lost. In a fictional reconciliation flow, record the instruction identifier, send attempt, receipt, result commit and settlement separately. A network timeout does not prove the earlier operation failed. Support needs to inspect business state before deciding to repeat an instruction. The requirement should describe the acceptable outcome at each failure boundary. This makes recovery decisions traceable rather than dependent on interpreting a single exception as the whole transaction outcome.
2. Distinguish broker deduplication from idempotency
Duplicate detection uses message identity within a configured window; a matching retry can be accepted on send and discarded by the broker. A new identifier on every retry can therefore defeat the intended protection. This feature does not turn an external database commit and settlement into one transaction. The consumer still needs an idempotency contract. In the local model, the processed identifier and business change are written in the same in-memory SQLite transaction. Failure before commit rolls both back; a retry after commit recognizes the identifier and avoids applying another change. The exercise demonstrates only that local boundary. An external effect, such as calling another API, needs an additional contract that this model does not provide. Review that boundary before claiming that duplicate business effects have been prevented.
3. Choose the ordering unit
Sessions group related messages for ordered processing, with one receiver holding the session lock. Choose SessionId from the unit needing order, such as a composite instruction or fictional account, rather than merely to simplify configuration. Different sessions can progress in parallel; a single session for the whole business can restrict throughput. Inside the consumer, preserve the order of dependent effects instead of launching concurrent tasks without coordination. A sequence number alone does not guarantee completion order. Do not assume resubmitting a message from the DLQ returns it to its original historical position either. Define how to reconcile a delayed step after later steps have produced results. Acceptance should include an interrupted sequence and explain whether recovery resumes, compensates or requires a business decision.
4. Size in-flight work against locks
Prefetch can reduce network waiting, but under PeekLock the lock clock starts when the message enters the buffer, before its handler processes it. A large buffer with slow processing can consume lock validity while waiting. Measure processing-time distribution, buffer age and settlement failures before increasing concurrency or renewing locks indiscriminately. In a fictional workload, handlers run in two seconds but some messages wait over a minute in the buffer; optimizing handler code alone does not explain the failure. Reducing prefetched work and adjusting capacity can be more appropriate. Check whether the selected SDK supports prefetch as well. A valid.NET client design should not be transcribed into another SDK without checking available capabilities. Use a representative load rehearsal to establish safe operating settings and alert thresholds.
5. Operate the DLQ as outstanding work
A dead-letter queue separates messages that could not follow normal processing. It is not a self-cleaning archive: define an owner, triage criteria and a procedure to complete or reprocess each message. Classify the cause with sufficient context while avoiding sensitive information in error descriptions. A schema failure requires fixing the contract or transforming the message with approval; a transient outage requires confirming dependency recovery. Resending everything without analysis can repeat the same failure or duplicate effects. For a topic, consider the affected subscriptions and their consumers. Operational reporting should connect message age with outstanding business outcomes rather than show only the accumulated DLQ count. A recovery exercise should demonstrate why a selected message is safe to retry and how completion is confirmed.
6. Close the contract between data and publication
When an application saves data and publishes a message in separate operations, a window exists where only one has completed. The outbox pattern records publication intent with the business change inside the same transaction boundary and uses a later process to deliver the message. Delivery can repeat, so consumer idempotency remains necessary. For Cosmos DB, respect the partition transaction scope instead of assuming atomicity across arbitrary documents. Include supported libraries and protocols in transition planning. Microsoft documentation gives 30 September 2026 as the support-retirement date for older Service Bus libraries and the SBMP protocol; new examples should use supported clients. The local exercise depends on none of those clients and proves no service connectivity. Project acceptance must separately cover deployed integration, monitoring and the operations team's recovery procedure.
import sqlite3
db = sqlite3.connect(":memory:")
db.executescript("CREATE TABLE done(id TEXT PRIMARY KEY); CREATE TABLE balance(n INTEGER); INSERT INTO balance VALUES(100);")
def apply(message_id, delta, fail=False):
with db:
inserted = db.execute("INSERT OR IGNORE INTO done VALUES(?)", (message_id,)).rowcount
if not inserted:
return "duplicate"
db.execute("UPDATE balance SET n=n+?", (delta,))
if fail:
raise RuntimeError("fictional failure before commit")
return "applied"
assert apply("instruction-A", 5) == "applied"
assert apply("instruction-A", 5) == "duplicate"
assert db.execute("SELECT n FROM balance").fetchone[0] == 105
try:
apply("instruction-B", 3, fail=True)
except RuntimeError:
pass
assert db.execute("SELECT COUNT(*) FROM done WHERE id='instruction-B'").fetchone[0] == 0
assert db.execute("SELECT n FROM balance").fetchone[0] == 105
assert apply("instruction-B", 3) == "applied"
assert db.execute("SELECT n FROM balance").fetchone[0] == 108
db.close
print("seven local transaction checks passed; no Service Bus request performed")
Fictional case: the consumer commits a position change, loses its connection before Complete and receives the message again. The recorded identifier prevents applying the change twice; the operator confirms the result and follows settlement.
Common pitfalls
Using a new MessageId for every retry; assuming deduplication makes an external database transactional with the broker; ignoring buffer time; replaying the DLQ without reconciling effects.
Related topics: Transactional outbox · Cosmos DB concurrency · Integration observability
The business contract must survive retries, interruptions and recovery-time reordering.
Reference: Prevent Service Bus message loss and duplicates · AZ-305 objectives 2026-04-17