A failed statement does not define the whole unit
The lab changes 50 to 60 inside a transaction and attempts to insert an existing token into a UNIQUE column. Under default ABORT policy, INSERT fails but the transaction remains active: the same connection reads 60 and another reads 50. The script executes ROLLBACK and confirms that the external value remains 50. In a second sequence, INSERT OR ROLLBACK ends the transaction on conflict. The constraint code is the same, but policy and the fate of prior changes differ. In a fictional batch whose two steps are indivisible, committing the earlier change after catching the exception violates the contract. Testing must check final state and the response to the consumer.
Incomplete checkpoint with a successful call
In another sequence, A retains a read of 50 while B commits 51, 52 and 53. The PASSIVE checkpoint returns zero status but copies fewer pages than are present in WAL. The old read still observes 50 while a new read sees 53. After ending the old read, TRUNCATE returns [0,0,0] and a fresh query retains 53. The exercise demonstrates the difference between finishing the call and completing all work. Retain all three fields and transaction state; do not reduce observation to success or error. The script disables automatic checkpoints only to control this sequence. This is not a universal configuration recommendation or authorization to remove an active WAL to reclaim space.
Turn a mechanism into acceptance conditions
Imagine a supplier delivering a patch after local tests while the target uses another engine and a transaction manager. Review should list what remains: actual driver rules, failure between steps, replay with changed data, reservation duration and prolonged readers. Define expected final state and evidence for each condition. Add stopping criteria, recovery and a decision owner. If cut-off leaves no time for the exercise, temporary job serialization can be assessed with explicit cost, deadline and risk. One successful batch under a workaround does not demonstrate permanent correction. Schedule pressure informs the risk decision but does not replace observation of the paths that failed.
Operational handover and checkable learning
In a fictional postmortem, separate a write held during an external call, missing concurrency tests and a message hiding distinct codes. These factors are related, but each action needs its own owner and outcome. Delivering a log field improves diagnosis; it does not prove a shorter transaction. An executed new test needs to show that it distinguishes faulty and corrected versions. Prepare a RUN guide with applicability conditions, retry limits, state checks and escalation destination. The proposed workshop asks participants to explain these limits to a supplier in English and record the decision. The guide exists, but no human session or independent specialist review has taken place; that evidence remains outstanding.
python3 content/labs/problem-transactions/run.py --output /tmp/dr-problem-atomicity.json
# Compare: abortOnlyFailedStatement / explicitRollbackOfUnit
# Compare: readerLimitsCheckpoint / checkpointAfterReaderEnds
# Check final state, not only whether the call raised an exception.A rejected INSERT leaves a previous UPDATE pending; a PASSIVE checkpoint can end with zero status while pages still remain to copy.
Common pitfalls
Confusing exceptions with full rollback, an error-free call with fully completed work or a local checkpoint with power-failure recovery.
Related topics: Residual risk and workarounds · RUN handover and prevention
Accept a correction using functional state and observed operational limits, including failures between steps and prolonged readers.
Reference: The ON CONFLICT Clause · Problem management practices 2026-09; ServiceNow Brazil examples with scoped plugins and properties