← REST APIs: integrate applications and diagnose failures
07 / 8 · 60 MIN

Recover requests without duplicating work

Distinguish attempt, operation and outcome; decide when to look up, retry or escalate.

1. A timeout leaves an unanswered question

A request crosses several boundaries: client, intermediary, application and persistence. If the response is lost after job creation, the client knows about a communication failure but not the service outcome. In the lab, the server creates a report-generation request and deliberately closes the connection before responding. The client receives RemoteDisconnected. That observation should keep the outcome unknown until further evidence exists. Do not mark the job failed or completed from the exception alone. During L3 support, record the operation reference, time window and layer where communication ended.

2. Identity belongs to intent

An attempt is a submission; an operation is the intent that submission seeks to realize. An idempotency key lets the contract connect several attempts to the same intent. In the fixture, identity includes tenant, endpoint and key. Repeating the same combination and payload returns the saved acceptance without creating another job. Changing the payload with a still-valid key receives 409. This is an explicit lab design choice. Stripe documentation illustrates a real contract with retention and parameter comparison, but does not make every provider’s POST operations idempotent.

3. Scope, access and expiration

The service must check current access before returning a saved response. A user may lose authorization after the first request; knowing the key does not restore that access. Distinct tenants may use the same key text without sharing results when the namespace includes the authenticated tenant. The lab simulates this context with X-Lab headers; these are test selectors, not usable authentication. Retention also limits the guarantee. After ten synthetic ticks, the fixture forgets the key but retains the job. Retrying then may create another job. Ten ticks are a teaching choice, not an HTTP or vendor rule.

4. Saved acceptance and current state

A 202 response may point to a resource where the consumer tracks work. In this exercise, POST saves an id and state=queued. The runner later changes the job state to succeeded, simulating worker completion. GET shows current state while POST replay still returns initial acceptance. Do not interpret replay as a return to the queue. Track job identity and query the appropriate resource under the contract. At work, distinguish API availability, acceptance rate and functional completion of the process the business expects.

5. One budget for all attempts

A recovery policy needs attempt limits, total time and authority to resubmit. In the guided example, 1.5 seconds remain and the response requests a two-second wait. The supplied policy requires honoring the wait and forbids starting after the deadline; use the defined recovery path without starting a new request. Consider layers too: three total client attempts, each with three gateway attempts, may produce nine service requests. Count retries where they actually occur. A local resilience improvement may increase load and delay other consumers without coordination.

6. Support exercise and shift handover

Analyze the case of a lost response whose key has expired. Prepare a handover containing the business reference, known or unknown outcome, last authorized lookup, retention window and owner of the resubmission decision. If lookup confirms a completed job, validate the expected result; if it shows failure, apply the approved procedure; if it remains ambiguous, keep that ambiguity explicit. The lab executes real HTTP only on loopback with in-memory state. It does not exercise banking APIs, financial effects, TLS, persistence after restart or cross-region recovery.

# Observações do fixture local, com dados sintéticos
POST /jobs Idempotency-Key: lost -> ligação fechada após criar job 1
POST /jobs Idempotency-Key: lost -> 202 {"id":"1","state":"queued"}
GET /jobs/1 -> 200 {"id":"1","state":"succeeded"}
# O POST repetido conserva a resposta inicial; GET mostra o estado atual.
IN PRACTICE

POST creates job 1, its response is lost, and retrying the same identity returns job 1. Count remains one; after key expiry, another submission creates job 3 in the experiment.

Common pitfalls

Change keys on every retry; confuse 202 with completion; ignore retention; treat test headers as authentication.

Related topics: HTTP and HTTPS · Concurrency control · Observability and recovery

Take this idea with you

Recover original intent using identity, lookup and explicit limits; do not infer outcome from transport alone.

Create account

Reference: Idempotent requests · HTTP semantics RFC9110; OpenAPI3.2.1; selected primary standards and provider contracts consulted2026-09-30