A queue provides limited headroom
In the fictional Tide case, a queue holds 200 requests with capacity 600. Arrivals are 120 requests per minute and departures eighty. In a constant model without loss or retries, net growth is forty and capacity is reached in (600−200)/40=10 minutes. This is an occupancy projection, not a promise of ten loss-free minutes. Maximum retry time may expire earlier, requests may have different sizes, and rates may change. Record the configured unit: requests, items, and bytes are not interchangeable. Occupancy of 240/600=40% in requests reveals neither span count nor consumed memory. In a real system, also observe ages, resources, and destination responses. Increasing queue_size may provide headroom but does not remove departures remaining below arrivals. Decide mitigation with responsible owners and make its effect on observability coverage explicit.
Persistence and pressure have boundaries
An in-memory-only queue can lose pending data during an abrupt restart. file_storage enables queue persistence but depends on storage capacity, permissions, and lifecycle. A directory on ephemeral storage does not guarantee recovery after instance replacement. A gateway WAL also cannot recover spans lost at the agent before arrival. Draw the path and identify protection at each hop. Persistence establishes neither exactly-once delivery nor freedom from retry timeout, full queues, or disk failure. The memory_limiter has a different role: it refuses data under pressure and depends on preceding components’ response. It is neither a persistent queue nor an absolute OOM guarantee. For each failure hypothesis, write which data may remain in memory, on disk, in transit, or outside any recoverable source. The architecture should declare tolerated loss or required protection according to purpose and authorized decisions.
Observe recovery without erasing gaps
A responding destination does not mean the queue has drained. With 600 pending requests, arrivals of one hundred per minute, and possible departures of 160, net draining is sixty: the model requires ten minutes. The calculation 600/160 would apply only if arrivals stopped. Confirm actual behavior and bound gaps from the affected period. Increasing max_elapsed_time after a drop does not reconstruct lost data. During upgrades, compare internal-metrics configuration and actually exposed names; names, suffixes, and levels may change. Converting an absent series to zero can produce a false healthy indication. A recovered Collector also does not demonstrate completion of the observed payment or batch. Close investigation using evidence of new flow, backlog, known losses, and useful service outcome, assigning open questions to appropriate owners. Do not remove inconvenient metrics to make reporting green.
Tide decision rehearsal
Use forty minutes with L3, the pipeline owner, a destination representative, and an observer. Spend ten drawing hops and indicator units, ten calculating headroom and retry limits, ten considering mitigations, and ten accepting recovery. Introduce two facts during discussion: the agent has memory only and gateway storage is ephemeral. Ask the group to revise its initially stated guarantees. Deliverables include facts, hypotheses, capacity, unknown age, loss risk, authorized action, owner, and closure evidence. The observer records whether the group distinguished model from measurement, availability from draining, and telemetry from business outcome. This is a fictional-data practice guide not yet executed with participants. Calculations were checked locally; persistence, saturation, retry, and failover still need their own rehearsals in an authorized environment. Do not attribute coverage to the preceding laboratory that it does not have.
TIDE GUIDE | 40 min: 10 + 10 + 10 + 10
Path and owners per hop:
Queue unit / capacity / occupancy / window:
Arrivals / departures / net growth:
Age and retry limit / missing information:
Memory / disk / volume lifecycle:
Dropped data / possible recovery source:
Mitigation / authority / coverage effect:
Draining and new-arrival evidence:
Business outcome to verify:
Open questions / owner / review:
Observation: not observed | with help | in rehearsal without helpCapacity: (600−200)/(120−80)=10 min. Draining: 600/(160−100)=10 min. The two results depend on different assumptions.
Common pitfalls
Queue as loss-free guarantee; one-hop WAL as total durability; restart as cure; available destination as drained backlog; zero constructed from absence.
Related topics: Capacity and backpressure · Persistence and retry · APS handover
Recovery requires understanding the path and checking what arrived, what is waiting, and what can no longer be recovered.
Reference: Collector resiliency · Observability 2026-09; selected OpenTelemetry, Prometheus and Dynatrace Classic concepts