Shutdown starts with the application contract
Before stopping a service, identify the effective signal, its recipient and work in progress. An image may define STOPSIGNAL, while creation or the operation may introduce overrides. In the fictional case, the batch needs 35 seconds and stop allows ten. The solution is more than increasing a number: review work admission, draining, window and completion criteria. A terminated process does not confirm file delivery to the consumer. Recovery needs to cover work left incomplete or with uncertain outcomes after the operation.
Connect timeline and likely cause
A fictional evidence sheet records stop at 02:00:00, termination at 02:00:10, code 137 and OOMKilled=false. Timing is consistent with expiry of a ten-second deadline, but the conclusion needs instance events and shutdown behavior. Avoid turning an exit code into a universal cause. Preserve identity, command, timestamps and recent change. If someone started a worker through docker exec, that process does not automatically become part of permanent container startup. Reconstruct the sequence including manual interventions that affected the instance and the work it performed.
Restart policy and maintenance intent
Policies respond to different events. On-failure uses an error exit and does not understand an incomplete business file despite zero exit. Always may start again after daemon restart even following manual stopping. Unless-stopped preserves stopped intent in that context. Also review external automation that may start the service. If mitigation uses docker update, reconcile effective configuration with repository-maintained Compose; recreation may reapply the old declaration. Document the decision before the next intervention, including owner and scope so that maintenance behavior matches the intended state.
Health differs from lifecycle
A healthcheck may fail while the main process keeps running. On standalone Engine without an additional controller, restart policy should not be treated as automatic recovery from every unhealthy state. Define who responds to the signal, which observation confirms impact and what mitigation may be authorized. A localhost check may also pass while an external client fails because it measures another path. Compare origin, listener, publication, authentication and business outcome. These decisions use documentation and original scenarios; this expansion did not exercise them against a local daemon.
Logs: capacity, loss and evidence
Non-blocking mode uses an intermediate buffer; new records may be lost when it fills. That protects the writer from some destination pressure but does not guarantee a complete trail. In a fictional case, 200 operations are reconciled while only 150 central events exist. The difference requires observability investigation, not automatic replay to manufacture missing records. Also review json-file rotation and retention and effective instance configuration: changing daemon defaults does not automatically update existing containers. Plan change application without deleting evidence needed to understand the incident.
Recovery workshop with uncertain outcomes
A fictional worker sends an operation, loses acknowledgement and exits with an error. Restart recovers a process but does not prove the recipient rejected the operation. Retain business identity, reconcile the outcome and decide replay according to the integration contract. Prepare the handover: “The process restarted; the previous operation’s outcome remains unconfirmed.” Add instance, timestamps, exits, policy, available logs and pending decisions. The aim is service recovery while preserving the distinction between being in execution and having completed the expected business work.
Fictional sheet: stop 02:00:00, deadline 10 s, exit 02:00:10/137, OOMKilled=false. The deadline-expiry hypothesis needs events and instance identity; this is not a recorded Docker execution.
Common pitfalls
Treating always as permanent suspension after stop; using unhealthy as exit; interpreting zero as business delivery; expecting exec-process recovery; taking missing logs as missing effects.
Related topics: Processes and images · Data and mounts · Compose: overrides, data and acceptance
Connect policy and timeline to actual work. A restart may recover execution without reconciling operations, fixing the cause or completing evidence needed to accept the service.
Reference: Start containers automatically · Docker Engine Linux containers, BuildKit and Compose; official documentation consulted 2026-09-30; version-dependent behavior explicitly scoped