← Kubernetes: operate workloads and recover services
10 / 12 · 70 MIN

Probes, restarts and running configuration

Distinguish traffic unavailability, container restart and a pre-start configuration block using cluster observations.

Readiness changes traffic eligibility

The API has separate readiness and liveness endpoints. Removing /tmp/ready makes the former return 503 while leaving the process and ordinary response running. The Pod stays Running, becomes unready and retains its UID and restart count. The EndpointSlice changes to ready=false and the ordinary Service request fails. Direct access to the Pod IP still responds. Readiness therefore does not act as a firewall: it informs mechanisms using that state while other paths can remain accessible. During an incident, interpret the condition together with Service configuration and the path actually used by the affected consumer.

A publication exception does not heal the application

The experiment temporarily enables publishNotReadyAddresses on the Service. The EndpointSlice ready condition becomes true although the Pod remains unready, and ordinary requests reach the API again. This demonstrates why reading a ready field requires identifying its source and Service options. The experiment does not recommend the change as a universal mitigation. In a real application, the probe may indicate inability to serve requests safely. The script restores the option to false, confirms destination exclusion and only recovers readiness after restoring the marker. Operational decisions should preserve the intended meaning of the probe instead of merely making traffic appear successful.

Recover without restarting

After restoring /tmp/ready, the script observes a ready Pod, an eligible destination and a response through the Service. The restart count remains unchanged. This sequence demonstrates recovery without replacing the container. In a real service, the cause may require repairing a dependency or completing initialization; the marker is only this exercise’s synthetic mechanism. Avoid deleting Pods before collecting evidence or when deletion does not address the cause. Record the transition and confirm functional outcome. A recovered probe alongside an accumulated backlog may justify keeping an incident open and monitoring work completion before declaring recovery.

Liveness and Pod identity

Removing /tmp/live produces enough liveness failures to restart the container. The script confirms that the Pod UID stays unchanged, the restart count increases and the API becomes ready again. Startup code recreates the markers, which specifically explains recovery here. Previous-container logs contain the synthetic startup message and can be requested with --previous. They do not constitute a complete archive of every instance or restart. Preserve relevant evidence before losing context and distinguish container restart from controller-driven Pod replacement. The observed recovery does not establish that restarting would repair every dependency failure in a real application.

Failure before the process starts

Another Pod requires key MODE from a ConfigMap that does not yet exist. Its container waiting state reports CreateContainerConfigError and the application has not started. Creating the ConfigMap with MODE=batch allows execution; the process prints batch and becomes ready without restarting a previously running process. The problem lies in configuration preparation rather than an overly strict liveness check. Inspect the message, reference and namespace before changing probes or increasing resources. In a real system, also confirm who supplies the object and whether delivery ordering makes it available when needed by the workload.

Declared and consumed configuration

After startup, the ConfigMap changes to MODE=online, but reading the environment inside the container still returns batch. The value was supplied as an environment variable at startup; changing the object does not rewrite the existing process environment. Define controlled consumer renewal and validate the effective value afterwards. The lab does not execute that additional rollout or measure ConfigMap volume propagation. Evidence covers the difference between object and environment in this specific configuration. Handover should include an update method, functional criteria, compatible rollback and outstanding limitations, alongside any independent review needed before using the procedure in production.

kubectl --context OWNED -n LAB get pod POD -o json
kubectl --context OWNED -n LAB describe pod POD
kubectl --context OWNED -n LAB logs POD -c api --previous
# Inspect Ready, container waiting reason, restartCount and the relevant probe.
# A prior-container log is not a complete historical log archive.
IN PRACTICE

Removing the readiness marker excluded the destination from ordinary traffic without restarting; removing the liveness marker restarted a container in the same Pod.

Common pitfalls

Running as service readiness; readiness as network isolation; bypassing readiness as recovery; updating a ConfigMap as changing process environment.

Related topics: Traffic and probes · Operations and recovery

Take this idea with you

Connect each observation to the decision it supports and confirm recovery through the consumer path while keeping experimental limits explicit.

Create account

Reference: Configure probes · Kubernetes v1.37 concepts; current official documentation consulted 2026-09-30; cluster versions and plugin capabilities must be confirmed

Kubernetes® is a registered trademark of The Linux Foundation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by The Linux Foundation. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.