Concept and mechanism
Startup, liveness, and readiness answer different questions. A startup probe allows startup time; liveness can restart a container that stopped functioning; readiness controls traffic eligibility. A slow remote dependency does not automatically justify restarting a healthy local process. If every replica checks that dependency through liveness, they can restart together and worsen the incident. Design checks for the service contract and assess whether useful degraded behavior exists.
Guided application
Start with status and events, then read logs from the relevant container and execution. --previous can retrieve the prior instance’s output when available. kubectl top depends on the metrics pipeline; its failure does not prove main API unavailability. For Pending with insufficient memory, compare requests and committed capacity on eligible nodes. When maintaining manifests, consult served versions and schema migrations: changing only apiVersion can leave incompatible fields. Record observations and the hypothesis each action checks so team handover preserves the reasoning.
If the event says FailedScheduling, investigate placement before probes. If the container restarted, inspect previous logs before another rollout loses evidence.
Common pitfalls
Using restart as universal diagnosis; reading absent metrics as zero usage; confusing authentication with a removed API.
Related topics: Configuration, secrets, and identity · Resources, security, and extensions
The next action should reduce a specific uncertainty about the failure.
Reference: Container probes · CKAD Kubernetes v1.35