← CKAD: Kubernetes applications in production
04 / 7 · 24 MIN

Application observability and maintenance

Choose evidence according to failure phase and expected behavior.

Concept and mechanism

Startup, liveness, and readiness answer different questions. A startup probe allows startup time; liveness can restart a container that stopped functioning; readiness controls traffic eligibility. A slow remote dependency does not automatically justify restarting a healthy local process. If every replica checks that dependency through liveness, they can restart together and worsen the incident. Design checks for the service contract and assess whether useful degraded behavior exists.

Guided application

Start with status and events, then read logs from the relevant container and execution. --previous can retrieve the prior instance’s output when available. kubectl top depends on the metrics pipeline; its failure does not prove main API unavailability. For Pending with insufficient memory, compare requests and committed capacity on eligible nodes. When maintaining manifests, consult served versions and schema migrations: changing only apiVersion can leave incompatible fields. Record observations and the hypothesis each action checks so team handover preserves the reasoning.

IN PRACTICE

If the event says FailedScheduling, investigate placement before probes. If the container restarted, inspect previous logs before another rollout loses evidence.

Common pitfalls

Using restart as universal diagnosis; reading absent metrics as zero usage; confusing authentication with a removed API.

Related topics: Configuration, secrets, and identity · Resources, security, and extensions

Take this idea with you

The next action should reduce a specific uncertainty about the failure.

Create account

Reference: Container probes · CKAD Kubernetes v1.35