← CKA: Kubernetes administration and troubleshooting
08 / 8 · 70 MIN

Diagnosis: Service, probes, and binding

Connect ports, conditions, and topology to the client-observed outcome.

Start with the affected contract

A fund API may have every Pod Running and still fail for its client. First define the affected operation, source client, destination, and time window. Separate four questions: did the process start, is the Pod ready, does the Service target the correct destination, and does the business operation finish with the expected result? Record context, namespace, revision, and evidence before changing configuration. During an incident involving financial operations, a positive technical response does not remove the need to check outcomes and duplicates. Observations in this lesson are fictional; commands are proposed reads for an authorized lab, not records of cluster execution.

From the public port to the listener

In the example the client uses 443 while the process listens on 8443. A targetPort of 8080 leaves an incorrect destination even if the Service has an address and the Pod is ready. Compare the client contract, Service configuration, EndpointSlice ports, and observed process listener. The containerPort field creates no listener. During migration a named targetPort of api allows different versions to declare different numbers. Diagnostic tooling must retain each slice-to-port association; selecting only the first number loses information. A request along the client path after correction is required to establish the outcome.

Read conditions without confusing presence and readiness

An IP being present is not a health confirmation. In the ordinary fixture, publishNotReadyAddresses=false and a non-terminating Pod has ready=false: the address may remain represented but is not a ready backend for ordinary selection. Another snapshot shows serving=true, terminating=true, and ready=false during shutdown. This combination does not prove corruption; it separates ability to serve from lifecycle phase. Do not infer every existing connection outcome solely from those conditions. With publishNotReadyAddresses=true, slice ready no longer establishes Pod readiness. Inspect the Service option and consumer semantics before feeding an availability dashboard.

Probes with different purposes

Startup provides initialization time before readiness and liveness execute. Afterwards readiness addresses ability to accept work; liveness addresses situations where restarting the process can help. For an application warming up for 95 seconds, set headroom using representative measurements and retain suitable steady-state detection. Do not treat readiness as an automatic liveness gate. If both query an external database, database failure can cause restarts that do not repair it and repeat expensive work. Review contracts, failure thresholds, and gradual recovery. A proposed configuration gains operational evidence only after startup, overload, and dependency-failure exercises.

Pending PVC: consumer and topology

With WaitForFirstConsumer, a claim without a consumer may legitimately wait. Once a consumer exists, read Pod, PVC, and StorageClass configuration and events together. nodeName bypasses the scheduler; nodeSelector preserves placement constraints without skipping its decision. In the proposed case replacing one with the other is necessary but insufficient: the requested zone still conflicts with permitted storage. Compare the sets and justify an approved requirements change. Binding does not establish mounting or read/write access. Do not delete the claim to remove the symptom before knowing whether it references required data; plan controlled consumer recreation when its design change requires it.

A local check has limited scope

A port-forward request ending in HTTP 200 shows that the particular request reached the application through the tunnel. It does not establish that the usual client reaches the ClusterIP or that the complete operation works. Record separately: name resolution, a PodIP request on the expected port, a Service request, and an operation through the representative client. Each step narrows hypotheses, but success on one path does not replace the next. If a direct ClusterIP request still times out, do not automatically attribute the cause to DNS. Compare network policy, endpoint selection, ports, and routing with authorized access and preserved evidence.

Guided recovery exercise

Consider six ready Pods: A1–A3 listen on 8080; B1–B3 on 9090; all correctly declare api. Numeric destination 8080 works only for the first group. In a trial forcing ten requests to each Pod, 30 of 60 work. This denominator is defined by the exercise and does not describe any load balancer’s actual distribution. Explain the hypothesis, change request to targetPort=api, and required subsequent evidence for both groups. Also define a rollback alternative, minimum capacity, and reconciliation. The diagnosis sheet should separate observed facts, inferences, and outstanding criteria instead of recording only “Kubernetes fixed”.

Summary and operational handover

Close the exercise with five evidence points: correct target and context, observed port mapping, interpreted endpoint conditions, available dependencies, and the representative operation outcome. For storage add binding, mounting, and data validation. Record approved changes, return criteria, regression signals, and the follow-up owner. Supplied local models check only arithmetic and explicit fixture rules. They execute no kubelet, scheduler, CSI, or proxies, measure no timings, and do not establish cluster availability. Connect this lesson with services, storage, workload diagnosis, and incident management before proceeding to supervised real practice.

kubectl --context=lab -n funds get service fund-api -o yaml
kubectl --context=lab -n funds get endpointslices -l kubernetes.io/service-name=fund-api -o yaml
kubectl --context=lab -n funds get pods -l app=fund-api -o wide
kubectl --context=lab -n funds describe pod recon-b
kubectl --context=lab -n funds describe pvc recon-data
kubectl --context=lab get storageclass funds-zonal -o yaml
IN PRACTICE

Six Pods receive ten requests each in a controlled trial. Three listen on 8080 and three on 9090; with destination 8080, 30 of 60 requests work. Changing to a shared port name requires confirmation across both groups.

Common pitfalls

Confusing Running with service recovery; using the first port across all slices; deleting claims without diagnosis; inferring real traffic from the fixture.

Related topics: Services and networking · Volumes and data · Incidents and recovery

Take this idea with you

Diagnosis concludes when the affected operation works again with integrity, evidence, and stability.

Create account

Reference: Diagnose Service connectivity · CKA Kubernetes v1.35

Kubernetes® and CKA are trademarks or registered trademarks of The Linux Foundation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by The Linux Foundation. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.