← Kubernetes: operate workloads and recover services
09 / 12 · 70 MIN

Diagnosing Services and EndpointSlices

Locate the failing layer by comparing DNS, selectors, endpoints, ports and HTTP requests in the synthetic application.

Start from a working reference

The lab runs a synthetic API listening on 8080 and returning identifier dr-synthetic-ledger. A client Pod makes HTTP requests with a timeout; both belong to a disposable namespace. Before injecting faults, the script confirms the response through the Pod IP, its Ready condition and restart count. This baseline distinguishes an introduced fault from an environment that never worked. Record the effective server and client versions as well. The experiment uses one local node: its result does not establish availability across nodes, external load-balancer behavior or capacity for a banking service’s workload. Those require their own representative observations.

Correct DNS with the wrong selector

Service ledger starts with selector app=wrong while the Deployment creates Pods labeled app=ledger. Name ledger resolves to the Service’s assigned ClusterIP, but the destination list is empty and the request fails. The API Pod remains ready. These facts are compatible: a normal Service’s DNS record does not require a working backend. Compare the selector with Pod labels in the same namespace. Avoid changing CoreDNS merely because the user describes the symptom as a name problem. In this experiment, correcting the selector is followed by a ready destination and a successful HTTP request through the Service.

Observe convergence

An API-accepted change and a configuration already affecting traffic are different observations. The script waits for a ready EndpointSlice destination and then retries the HTTP request within a bound. The controller and traffic routing need not change at the instant the patch command returns. Define an operational deadline and observe transitions without declaring failure at the first intermediate state. In production, preserve change time, relevant state and the consumer-visible result. If the deadline expires, investigate where convergence stopped. Avoid extending a wait indefinitely when it conceals lack of progress or postpones an agreed recovery decision.

Ready destination with the wrong port

The second fault keeps the correct selector and changes targetPort to 9090. The EndpointSlice still reports a ready destination, now with that port. The application continues listening on 8080; a direct request to IP:8080 works and the Service request fails. Ready state comes from the configured probe, not from testing every port published by the Service. Compare the actual listener, Service port and resolved destination port. The experiment recovers when targetPort returns to name http, associated with containerPort 8080. Declaring containerPort does not itself start a listener; the Python process opens it in this exercise.

Compare paths without overstating the result

A successful direct Pod request reduces the likelihood that the application is entirely unavailable at that instant. It does not establish that every consumer can reach the Service, nor test Ingress, TLS, authentication or a business transaction. This experiment’s client and API share one local cluster and use plain HTTP. When transferring the method to an incident, identify request origin, protocol, name, port and expected result. Two tests from different origins may encounter different policies or routes. Preserve these conditions in the incident record so the comparison remains interpretable and repeatable rather than becoming an unsupported claim about overall service health.

Hand useful evidence to the next team

In a fictional funds-processing support case, the functional owner needs to know whether the batch consumer’s channel works again. A useful summary distinguishes resolved name, selected destinations, observed port and application response. Add the change made, its observed result and outstanding validation. The lab removes resources it created in its namespace; it does not operate on personal contexts or external environments. This discipline makes diagnosis reproducible and bounds impact. The synthetic request is a small technical criterion that should be complemented by a representative functional operation before closing a real incident or accepting a production handover.

# Use an explicitly authorized context and namespace.
kubectl --context OWNED -n LAB get service ledger -o yaml
kubectl --context OWNED -n LAB get pods -l app=ledger --show-labels
kubectl --context OWNED -n LAB get endpointslices \
 -l kubernetes.io/service-name=ledger -o yaml
# Compare Service port, resolved target port, Pod readiness and a bounded request.
IN PRACTICE

DNS returned the ClusterIP although the wrong selector produced zero destinations; correcting the selector restored the Service request.

Common pitfalls

Resolving a name as proof of health; a ready endpoint as proof of the correct port; direct Pod access as validation of the entire path.

Related topics: Traffic and probes · Operations and recovery

Take this idea with you

A controlled comparison between paths reduces hypotheses; recovery requires repeating the request through the consumer’s actual path.

Create account

Reference: Debug Services · Kubernetes v1.37 concepts; current official documentation consulted 2026-09-30; cluster versions and plugin capabilities must be confirmed

Kubernetes® is a registered trademark of The Linux Foundation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by The Linux Foundation. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.