← Production Support L3: investigate and recover
09 / 12 · 50 MIN

Linux, JVM, and connectivity diagnosis

Choose observations distinguishing resource shortage, application waiting, and connection failures.

Bytes, inodes, and files still open

On a Linux filesystem, byte capacity and inodes are different resources. Use df -h and df -i on the correct mount to compare both. If bytes remain free but inodes are exhausted, looking only for large files misses the diagnosis. In another case, deleting a log’s name may not release its blocks while a process keeps it open. Entries in /proc/PID/fd help identify descriptors, subject to permissions. Investigate first; application-supported rotation or reopening may suit the situation better than termination. Verify the effect and preserve logs needed for the incident.

Logs for the right process and window

A restart separates executions that may share a service name. With systemd, journalctl can filter by unit, boot, and interval; --utc standardizes displayed time. In an exercise, failure happened during the previous boot while the operator inspects only the current one. Missing lines do not exclude the failure. Check the window, relevant boot, retention, and permissions before concluding evidence does not exist. Correlate PID and execution when available. Converting to UTC does not correct a clock that was wrong; record known skew when reconstructing cross-machine events.

Resource pressure with low CPU

PSI, when available in the kernel, measures time tasks are stalled by CPU, memory, or I/O pressure. In /proc/pressure/io, some includes periods when at least some tasks wait; full covers simultaneous stalls of all non-idle tasks. In a fictional scenario, latency rises, CPU remains low, and I/O pressure increases. That supports investigating waiting but does not independently identify a failed disk or responsible query. Correlate the window with throughput, devices, queues, and dependencies. Avoid treating CPU percentage as a complete capacity measurement.

JVM samples and tool boundaries

A thread waiting at one instant may be operating normally. Compare samples across the incident, stack groups, duration, and pool metrics. HotSpot JDK 25 documentation describes using jcmd to inspect supported commands and obtain Thread.print; collection impact depends on thread count. Confirm JVM, version, PID, and permissions before selecting a tool. Do not automatically apply a HotSpot command to every WebSphere runtime. In an example, three samples show JDBC waiting while the DBA observes blocked transactions: adding threads can expand the queue without freeing the dependency.

TLS: destination, name, and trust

An established TCP connection does not establish TLS identity. For a fictional endpoint, OpenSSL 3.5 supports destination through -connect, SNI through -servername, and expected name through -verify_hostname. The -verify_return_error option fails the operation when verification returns an error; otherwise s_client may continue despite trust problems. Use the appropriate chain and trust store for the test and identify TLS termination. A successful laptop check does not establish that the application has the same trust store. Retain name, destination, time, and result without putting private keys in the ticket.

DNS in the affected context

A Pod’s resolver and search path can differ from those of a laptop or node. If only the application fails resolution, observe the affected context: queried name, namespace, /etc/resolv.conf, response, and DNS-service reachability. Kubernetes documentation describes controlled tests from a diagnostic Pod. Use approved tools and permissions; do not assume exec is always available. In an example, a short name works in one namespace and fails in another. Comparing the full name and search path is more discriminating than immediately changing organization-wide DNS.

IN PRACTICE

In a fictional funds service, CPU is at 20%, requests wait for connections, and I/O pressure has risen. Collect measurements over the same interval and avoid increasing concurrency without understanding the bottleneck.

Common pitfalls

Disk as bytes alone; low CPU as no saturation; one dump as proof of deadlock; handshake as validated identity.

Related topics: Triage and organize the response · Investigate hypotheses and dependencies · Recover batch chains without falsifying success

Take this idea with you

Define a hypothesis, a discriminating observation, and the conclusion’s limit.

Create account

Reference: Pressure Stall Information · DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks