← Production Support L3: investigate and recover
02 / 12 · 20 MIN

Investigate hypotheses and dependencies

Build a discriminating diagnosis across Linux, JVM, networking, and databases.

Symptom, hypothesis, and observation

A symptom is what was observed; a hypothesis is a possible explanation. “Requests take 30 seconds” is a symptom. “They wait for a JDBC connection” is a hypothesis to compare with pool metrics and stacks. Choose observations that distinguish competing explanations. When evidence contradicts a hypothesis, update it instead of repeating the same action.

Correlate layers over the same interval

Connect the business operation to requests, processes, and dependencies. Compare errors and latency with CPU, memory, I/O, pools, external calls, and database activity over the same period. Low CPU does not rule out saturation elsewhere. A full pool does not prove its limit is too small. A timeout may occur at different stages. Retain the dependency map and the origin of each measure.

Collect and escalate with a useful package

When involving a specialist team, provide onset, impact, affected operation, recent changes, identifiers, and observations already made. Distinguish lack of evidence from evidence of absence. Database blocking and slow queries require DBA or appropriately authorized analysis; terminating a session can cause lengthy rollback and business effects. Do not replace analysis with an irreversible action merely because it is available.

Workplace application

For a slow export, compare an affected run with a reference of similar volume. Align application logs, dependency timings, and recent changes. A useful hypothesis predicts an observation that could disprove it. If small requests work and large ones fail, preserve that distinction in escalation rather than reporting only that the whole application is slow.

IN PRACTICE

WebSphere waits for connections while the database shows blocking behind an old transaction. The escalation package includes time, correlation IDs, relevant sessions obtained through authorized means, and the preceding change. The team chooses mitigation with the transaction risk understood.

Common pitfalls

Confusing symptoms with cause; collecting without a hypothesis or context.

Related topics: Recover batch chains without falsifying success · Validate transfers and reconciliation

Take this idea with you

Every diagnostic step should reduce uncertainty and preserve context for the next team.

Create account

Reference: Effective troubleshooting · DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks