Start with interval and identity
In a fictional operations portal, latency rises during an import. Before collecting more data, record the impact interval, affected members, and request type. Tie each file to host, profile, PID, and process instance. Two dumps with the same server name do not necessarily share an origin; a restart can reuse thread names and even identifiers. Keep times and time zones consistent and mark samples that are not simultaneous. The first question is whether the data can be compared. Only then calculate differences. This discipline avoids a convincing explanation based on counters belonging to different process lifetimes.
Use differences rather than old totals
Accumulated CPU describes work since the counter started. For a continuous thread, subtract the initial value from the final value and divide by elapsed time. In the example, 14.4 minus 12.0 gives 2.4 CPU seconds; across four seconds, that represents 60% of one logical core. State that normalization, since it does not equal 60% of the whole machine. A process with multiple threads can consume eight CPU seconds during five wall-clock seconds by using cores in parallel. If observed thread totals do not match the process, check coverage and alignment. Do not automatically assign the remainder to a leak, GC, or an unidentified thread.
A stack locates work without proving cause
Three snapshots show parseDocument in the same thread with a large CPU increment. That justifies investigating the path but does not prove an infinite loop: a large document can require legitimate work throughout the interval. Seek request progress, input size, expected duration, and other evidence distinguishing hypotheses. Runnable state alone also measures neither actual usage nor functional progress. Conversely, low CPU does not mean a healthy service: threads may be waiting for a database, network, or another resource. Correlate JDBC reading with backend sessions and timing. If the database confirms a common lock, that observation guides joint investigation.
Collect to answer a hypothesis
Before requesting another collection, write the hypothesis and what would weaken it. For CPU, a comparable series can show whether the same work remains active. For JDBC waiting, the database team can confirm duration, a blocker, or query completion. For a reproducible failure, targeted tracing in a short window can link a request to a component, with volume limits and defined disabling. There is no universal overhead: instrumentation can disturb the system being measured. Identify runtime and maintenance level before choosing commands or options. This lesson uses synthetic tables; it executes no WebSphere commands, collects no dumps, and demonstrates no safe collection window for your environment.
Laboratory with synthetic data
Read the table below without first looking for a solution. Thread import-A has the largest increment in this window; report-B has the largest historical total. Calculate 60% and 2.5% of one core respectively. Then imagine the process instance changes between t0 and t1: neither calculation remains valid merely because the thread name matches. Also change the final counter below the initial value and mark discontinuity rather than taking absolute value or substituting zero. The laboratory checks arithmetic and evidence interpretation. It is not a javacore parser, does not measure a real process, and does not identify which business routine caused the incident.
Communicate mitigation and validate recovery
Removing a slow member can lower global latency while that member’s problem remains intact. Communicate what was observed: reduced impact, remaining capacity validated for current load, and cause still under investigation. Compare volume, errors, and latency per member and request class over equivalent windows. A better aggregate can merely reflect changed population. Define reentry criteria with service operators: functional checks, dependencies, capacity, monitoring, and rollback thresholds. JVM started state is insufficient. Close diagnosis only when evidence supports the conclusion, and preserve the distinction between hypothesis, mitigation, correction, and validation in the incident record.
SYNTHETIC EVIDENCE; not output from a WebSphere instance
process_instance: node-A/server1/start-2026-10-01T09:00Z
elapsed_seconds: 4
thread_identity | cpu_seconds_t0 | cpu_seconds_t1
import-A | 12.0 | 14.4
report-B | 500.0 | 500.1
one_core_percent = (cpu_seconds_t1 - cpu_seconds_t0) / elapsed_seconds * 100
A historical counter of 500 seconds looks alarming but adds only 0.1 seconds. The thread with a lower total is using more CPU now.
Common pitfalls
Avoid sorting only by accumulated CPU, subtracting different process counters, equating a repeated frame with a loop, or declaring recovery from an aggregate alone.
Related topics: JDBC, pools, and transaction outcomes · JVM, logs, and time-based diagnosis · Maintenance, recovery, and RUN handover
A good timeline links identity, counter deltas, and functional progress; conclusions must not exceed the observed evidence.
Reference: J9 javacore and memory diagnosis · BigSavant WebSphere traditional ND 9.0.5; maintenance baseline 9.0.5.29 (2026-09-08); documentation reviewed 2026-09-30