← WebSphere Administration
07 / 9 · 60 MIN

Diagnosis through a timeline

Connect CPU counters, thread identity, dependencies, and impact before choosing mitigation.

Start with interval and identity

In a fictional operations portal, latency rises during an import. Before collecting more data, record the impact interval, affected members, and request type. Tie each file to host, profile, PID, and process instance. Two dumps with the same server name do not necessarily share an origin; a restart can reuse thread names and even identifiers. Keep times and time zones consistent and mark samples that are not simultaneous. The first question is whether the data can be compared. Only then calculate differences. This discipline avoids a convincing explanation based on counters belonging to different process lifetimes.

Use differences rather than old totals

Accumulated CPU describes work since the counter started. For a continuous thread, subtract the initial value from the final value and divide by elapsed time. In the example, 14.4 minus 12.0 gives 2.4 CPU seconds; across four seconds, that represents 60% of one logical core. State that normalization, since it does not equal 60% of the whole machine. A process with multiple threads can consume eight CPU seconds during five wall-clock seconds by using cores in parallel. If observed thread totals do not match the process, check coverage and alignment. Do not automatically assign the remainder to a leak, GC, or an unidentified thread.

A stack locates work without proving cause

Three snapshots show parseDocument in the same thread with a large CPU increment. That justifies investigating the path but does not prove an infinite loop: a large document can require legitimate work throughout the interval. Seek request progress, input size, expected duration, and other evidence distinguishing hypotheses. Runnable state alone also measures neither actual usage nor functional progress. Conversely, low CPU does not mean a healthy service: threads may be waiting for a database, network, or another resource. Correlate JDBC reading with backend sessions and timing. If the database confirms a common lock, that observation guides joint investigation.

Collect to answer a hypothesis

Before requesting another collection, write the hypothesis and what would weaken it. For CPU, a comparable series can show whether the same work remains active. For JDBC waiting, the database team can confirm duration, a blocker, or query completion. For a reproducible failure, targeted tracing in a short window can link a request to a component, with volume limits and defined disabling. There is no universal overhead: instrumentation can disturb the system being measured. Identify runtime and maintenance level before choosing commands or options. This lesson uses synthetic tables; it executes no WebSphere commands, collects no dumps, and demonstrates no safe collection window for your environment.

Laboratory with synthetic data

Read the table below without first looking for a solution. Thread import-A has the largest increment in this window; report-B has the largest historical total. Calculate 60% and 2.5% of one core respectively. Then imagine the process instance changes between t0 and t1: neither calculation remains valid merely because the thread name matches. Also change the final counter below the initial value and mark discontinuity rather than taking absolute value or substituting zero. The laboratory checks arithmetic and evidence interpretation. It is not a javacore parser, does not measure a real process, and does not identify which business routine caused the incident.

Communicate mitigation and validate recovery

Removing a slow member can lower global latency while that member’s problem remains intact. Communicate what was observed: reduced impact, remaining capacity validated for current load, and cause still under investigation. Compare volume, errors, and latency per member and request class over equivalent windows. A better aggregate can merely reflect changed population. Define reentry criteria with service operators: functional checks, dependencies, capacity, monitoring, and rollback thresholds. JVM started state is insufficient. Close diagnosis only when evidence supports the conclusion, and preserve the distinction between hypothesis, mitigation, correction, and validation in the incident record.

SYNTHETIC EVIDENCE; not output from a WebSphere instance
process_instance: node-A/server1/start-2026-10-01T09:00Z
elapsed_seconds: 4
thread_identity | cpu_seconds_t0 | cpu_seconds_t1
import-A | 12.0 | 14.4
report-B | 500.0 | 500.1

one_core_percent = (cpu_seconds_t1 - cpu_seconds_t0) / elapsed_seconds * 100
IN PRACTICE

A historical counter of 500 seconds looks alarming but adds only 0.1 seconds. The thread with a lower total is using more CPU now.

Common pitfalls

Avoid sorting only by accumulated CPU, subtracting different process counters, equating a repeated frame with a loop, or declaring recovery from an aggregate alone.

Related topics: JDBC, pools, and transaction outcomes · JVM, logs, and time-based diagnosis · Maintenance, recovery, and RUN handover

Take this idea with you

A good timeline links identity, counter deltas, and functional progress; conclusions must not exceed the observed evidence.

Create account

Reference: J9 javacore and memory diagnosis · BigSavant WebSphere traditional ND 9.0.5; maintenance baseline 9.0.5.29 (2026-09-08); documentation reviewed 2026-09-30

WebSphere® is a registered trademark of International Business Machines Corporation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by IBM. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.