Concept and mechanism
Problem management seeks to understand and reduce factors that cause incidents or could cause them. Urgent service recovery remains an incident-response responsibility. A restart that allows a batch to finish may be useful without explaining the failure mechanism. Preserve that distinction in the record: restored service, cause still under investigation, and available workaround describe different outcomes. There is no requirement to wait for three incidents before investigating. A capacity trend, near miss, or architecture review may reveal sufficient exposure to act. Counting tickets is also insufficient: several timeout messages may result from independent failures despite sharing the same text.
Guided application
In a fictional banking-production example, compare a rare failure that could prevent closing with frequent low-impact interruptions. Consider consequence, exposure, recurrence, effectiveness of temporary measures, and investigation effort. The problem owner coordinates specialists and next steps; the process owner oversees the practice as a whole; the service owner participates in outcome and risk decisions. Roles may be combined when responsibility remains explicit. Link related incidents without deleting individual timelines and impacts. Describe the problem using expected and observed behavior, scope, versions, and available evidence. Writing slow database as the cause before measuring the dependency prematurely constrains the investigation.
Three similar timeouts may involve different DNS, capacity, or contention mechanisms.
Common pitfalls
Count as priority; third incident as a prerequisite; recovery as cause removal.
Related topics: Investigation and evidence · Workarounds and known errors · Correction and accepted risk
Prioritize exposure and consequences and keep an owned investigation.
Reference: Problem investigation and contributing causes · Problem management practices 2026-09; ServiceNow Brazil examples with scoped plugins and properties