Concept and mechanism
Reliability should be discussed through user experience and business consequences. A service indicator measures an observed property; an objective defines the intended level over a period. An error-budget policy helps decide when to emphasize stability or evolution, but must be agreed in the service context. The manager supports decisions with appropriate evidence and people without inventing a universal rule to stop every change. During incidents, respect defined coordination roles and communication channels. Taking technical command solely because one is the manager can duplicate instructions and distract those restoring the service.
Guided application
After a fictional batch incident, the first reaction is to blame whoever ran the last command. A useful postmortem reconstructs what was known, the sequence, contributing conditions, and missing protections. Blameless learning does not remove facts, follow-up ownership, or the need to improve. Agree concrete actions, owners, and completion criteria, prioritized by expected effect. Check whether actions actually change the ability to detect, recover, or prevent recurrence. “Be more careful” does not describe a verifiable system change. Protect the ability to report doubts and early signals; if bad news is punished, management may receive reassuring reports while risk grows.
An action should change a condition and allow its result to be checked.
Common pitfalls
Manager as parallel command; blame treated as a complete cause; ownerless actions; invented reliability rules.
Related topics: Mandate, autonomy, and delegation · Feedback and skill development · Capacity, operational load, and toil
Support coordinated recovery and turn learning into verifiable changes.
Reference: Postmortem culture learning from failure · GitLab Handbook 2026; DORA current five-metric model; SRE and engineering guidance reviewed 2026-09-30