Concept and mechanism
An architecture with two instances may appear resilient while both depend on the same unavailable component. A technical analyst looks for failure mechanisms, activating conditions, and service consequences. In a fictional project, a batch must recover after interruption without losing or duplicating instructions. The risk is not adequately described as server down: it includes partially stored state, lost acknowledgment, and repeated operations. Design the investigation with people who understand application, infrastructure, and operations. Distinguish product risk, such as duplicate results, from project risk, such as the environment being unavailable on time. The second situation can prevent evidence about the first and should appear in the plan.
Guided application
For each investigation, define expected behavior, environment, data, observation, and limits. A requirement such as fast does not establish an oracle; percentile, load, duration, and error rate need context. For recovery, agree what restored service means and how data integrity will be checked. Use authorized environments and artificial data in exercises. Before a failure experiment, identify who can stop it, how conditions will be restored, and which dependencies must remain protected. The final decision should consider evidence, representativeness, and residual risk. A successful experiment in a reduced configuration does not automatically demonstrate the same production outcome, especially when concurrency, volume, and dependencies differ.
Two instances with the same dependency may share a failure mode.
Common pitfalls
Redundancy treated as proof; unmeasured requirement; reduced environment treated as equivalent.
Related topics: White-box logical coverage · Static and dynamic analysis · Performance and workload profile
Connect risk, experiment, and operational decision.
Reference: Google SRE testing for reliability · CTAL-TTA v4.0 (2021)