From suspicion to a testable hypothesis
The exercise hypothesis is that failures followed by success can justify investigating an account. Define population, sources, time window, analysis unit and alternative explanations before running the search. The chosen rule looks for at least three distinct failure events for the same tenant and user within the inclusive 300 seconds before success. It does not prove credential theft: a legitimate user can forget a password. An attacker can also use valid credentials without earlier failures. Investigation needs additional authorized context, such as device, later actions and privilege changes. Keep conclusions bounded by the hypothesis and available data; a negative result does not exclude behavior outside the predicate.
The denominator is part of the evidence
The lab has six cases with known synthetic classifications. Alpha, gamma and late represent malicious activity in the scenario; beta, delta and shared in funds-b represent benign activity. This truth was defined for the exercise, not inferred from the rule itself. After full replay, alpha and late alert: two true positives. Beta also alerts: one false positive. Gamma does not alert because it has only two failures: one false negative. Delta and shared correctly do not alert: two true negatives. Time-boundary, clock and missing-telemetry cases support other tests and do not enter this matrix. Mixing unlabeled cases into the denominator would create unsupported apparent precision. In production, document how classifications were obtained and which cases remain unknown.
Precision and recall answer different questions
Precision is TP/(TP+FP): in this set, two of three labeled alerts correspond to the defined malicious activity. Recall is TP/(TP+FN): two of three known malicious cases were found. Both equal 2/3, but that coincidence does not make the metrics equivalent. With no alerts, precision has a zero denominator; the script returns null rather than claiming 100%. With no labeled positive cases, recall is also undefined. Adding many benign cases without alerts can increase accuracy without finding the missed attack. Therefore retain absolute counts, sample composition and analysis period. These six teaching cases support neither a statistical forecast of real effectiveness nor a forecast of SOC staffing capacity.
Tuning a rule changes the trade-off
Lowering the threshold from three to two finds gamma in this set. That does not establish that the new threshold is better in every population; other accounts may create more noise. In a second trial, the team suppresses alpha and beta to reduce volume. The only remaining labeled alert is late: precision rises to 1 while recall falls to 1/3. Dashboard improvement concealed removal of a malicious case. An operational exception needs a reason, scope, owner, expiry and regression test. Before promoting a change, compare positive, negative and boundary cases as well as execution cost and analysis delay. A rule that detects correctly during replay may arrive too late for the intended operational decision.
Turn investigation into a sustainable control
The search result should let another person reproduce the decision: retain query version, parameters, sources, period, selected clock, labeled cases and limitations. For handover, assign a rule owner and triage procedure. If a source stops reporting, coverage should visibly degrade. If replay finds the same account again, define whether it updates an investigation or opens a new alert; DISTINCT over accounts in the script does not solve incident lifecycle management. A fictional payments project should accept the control only after proving the path in the authorized environment and agreeing response, access and retention. The lab runs real SQL over invented data without a SIEM or actual attack, preparing the learner for those trials.
# Six explicitly labeled synthetic cases only:
TP, FP, FN, TN = 2, 1, 1, 2
precision = TP / (TP + FP) if TP + FP else None
recall = TP / (TP + FN) if TP + FN else None
# Suppressing alpha and beta: TP=1, FP=0, FN=2, TN=3.
# More precise alerts can coexist with more missed malicious cases.
Suppressing two accounts improves precision from 2/3 to 1 but reduces recall to 1/3 because one account represented a malicious case.
Common pitfalls
Using the rule to create its own labels; excluding false negatives; celebrating fewer alerts without coverage; confusing replay with timely response.
Related topics: Telemetry quality · Governance and residual risk
A useful change improves decisions with explicit evidence about what it finds, misses and costs.
Reference: SecurityX CAS-005 objectives · CAS-005 / SecurityX V5; objectives 3.0; launched 2024-12-17