Concept and mechanism
The metric should answer the intended usage decision. In a system flagging jobs for review, precision measures how many alerts genuinely match the event, while recall measures how many events were found. With 18 true positives, six false positives, and 12 false negatives, precision is 18/24 = 75% and recall is 18/30 = 60%. F1 combines precision and recall harmonically; it is not the overall percentage of correct predictions. High accuracy can hide a rare class that is never detected. Analyze thresholds, review volume, and error consequences. ROC AUC describes ranking across thresholds but does not itself define an operational threshold or establish calibrated probabilities.
Guided application
For classification probabilities, log loss penalizes confident wrong predictions. For duration regression, MAE uses target units and RMSE gives more weight to large errors. With absolute errors of 2, 4, and 9 minutes, MAE is five minutes. Also compare against a baseline and relevant operational groups; R² can be negative when a model performs worse than predicting the reference mean. If training used log1p of the target, interpret results on the original scale using the appropriate inverse transformation. Small training error and large validation error suggest investigating overfitting or a data change; they do not prove one unique cause.
Choosing the highest accuracy may increase missed critical events.
Common pitfalls
Global average treated as an SLA; probability treated as class; logarithmic error presented in minutes.
Related topics: Environment, AutoML, and reproducibility · Temporal features and consistency · MLflow, registry, and promotion
Show the denominator, scale, and error cost before promotion.
Reference: Model evaluation metrics · 2025-03-01