Concept and mechanism
Before choosing an algorithm, inspect types, counts, missing values, and distributions. PySpark summary includes statistics and approximate percentiles; it is not a business-validity test. An extreme value may be an error or an important operational event. With a first quartile of 20 and a third quartile of 36, the interquartile range is 16; the exploratory upper rule Q3 + 1.5 × IQR gives 60. Investigate values above that boundary before deleting them. A mean is sensitive to extreme values; a median may be an alternative for a skewed numeric variable. For categories, a mode or an explicit missing category may be more appropriate, depending on the contract and model.
Guided application
Split training and evaluation before learning preparation statistics. Fit imputation, scaling, and feature selection only on the training data within each fold; apply the fitted transformation to validation. A pipeline helps maintain this discipline. One-hot encoding avoids imposing an arbitrary numeric order on application names. Define how new categories are handled: OneHotEncoder defaults to an error, while ignore produces zeros for the unknown feature without automatically learning a category. For a model predicting future days, use validation that preserves temporal direction. If several rows belong to the same incident, also consider separation by group.
A median computed before splitting has already used the held-out set.
Common pitfalls
Deleting every extreme; fitting validation data; category codes treated as quantities.
Related topics: Environment, AutoML, and reproducibility · Temporal features and consistency · MLflow, registry, and promotion
Learn transformations on training data and preserve evaluation independence.
Reference: Prevent inconsistent preprocessing and data leakage · 2025-03-01