← AWS Machine Learning Engineer: prepare and operate models
02 / 8 · 40 MIN

Features, splits, and leakage

Maintain training-serving consistency and evaluate genuinely independent data.

Concept and mechanism

A feature should represent information available at prediction time. Querying a customer’s current value to reconstruct an old decision can introduce future knowledge. Feature Store provides online storage for recent low-latency values and offline storage for history and training; selecting the correct historical state remains a pipeline responsibility. Keep transformations consistent: units, categories, missing-value handling, and normalization should follow the version used by the model. Fit transformation parameters on training data so testing does not influence preparation. Separate training, validation for selection, and a reserved test set for final evaluation. Split percentages are contextual choices; independence and representativeness matter more than repeating a familiar ratio.

Guided application

In a support case, one incident generates five nearly identical tickets. A random row split spreads copies across training and testing and creates apparent generalization. Group related examples and address duplication before interpreting results. For temporal prediction, respect ordering and check that evaluation represents the usage period. A rare class may need suitable sampling or weighting without changing the test to hide actual prevalence. Compare performance across relevant segments. In a review calculation, 9,000 records minus 600 duplicates leave 8,400 distinct examples; reserving20% of those gives1,680 rather than1,800. This calculation is valid only after confirming the split unit and grouping constraints.

IN PRACTICE

A row-level split can leave the same incident on both sides of evaluation.

Common pitfalls

Future values in historical training; divergent transforms; testing used for tuning; duplicates treated as new examples.

Related topics: Ingestion, storage, and quality · Training, tuning, and reproducible experiments · Foundation models and RAG quality

Take this idea with you

Evaluate generalization with time, entity, and data boundaries preserved.

Create account

Reference: AWS MLOps splits and data leakage · MLA-C02