Concept and mechanism
Define what one row represents before transforming. A movement may have several messages, and an account may have several reference versions. A join on an incomplete key multiplies rows; changing INNER to LEFT retains unmatched rows but does not prevent multiplication by several matches. Use the keys and validity period required by the model. In PySpark, transformations return new DataFrames and are evaluated when an action requires results. Keep the returned reference from filter or drop; the previous object retains its plan. Avoid collecting large results to the driver for tasks that can remain distributed. A small preview does not establish behavior at real volume.
Guided application
For a fictional reconciliation file, compare counts before and after each step. try_cast can turn invalid text into NULL for a supported scalar conversion; handle that result instead of treating it as zero in totals. count(*) counts rows, whereas count(column) excludes null values. UNION ALL preserves repetitions; UNION removes identical complete rows without selecting the latest version per key. explode generates rows from arrays and produces no rows for a NULL array; choose the outer variant when that case must be retained. For a frequently queried gold aggregation, a materialized view can reduce read work, with freshness dependent on refresh and possible full recomputation.
A join with two matches per movement doubles the aggregated amount.
Common pitfalls
DISTINCT hiding a wrong join; NULL treated as zero; unassigned transformation; assumed instant refresh.
Related topics: Platform, compute, and data contracts · Incremental ingestion, state, and schema · Jobs, dependencies, and recovery
Validate the grain and reconcile results as well as checking syntax.
Reference: PySpark DataFrames and evaluation · 2026-05-04