Concept and mechanism
An incremental load needs a boundary representing committed work. With a reliable marker advancing on each change, capture the previous value and a new upper bound, then extract the left-open, right-closed interval. Updating the marker before completing writes can skip data after failure. Repeating the interval is also insufficient when the target appends: previously written rows can duplicate. Use staging, upsert, or reconciliation suited to the key and business rule. Physical deletes require another mechanism, such as Change Tracking or controlled comparison; a vanished row supplies no LastModified value to an incremental query.
Guided application
The integration runtime determines where copying can execute and which network it can reach. Self-hosted IR supports private sources reachable from its host; mapping data flows execute on managed Azure compute. For contiguous stateful windows with dependencies, assess tumbling window triggers. Before MERGE, resolve multiple rows matching the same target. Duplicate detection changed between Databricks Runtime15.4 LTS and16.0+, so cases must state conditions and version. Never apply absence-based deletion to a delta as though it were a complete snapshot. Schema drift accepts structural changes, but units, types, and amounts still need validation. Record rejects and reconcile them before publishing results.
240 of300 rows were written. Retain the old marker, replay the interval through upsert, and establish all300 rows before publication.
Common pitfalls
Advanced marker; append during recovery; delta treated as snapshot; success with rejects treated as completeness.
Related topics: Storage, distribution, and exploration · Streams, time, and external effects · Lake access and secret management
A load ends when target and control state satisfy the data contract.
Reference: Incremental loading and change tracking · DP-203 objectives 2024-10-24; retired 2025-03-31