Concept and mechanism
Lakeflow Jobs organizes tasks in a dependency graph. A task can execute a notebook, SQL, or a pipeline; conditions support branching, and for each repeats parameterized work. Design dependencies around data and publication criteria. Two tasks scheduled close together do not form a robust dependency. A scheduled trigger responds to the clock; file arrival responds to arriving files; table update detects table changes. Selecting any or all for several tables controls the triggering event but does not establish that data belongs to the same business date. Add explicit period, completeness, and quality checks before releasing results.
Guided application
In a fictional incident, ingestion completes, transformation writes partial output and fails, and publication is blocked. Investigate the first failure and persisted effects before repairing. Repair can rerun unsuccessful tasks and dependants using current configuration, but restarts the task and does not make an append idempotent. Define suitable keys, partitions, or merge behavior for duplicate-free recovery. Confirm that a code change between the first attempt and repair remains compatible with existing input. Notifications should identify the job, run, task, impact, owner, and next action. APS handover includes criteria for stopping retries and escalating while retaining evidence for later investigation.
Repairing a job does not erase effects from its previous attempt.
Common pitfalls
Scheduling treated as dependency; all tables treated as aligned dates; retry treated as idempotency.
Related topics: Platform, compute, and data contracts · Incremental ingestion, state, and schema · Transformation, grain, and quality
Tie publication to data criteria and reconcile effects before retrying.
Reference: Lakeflow Jobs orchestration · 2026-05-04