Concept and mechanism
Start with run history and separate waiting, startup, execution, and dependencies. A skipped task may result from an earlier failure. On compute offering Spark UI, follow the job to its longest stage and inspect duration distribution, shuffle, and spill. Spill indicates execution-memory pressure with movement to disk; it does not necessarily mean insufficient free disk space. Uneven work across tasks can indicate skew, whereas a uniform regression may have another cause. Import errors before reading point to libraries or the environment rather than immediately to data size. Compare the first useful error with recent changes and preserve context.
Guided application
In an exercise, selecting fewer columns before a join reduces shuffle without changing the result. Confirm counts and business measures and repeat with comparable input before accepting the improvement. Liquid clustering organizes data around access patterns and replaces partitioning and ZORDER for that table; check runtime and client compatibility. Predictive optimization automates maintenance for Unity Catalog managed tables using billed serverless compute. It does not apply in the same way to external tables. Set retention before enabling maintenance that includes VACUUM, considering recovery and time travel. Monitor cost per useful delivery and time to the consumer; a cheaper run that misses its deadline may fail the objective.
Less time is an improvement only if output remains complete and correct.
Common pitfalls
Scaling without diagnosis; deleting data to speed up; comparing different volumes; assuming free maintenance.
Related topics: Platform, compute, and data contracts · Incremental ingestion, state, and schema · Transformation, grain, and quality
Change one hypothesis at a time and confirm performance, cost, and correctness.
Reference: Diagnose performance with Spark UI · 2026-05-04