Diagnosing skew, shuffle, and spill from Spark UI stage metrics
Spark UI stage-level metrics reveal where jobs slow down: task duration skew, shuffle read/write volume, and spill to disk all show up in the Stages tab's task summary and event timeline. Data skew appears as a small number of tasks taking far longer than the median, shuffling shows up as large shuffle read/write bytes, and spilling shows as memory spill and disk spill columns indicating a partition didn't fit in executor memory.
1 · Learn the must-know
- Data skew is diagnosed by comparing min/median/max task duration in the Stages tab; a max far above the median (long-tail tasks) indicates one or a few partitions hold disproportionate data.
- Shuffle read/write size and shuffle spill (memory) and shuffle spill (disk) are visible per-stage summary metrics; wide transformations (groupBy, join, distinct, repartition) trigger shuffles and are the usual root cause of skew and spill.
- Disk spilling occurs when a task's data exceeds available executor memory during a shuffle or aggregation and Spark writes intermediate data to disk, which sharply increases stage duration and I/O.
- Salting skewed join/group-by keys, using Adaptive Query Execution (AQE) to auto-optimize skewed joins and coalesce shuffle partitions, or increasing shuffle partition count are common fixes for skew and spill.
- The Spark UI SQL tab and stage DAG visualization help pinpoint which operator (join, aggregate, sort) caused the shuffle, complementing the raw stage metrics.
- A stage with many small tasks (over-partitioning) or very few large tasks (under-partitioning) both show up as inefficiency in the Stages tab and often correlate with skew or spill symptoms.
2 · Check your understanding
In the Spark UI, a stage shows 200 tasks total. The event timeline shows 199 tasks finish in under 10 seconds each, while 1 task runs for 12 minutes. The Summary Metrics table shows a huge gap between the 75th percentile and max for Duration and Shuffle Read Size. What does this pattern most strongly indicate?
What you have tried across Databricks DEA's objectives, not a readiness score.
Databricks Intelligence Platform6% of the exam0 of 2 tried
Data Ingestion and Loading21% of the exam0 of 7 tried
Data Transformation and Modeling22% of the exam0 of 7 tried
Working with Lakeflow Jobs16% of the exam0 of 4 tried
Implementing CI/CD10% of the exam0 of 4 tried
Troubleshooting, Monitoring, and Optimization10% of the exam0 of 5 tried
Governance and Security15% of the exam0 of 4 tried
3 · Keep going
Ready for more? Take a weighted mock or try free practice questions.