TL;DR : spark.read.parquet(path) reads almost nothing. The real work starts at the first action , when Spark turns your DataFrame into four plans, slices your files into tasks, ships those tasks to executors and, depending on what you asked for, shuffles data across the network. This article follows one DataFrame through all of that, in five scenarios: a plain read, a read with filters, a sort, a join, and finally the whole pipeline end to end. ⏱️ 30-minute read. Every scenario comes ...