Tech
What Actually Happens When You Call spark.read? One Line of Python, a Thousand Tasks
TL;DR : spark.read.parquet(path) reads almost nothing. The real work starts at the first action , when Spark
turns your DataFrame into four plans, slices your files into tasks, ships those tasks to executors and, depending on
what you asked for, shuffles data across the network. This article follows one DataFrame through all of that, in
five scenarios: a plain read, a read with filters, a sort, a join, and finally the whole pipeline end to end.
⏱️ 30-minute read. Every scenario comes ...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to