ML Engineer MasterClass (October) | 4 seats left

Spark · Read an execution plan
AmazonAmazon Analytics
Amazon · How Spark Executes

Read an execution plan

Inspect live Spark plans without treating node numbers or optimizer choices as fixed.

Step 1 of 6 · Learn

Follow a physical plan

explain("formatted") prints a physical-plan outline and operator details. Read scans at the leaves, then follow Filter, Project, and other parent nodes toward the output. Node numbers are references, not task counts. explain itself does not run the query; this studio separately collects result rows to check the exercise. Plan text appears below the result table.

Lesson reference: PySpark and Scala

Follow a physical plan

explain("formatted") prints a physical-plan outline and operator details. Read scans at the leaves, then follow Filter, Project, and other parent nodes toward the output. Node numbers are references, not task counts. explain itself does not run the query; this studio separately collects result rows to check the exercise. Plan text appears below the result table.

PySpark example

from pyspark.sql.functions import col, sum, count, lit

result = orders.filter(col("total") >= 100).select("order_id", "total").orderBy("order_id")
result.explain("formatted")

Scala example

import org.apache.spark.sql.functions._

val result = orders.filter(col("total") >= 100).select("order_id", "total").orderBy("order_id")
result.explain("formatted")

Find aggregation and exchange

A grouped sum often has partial and final aggregation with an Exchange between them. The exchange brings matching grouping keys together. AdaptiveSparkPlan means Spark may revise physical choices using runtime information. Do not assume its initial outline is the final executed plan, or interpret each operator as a separate stage.

PySpark example

from pyspark.sql.functions import col, sum, count, lit

result = orders.filter(col("total") >= 100).groupBy("status").agg(
    sum("total").alias("total_value")
).orderBy("status")
result.explain("formatted")

Scala example

import org.apache.spark.sql.functions._

val result = orders.filter(col("total") >= 100).groupBy("status").agg(
    sum("total").alias("total_value")
).orderBy("status")
result.explain("formatted")

Compare logical and physical plans

explain("extended") adds parsed, analyzed, and optimized logical plans before the physical plan. Logical plans describe the requested transformations; the physical plan selects execution operators. Optimizer rewrites can reorder or combine operations while preserving results. Compare the filter and selected columns rather than memorizing expression IDs.

PySpark example

from pyspark.sql.functions import col, sum, count, lit

result = orders.filter(col("total") >= 100).select("order_id", "total").orderBy("order_id")
result.explain("formatted")

Scala example

import org.apache.spark.sql.functions._

val result = orders.filter(col("total") >= 100).select("order_id", "total").orderBy("order_id")
result.explain("formatted")
example.pyPySpark
1. Read from the leavesStart at the scan, then follow its parent operators toward the result.
2. Separate local work from movementFilter and Project process rows; Exchange redistributes them.
3. Inspect the current planNode IDs and adaptive choices can vary. The diagram is a guide, not a captured plan.
Source orders6 rows
order_idcustomer_idstatustotalitem_count
1001101Delivered89.52
1002102Shipped1493
1003101Cancelled351
1004103Delivered219.994
1005104Delivered49.991
1006105Shipped1202
Follow a physical plan
Result3 rows
order_idtotal
1002149
1004219.99
1006120
Look for the filter and projected columns. The final orderBy can add Sort and Exchange nodes.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.