Read nested fields
The setup packs status and total into a details struct. A struct holds named fields, potentially of different types. Dot notation reads one nested field; alias controls the flattened output name. Reading details.status does not change the nested source.
Lesson reference: PySpark and Scala
Read nested fields
The setup packs status and total into a details struct. A struct holds named fields, potentially of different types. Dot notation reads one nested field; alias controls the flattened output name. Reading details.status does not change the nested source.
PySpark example
from pyspark.sql.functions import col, lit, struct, array, size, element_at, filter, transform
nested_orders = orders.select(
"order_id", struct("status", "total").alias("details")
)
result = nested_orders.select(col("order_id"), col("details.status").alias("status")).orderBy("order_id")
result.show(truncate=False)
Scala example
import org.apache.spark.sql.functions._
val nested_orders = orders.select(
col("order_id"), struct("status", "total").alias("details")
)
val result = nested_orders.select(col("order_id"), col("details.status").alias("status")).orderBy("order_id")
result.show(false)
Construct a struct
struct groups column expressions into one nested column while keeping one row per order. The field names come from the input columns or their aliases. This example builds details from status and total; adding a field extends the nested schema.
PySpark example
from pyspark.sql.functions import col, lit, struct, array, size, element_at, filter, transform
result = orders.select(
col("order_id"), struct(col("status"), col("total")).alias("details")
).orderBy("order_id")
result.show(truncate=False)
Scala example
import org.apache.spark.sql.functions._
val result = orders.select(
col("order_id"), struct(col("status"), col("total")).alias("details")
).orderBy("order_id")
result.show(false)
Update one field
withField returns a struct with the named field added or replaced, preserving the other fields. withColumn then replaces the top-level details column with that new struct. The example adds 1 to each nested total and keeps status. The original orders DataFrame is unchanged.
PySpark example
from pyspark.sql.functions import col, lit, struct, array, size, element_at, filter, transform
nested_orders = orders.select(
"order_id", struct("status", "total").alias("details")
)
result = nested_orders.withColumn(
"details", col("details").withField("total", col("details.total") + lit(1))
).orderBy("order_id")
result.show(truncate=False)
Scala example
import org.apache.spark.sql.functions._
val nested_orders = orders.select(
col("order_id"), struct("status", "total").alias("details")
)
val result = nested_orders.withColumn(
"details", col("details").withField("total", col("details.total") + lit(1))
).orderBy("order_id")
result.show(false)