ML Engineer MasterClass (October) | 4 seats left

Spark · Structs
AmazonAmazon Analytics
Amazon · Work with Nested Data

Structs

Read, construct, and update named fields inside a struct.

Step 1 of 6 · Learn

Read nested fields

The setup packs status and total into a details struct. A struct holds named fields, potentially of different types. Dot notation reads one nested field; alias controls the flattened output name. Reading details.status does not change the nested source.

Lesson reference: PySpark and Scala

Read nested fields

The setup packs status and total into a details struct. A struct holds named fields, potentially of different types. Dot notation reads one nested field; alias controls the flattened output name. Reading details.status does not change the nested source.

PySpark example

from pyspark.sql.functions import col, lit, struct, array, size, element_at, filter, transform

nested_orders = orders.select(
    "order_id", struct("status", "total").alias("details")
)

result = nested_orders.select(col("order_id"), col("details.status").alias("status")).orderBy("order_id")
result.show(truncate=False)

Scala example

import org.apache.spark.sql.functions._

val nested_orders = orders.select(
  col("order_id"), struct("status", "total").alias("details")
)

val result = nested_orders.select(col("order_id"), col("details.status").alias("status")).orderBy("order_id")
result.show(false)

Construct a struct

struct groups column expressions into one nested column while keeping one row per order. The field names come from the input columns or their aliases. This example builds details from status and total; adding a field extends the nested schema.

PySpark example

from pyspark.sql.functions import col, lit, struct, array, size, element_at, filter, transform


result = orders.select(
    col("order_id"), struct(col("status"), col("total")).alias("details")
).orderBy("order_id")
result.show(truncate=False)

Scala example

import org.apache.spark.sql.functions._


val result = orders.select(
    col("order_id"), struct(col("status"), col("total")).alias("details")
).orderBy("order_id")
result.show(false)

Update one field

withField returns a struct with the named field added or replaced, preserving the other fields. withColumn then replaces the top-level details column with that new struct. The example adds 1 to each nested total and keeps status. The original orders DataFrame is unchanged.

PySpark example

from pyspark.sql.functions import col, lit, struct, array, size, element_at, filter, transform

nested_orders = orders.select(
    "order_id", struct("status", "total").alias("details")
)

result = nested_orders.withColumn(
    "details", col("details").withField("total", col("details.total") + lit(1))
).orderBy("order_id")
result.show(truncate=False)

Scala example

import org.apache.spark.sql.functions._

val nested_orders = orders.select(
  col("order_id"), struct("status", "total").alias("details")
)

val result = nested_orders.withColumn(
    "details", col("details").withField("total", col("details.total") + lit(1))
).orderBy("order_id")
result.show(false)
example.pyPySpark
Named fieldsStruct expressionNested or flat output
Source nested_orders6 rows
order_iddetails
1001{"status":"Delivered","total":89.5}
1002{"status":"Shipped","total":149}
1003{"status":"Cancelled","total":35}
1004{"status":"Delivered","total":219.99}
1005{"status":"Delivered","total":49.99}
1006{"status":"Shipped","total":120}
Read nested fields
Result6 rows
order_idstatus
1001Delivered
1002Shipped
1003Cancelled
1004Delivered
1005Delivered
1006Shipped
Compare the source columns with the transformed result.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.