Classify each line
when checks each order line; otherwise supplies the fallback. The example labels line_total of at least 100 as High. line_total includes every unit on that line, not the entire order or a single unit. These labels are rules for this exercise.
Lesson reference: PySpark and Scala
Classify each line
when checks each order line; otherwise supplies the fallback. The example labels line_total of at least 100 as High. line_total includes every unit on that line, not the entire order or a single unit. These labels are rules for this exercise.
PySpark example
from pyspark.sql.functions import col, when
result = order_items.withColumn(
"line_band", when(col("line_total") >= 100, "High").otherwise("Standard")
).select("order_item_id", "line_band").orderBy("order_item_id")
result.show()
Scala example
import org.apache.spark.sql.functions._
val result = order_items.withColumn(
"line_band", when(col("line_total") >= 100, "High").otherwise("Standard")
).select("order_item_id", "line_band").orderBy("order_item_id")
result.show(false)
Order your conditions
Chained when uses the first matching branch. Quantity at least three is Multi-unit; among other lines, totals at least 120 are Review and the rest Standard. Line 6 matches both conditions, so the first branch determines its label.
PySpark example
from pyspark.sql.functions import col, when
result = order_items.withColumn(
"packing", when(col("quantity") >= 3, "Multi-unit")
.when(col("line_total") >= 120, "Review")
.otherwise("Standard")
).select("order_item_id", "packing").orderBy("order_item_id")
result.show()
Scala example
import org.apache.spark.sql.functions._
val result = order_items.withColumn(
"packing", when(col("quantity") >= 3, "Multi-unit")
.when(col("line_total") >= 120, "Review")
.otherwise("Standard")
).select("order_item_id", "packing").orderBy("order_item_id")
result.show(false)
Choose a fallback
Without otherwise, unmatched rows receive null. The example labels quantity one as Single and everything else Other. All quantities here are positive; in a dataset with nulls, a null condition would also take the fallback.
PySpark example
from pyspark.sql.functions import col, when
result = order_items.withColumn(
"quantity_label", when(col("quantity") == 1, "Single").otherwise("Other")
).select("order_item_id", "quantity_label").orderBy("order_item_id")
result.show()
Scala example
import org.apache.spark.sql.functions._
val result = order_items.withColumn(
"quantity_label", when(col("quantity") === 1, "Single").otherwise("Other")
).select("order_item_id", "quantity_label").orderBy("order_item_id")
result.show(false)