ML Engineer MasterClass (October) | 4 seats left

Spark · Columns and expressions
AmazonAmazon Analytics
Amazon · Meet Spark

Columns and expressions

Shape order lines, calculate per-unit values, and export identifiers with deliberate types.

Step 1 of 6 · Learn

Select and rename

select chooses output fields and alias names a column without changing the original table. Each order_items row is one order line. line_total covers every unit on that line; it is not an individual unit price. The example renames it to amount.

Lesson reference: PySpark and Scala

Select and rename

select chooses output fields and alias names a column without changing the original table. Each order_items row is one order line. line_total covers every unit on that line; it is not an individual unit price. The example renames it to amount.

PySpark example

from pyspark.sql.functions import col, round

result = order_items.select(col("order_item_id"), col("line_total").alias("amount")).orderBy("order_item_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = order_items.select(col("order_item_id"), col("line_total").alias("amount")).orderBy("order_item_id")
result.show(false)

Calculate a unit value

Divide each line_total by its quantity to obtain unit_value. withColumn adds the derived expression; round controls its displayed precision. Every quantity here is positive. Real datasets need a policy for missing or zero quantities before division.

PySpark example

from pyspark.sql.functions import col, round

result = order_items.withColumn(
    "unit_value", round(col("line_total") / col("quantity"), 1)
).select("order_item_id", "unit_value").orderBy("order_item_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = order_items.withColumn(
    "unit_value", round(col("line_total") / col("quantity"), 1)
).select("order_item_id", "unit_value").orderBy("order_item_id")
result.show(false)

Cast an identifier

cast("string") exports an identifier as text. The example exports product_id as reference alongside quantity. An order line has several identifiers with different meanings; use order_item_id when the exported reference must identify the line itself. Sort before projecting away the sort key.

PySpark example

from pyspark.sql.functions import col, round

result = order_items.orderBy("order_item_id").select(
    col("product_id").cast("string").alias("reference"), col("quantity")
)
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = order_items.orderBy("order_item_id").select(
    col("product_id").cast("string").alias("reference"), col("quantity")
)
result.show(false)
example.pyPySpark
1. Read order_itemsOne row per order line.
2. Select and renameApply the displayed expression without changing the source DataFrame.
3. Compare the resultNine order lines remain nine rows; selecting columns does not aggregate them.
Source order_items9 rows
order_item_idorder_idproduct_idquantityline_total
11001201159.5
21001202130
31002203299
41002206150
51003204135
610042033149.97
71004205170.02
81005203149.99
910062012120
Select and rename
Result9 rows
order_item_idamount
159.5
230
399
450
535
6149.97
770.02
849.99
9120
Nine order lines remain nine rows; selecting columns does not aggregate them.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.