ML Engineer MasterClass (October) | 4 seats left

Spark · Missing values
AmazonAmazon Analytics
Amazon · Clean and Shape Data

Missing values

Handle missing product categories without silently losing catalog records.

Step 1 of 6 · Learn

Detect nulls

The setup creates prepared_products, keeping categories only for IDs below 204. when without otherwise leaves the other categories null; products stays unchanged. isNull tests missing values and isNotNull tests present values. Neither an empty string nor the text "NULL" is a null.

Lesson reference: PySpark and Scala

Detect nulls

The setup creates prepared_products, keeping categories only for IDs below 204. when without otherwise leaves the other categories null; products stays unchanged. isNull tests missing values and isNotNull tests present values. Neither an empty string nor the text "NULL" is a null.

PySpark example

from pyspark.sql.functions import col, when

prepared_products = products.withColumn(
    "category", when(col("product_id") < 204, col("category"))
)
result = prepared_products.filter(col("category").isNull()).select("product_id").orderBy("product_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val prepared_products = products.withColumn(
    "category", when(col("product_id") < 204, col("category"))
)
val result = prepared_products.filter(col("category").isNull).select("product_id").orderBy("product_id")
result.show(false)

Fill missing categories

na.fill replaces nulls only in the named column and preserves existing categories. Unknown is a placeholder, not an inferred category. The setup is repeated so each example can run independently.

PySpark example

from pyspark.sql.functions import col, when

prepared_products = products.withColumn(
    "category", when(col("product_id") < 204, col("category"))
)
result = prepared_products.na.fill({"category": "Unknown"}).select("product_id", "category").orderBy("product_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val prepared_products = products.withColumn(
    "category", when(col("product_id") < 204, col("category"))
)
val result = prepared_products.na.fill(Map("category" -> "Unknown")).select("product_id", "category").orderBy("product_id")
result.show(false)

Drop missing categories

na.drop restricted to category removes rows missing that field; unrelated columns are not checked. A category-only report may require this, but it excludes three products from the working copy. Here the example previews two retained records.

PySpark example

from pyspark.sql.functions import col, when

prepared_products = products.withColumn(
    "category", when(col("product_id") < 204, col("category"))
)
result = prepared_products.na.drop(subset=["category"]).orderBy("product_id").limit(2)
result.show()

Scala example

import org.apache.spark.sql.functions._

val prepared_products = products.withColumn(
    "category", when(col("product_id") < 204, col("category"))
)
val result = prepared_products.na.drop(Seq("category")).orderBy("product_id").limit(2)
result.show(false)
example.pyPySpark
1. Read prepared_productsSetup leaves three category values null in a separate DataFrame.
2. Detect nullsApply the displayed expression to the source records.
3. Inspect the resultThe source shown is the prepared copy after setup, not the unchanged products table.
Source prepared_products6 rows
product_idproduct_namecategory
201Wireless keyboardElectronics
202Laptop standOffice
203Desk lampHome
204USB-C hubNULL
205Travel backpackNULL
206Notebook setNULL
Detect nulls
Result3 rows
product_id
204
205
206
The source shown is the prepared copy after setup, not the unchanged products table.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.