ML Engineer MasterClass (October) | 4 seats left

Spark · Filter and sort
AmazonAmazon Analytics
Amazon · Meet Spark

Filter and sort

Explore the product catalog with category filters, combined conditions, and deterministic sorted previews.

Step 1 of 6 · Learn

Filter a category

filter retains rows whose condition is true without removing columns. PySpark uses == for a column equality comparison; Scala uses ===. Here the example selects Electronics products. The original products DataFrame remains unchanged.

Lesson reference: PySpark and Scala

Filter a category

filter retains rows whose condition is true without removing columns. PySpark uses == for a column equality comparison; Scala uses ===. Here the example selects Electronics products. The original products DataFrame remains unchanged.

PySpark example

from pyspark.sql.functions import col, round

result = products.filter(
    col("category") == "Electronics"
).orderBy("product_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = products.filter(
    col("category") === "Electronics"
).orderBy("product_id")
result.show(false)

Combine row conditions

Combine category and identifier conditions with & in PySpark or && in Scala. Parenthesize each comparison. Both conditions must hold. An identifier threshold is a row-selection rule here; it does not imply that a higher identifier means a newer or more valuable product.

PySpark example

from pyspark.sql.functions import col, round

result = products.filter(
    (col("category") == "Office") & (col("product_id") >= 205)
).orderBy("product_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = products.filter(
    (col("category") === "Office") && (col("product_id") >= 205)
).orderBy("product_id")
result.show(false)

Sort before limiting

Exclude Travel products, sort names alphabetically, then limit the preview. product_id is a tie-breaker for equal names. A limit without a sort does not mean the alphabetically first products. Sorting by a name is not a ranking by price or popularity.

PySpark example

from pyspark.sql.functions import col, round

result = products.filter(col("category") != "Travel").orderBy(
    col("product_name").asc(), col("product_id").asc()
).limit(3)
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = products.filter(col("category") =!= "Travel").orderBy(
    col("product_name").asc(), col("product_id").asc()
).limit(3)
result.show(false)
example.pyPySpark
1. Read productsOne row per product.
2. Filter a categoryApply the displayed expression without changing the source DataFrame.
3. Compare the resultElectronics retains product IDs 201 and 204. Other catalog rows remain in the source.
Source products6 rows
product_idproduct_namecategory
201Wireless keyboardElectronics
202Laptop standOffice
203Desk lampHome
204USB-C hubElectronics
205Travel backpackTravel
206Notebook setOffice
Filter a category
Result2 rows
product_idproduct_namecategory
201Wireless keyboardElectronics
204USB-C hubElectronics
Electronics retains product IDs 201 and 204. Other catalog rows remain in the source.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.