ML Engineer MasterClass (October) | 4 seats left

Spark · Your first DataFrame
AmazonAmazon Analytics
Amazon · Meet Spark

Your first DataFrame

Inspect the customer directory: preview records, read its schema, and count customers.

Step 1 of 6 · Learn

Preview customers

A DataFrame stores rows with named, typed columns. customers is already loaded from the Amazon database. Sort by the unique customer_id before limit so the preview is repeatable. show displays rows; it does not modify the source.

Lesson reference: PySpark and Scala

Preview customers

A DataFrame stores rows with named, typed columns. customers is already loaded from the Amazon database. Sort by the unique customer_id before limit so the preview is repeatable. show displays rows; it does not modify the source.

PySpark example

from pyspark.sql.functions import col, round

result = customers.orderBy("customer_id").limit(2)
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = customers.orderBy("customer_id").limit(2)
result.show(false)

Read the customer schema

printSchema describes column types and nullability, not the values of a single record. customer_id is a long identifier. customer_name and city are strings. The example inspects only the identifier; the exercise inspects all three fields. All fields in this directory are non-nullable.

PySpark example

customers.select("customer_id").printSchema()

Scala example

customers.select("customer_id").printSchema()

Count customers

count returns a number, not a DataFrame. The example counts a preview limited to two customers. Counting the full directory includes Fran, even though Fran has no orders. Customer count and order count answer different questions.

PySpark example

customer_count = customers.limit(2).count()
print(customer_count)

Scala example

val customer_count = customers.limit(2).count()
println(customer_count)
example.pyPySpark
1. Read customersOne row per customer.
2. Preview customersApply the displayed expression without changing the source DataFrame.
3. Compare the resultThe first two customer IDs are 101 and 102. A preview does not remove the other customers from the source.
Source customers6 rows
customer_idcustomer_namecity
101AriSeattle
102BoAustin
103CamChicago
104DeeBoston
105EliDenver
106FranPortland
Preview customers
Result2 rows
customer_idcustomer_namecity
101AriSeattle
102BoAustin
The first two customer IDs are 101 and 102. A preview does not remove the other customers from the source.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.