Preview customers
A DataFrame stores rows with named, typed columns. customers is already loaded from the Amazon database. Sort by the unique customer_id before limit so the preview is repeatable. show displays rows; it does not modify the source.
Join ML Engineer Interview MasterClass (October Cohort) led by FAANG Data Scientists | Just 4 seats remaining...
ML Engineer MasterClass (October) | 4 seats left
Inspect the customer directory: preview records, read its schema, and count customers.
Step 1 of 6 · Learn
A DataFrame stores rows with named, typed columns. customers is already loaded from the Amazon database. Sort by the unique customer_id before limit so the preview is repeatable. show displays rows; it does not modify the source.
A DataFrame stores rows with named, typed columns. customers is already loaded from the Amazon database. Sort by the unique customer_id before limit so the preview is repeatable. show displays rows; it does not modify the source.
from pyspark.sql.functions import col, round
result = customers.orderBy("customer_id").limit(2)
result.show()import org.apache.spark.sql.functions._
val result = customers.orderBy("customer_id").limit(2)
result.show(false)printSchema describes column types and nullability, not the values of a single record. customer_id is a long identifier. customer_name and city are strings. The example inspects only the identifier; the exercise inspects all three fields. All fields in this directory are non-nullable.
customers.select("customer_id").printSchema()customers.select("customer_id").printSchema()count returns a number, not a DataFrame. The example counts a preview limited to two customers. Counting the full directory includes Fran, even though Fran has no orders. Customer count and order count answer different questions.
customer_count = customers.limit(2).count()
print(customer_count)val customer_count = customers.limit(2).count()
println(customer_count)customers6 rows| customer_id | customer_name | city |
|---|---|---|
| 101 | Ari | Seattle |
| 102 | Bo | Austin |
| 103 | Cam | Chicago |
| 104 | Dee | Boston |
| 105 | Eli | Denver |
| 106 | Fran | Portland |
| customer_id | customer_name | city |
|---|---|---|
| 101 | Ari | Seattle |
| 102 | Bo | Austin |
The example is loaded in the editor. Run it as written, then try a small change.
Runs on the Spark backend. First startup may take a moment.
Run your code to see the result.