Interactive data engineering course
Apache Spark with PySpark and Scala
Learn Apache Spark through 46 hands-on lessons in PySpark and Scala. Analyze orders and products with DataFrames, joins, window functions, and streaming.
What will you learn?
Start with selecting columns, filtering rows, and cleaning data. Build reports with aggregates, joins, and window functions. Then read execution plans, handle data skew, and work with streaming windows, watermarks, and checkpoints.
How does practice work?
Each lesson alternates between three worked examples and three exercises. Read the explanation, trace the tables, then run the example. The next step asks you to adapt it. Switch between PySpark and Scala throughout the course.
The Amazon analysis scenario uses four related tables: orders, order_items, products, and customers. DataInterview produces this course independently; it is not affiliated with Amazon or the Apache Software Foundation.
What should you know first?
Basic Python or Scala: variables, function calls, and comparisons. Prior Spark experience is not required. SQL knowledge helps with joins and aggregation but is not a prerequisite.
What is free?
Every lesson is readable. Sign in to run code and use hints and solutions in Your first DataFrame, Columns and expressions, Group by, and Inner joins. Premium unlocks practice in all other lessons.
Apache Spark curriculum
Meet Spark
- 1Your first DataFrameInspect the customer directory: preview records, read its schema, and count customers.
- 2Columns and expressionsShape order lines, calculate per-unit values, and export identifiers with deliberate types.
- 3Filter and sortExplore the product catalog with category filters, combined conditions, and deterministic sorted previews.
- 4Transformations and actionsBuild a customer-directory plan and execute it with show or count.
Clean and Shape Data
- 5Missing valuesHandle missing product categories without silently losing catalog records.
- 6Conditional columnsLabel order lines with explicit boundaries, branch precedence, and fallbacks.
- 7Strings and patternsClean catalog names, extract SKU parts, and filter product-name prefixes.
- 8Dates and timestampsParse customer signup text, filter calendar dates, and calculate days since signup.
- 9DuplicatesSeparate replayed order lines from legitimate repeat purchases and audit their counts.
Aggregate and Analyze
- 10Aggregate functionsCount order lines and distinct products, then summarize line values.
- 11Group bySummarize order lines by product, order, or both keys.
- 12Filter grouped resultsDistinguish row filters from filters on aggregate values.
- 13Pivot tablesCompare products across line quantities using explicit pivot columns.
- 14Multiple grouping levelsBuild product and order subtotals with rollup and cube.
Combine DataFrames
- 15Inner joinsEnrich order lines with names and categories from the product catalog.
- 16Left and full joinsPreserve unmatched rows and recognize nulls introduced by a join.
- 17Semi and anti joinsCheck whether a match exists without adding right-side columns.
- 18Join cardinalityTrace repeated keys, count matches, and handle duplicate lookup rows deliberately.
- 19Union and schema alignmentStack rows safely, align fields by name, and handle missing columns.
Work with Nested Data
Window Functions
- 24Partition and orderDefine which rows a window sees and the sequence used within each group.
- 25Rank and deduplicateHandle ties, choose one row per key, and return deterministic top results.
- 26Previous and next rowsRead neighboring rows and calculate differences with explicit boundary handling.
- 27Running totalsAccumulate values with explicit row frames and partition-specific resets.
- 28Window framesChoose which neighboring rows or ordering values contribute to a window calculation.
How Spark Executes
- 29Partitions and tasksInspect physical partitions and distinguish them from rows, task slots, and complete jobs.
- 30Narrow and wide transformationsIdentify local row operations and cross-partition exchanges in a transformation pipeline.
- 31ShufflesRecognize when rows must move and reduce unnecessary data before an exchange.
- 32Read an execution planInspect live Spark plans without treating node numbers or optimizer choices as fixed.
- 33Repartition and coalesceChoose how to change partition counts without confusing fewer tasks with faster execution.
Tune Spark Workloads
- 34Broadcast joinsUse small lookup tables without assuming every join should broadcast.
- 35Data skewMeasure uneven keys and understand where a two-stage aggregation can help.
- 36Cache and persistReuse a DataFrame across actions and release storage deliberately.
- 37Built-in functions and UDFsChoose native expressions first and define null-safe UDFs when custom logic is needed.
- 38Adaptive executionInspect runtime plan adaptation without relying on fixed join strategies or partition counts.
Read and Write Data
- 39Files and schemasRead JSON and CSV with deliberate schemas while keeping temporary-file examples self-contained.
- 40Partition pruning and column selectionInspect directory partition filters and the columns requested by a Parquet scan.
- 41Write modesUnderstand append, overwrite, and ignore using disposable paths only.
- 42Small filesControl tiny teaching writes and understand why production file sizing needs measurement.
Structured Streaming
- 43Streaming DataFramesTransform a finite file stream and manage the lifecycle of a teaching query.
- 44Event-time windowsGroup events by their timestamps using tumbling and sliding windows.
- 45WatermarksRelate event-time progress, accepted late events, and finalized append-mode windows.
- 46Checkpoints and recoveryResume bounded file streams, verify committed output, and distinguish recovery from replay.