ML Engineer MasterClass (October) | 4 seats left

Spark · Built-in functions and UDFs
AmazonAmazon Analytics
Amazon · Tune Spark Workloads

Built-in functions and UDFs

Choose native expressions first and define null-safe UDFs when custom logic is needed.

Step 1 of 6 · Learn

Prefer a built-in expression

Built-in functions operate on Spark columns and give the optimizer more information than opaque custom code. upper changes customer names to uppercase without a Python callback per row. Built-ins are a good first choice when they express the required logic; measure real workloads rather than assuming a universal speedup.

Lesson reference: PySpark and Scala

Prefer a built-in expression

Built-in functions operate on Spark columns and give the optimizer more information than opaque custom code. upper changes customer names to uppercase without a Python callback per row. Built-ins are a good first choice when they express the required logic; measure real workloads rather than assuming a universal speedup.

PySpark example

from pyspark.sql.functions import col, upper

result = customers.select(
    "customer_id", upper("customer_name").alias("label")
).orderBy("customer_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val result = customers.select(
    col("customer_id"), upper(col("customer_name")).alias("label")
).orderBy("customer_id")
result.show(false)

Declare custom behavior

This scalar UDF appends a suffix and returns null for a null input. PySpark declares StringType explicitly; Scala infers the return type from the typed function. Python UDFs can add serialization and execution overhead. This simple operation has a built-in alternative, used in the next pair; avoid external side effects in functions Spark may retry.

PySpark example

from pyspark.sql.functions import col, udf, concat, lit
from pyspark.sql.types import StringType

label_name = udf(lambda value: None if value is None else value + "!", StringType())
result = customers.select(
    "customer_id", label_name(col("customer_name")).alias("label")
).orderBy("customer_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val label_name = udf((value: String) => if (value == null) null else value + "!")
val result = customers.select(
    col("customer_id"), label_name(col("customer_name")).alias("label")
).orderBy("customer_id")
result.show(false)

Replace a UDF with native operations

concat with a literal can express this suffix operation without a custom UDF. For this expression, a null name still produces null. Equivalent output is essential when replacing code for performance: check edge cases and types, not just the common non-null records.

PySpark example

from pyspark.sql.functions import col, udf, concat, lit
from pyspark.sql.types import StringType

label_name = udf(lambda value: None if value is None else value + "!", StringType())
result = customers.select(
    "customer_id", label_name(col("customer_name")).alias("label")
).orderBy("customer_id")
result.show()

Scala example

import org.apache.spark.sql.functions._

val label_name = udf((value: String) => if (value == null) null else value + "!")
val result = customers.select(
    col("customer_id"), label_name(col("customer_name")).alias("label")
).orderBy("customer_id")
result.show(false)
example.pyPySpark
1. Read each namePreserve customer_id and handle missing names.
2. Apply the expressionThe displayed example determines whether it is native or a UDF.
3. Return labelsCompare values and null behavior before changing implementation.
Source customers6 rows
customer_idcustomer_namecity
101AriSeattle
102BoAustin
103CamChicago
104DeeBoston
105EliDenver
106FranPortland
Prefer a built-in expression
Result6 rows
customer_idlabel
101ARI
102BO
103CAM
104DEE
105ELI
106FRAN
The expression transforms each name independently; the customer rows are unchanged.

The example is loaded in the editor. Run it as written, then try a small change.

Runs on the Spark backend. First startup may take a moment.

Run your code to see the result.